Skip to content

Signer disappears

A signer outage can make application metrics look quiet precisely because the client stopped sending paid creation requests. The application does not call the signer during /shorten; the load client does. This failure scenario tests isolation and signer recovery, and its interpretation must preserve that distinction.

The bounded hypothesis is that a selected staging signer pod becomes unavailable for 60s, the application avoids restarts and /shorten 5xx responses attributable to the observed run, no application replay increment is observed, and the signer target recovers. That is narrower than a guarantee of uninterrupted paid throughput.

The signer owns up to three configured payer wallet slots. At startup it validates settings, verifies the configured RPC chain, reads balance and Permit2 allowance, and sends an approval transaction only for slots whose allowance is below the configured payment amount. A successful slot remains available even when another slot’s bootstrap fails.

Service readiness is true when at least one slot has bootstrapped. The paid load setup imposes a stronger condition: enough configured and bootstrapped slots for its virtual users, plus cached balance sufficient for the bounded run. A 200 /health by itself therefore does not establish that a three-user payment workload is ready.

POST /sign-permit2 selects a slot, enforces the configured amount and allowed deadline, and creates a local EIP-712 signature. Its random uint256 Permit2 nonce needs no request-time RPC lookup. Signing does not settle a payment. An RPC outage after successful startup is therefore not equivalent to taking away the signer service itself.

/health and /wallets return boot-cached state rather than querying balances or allowances anew. They do not continuously verify funds, external chain availability, or successful settlement. Signer API tests verify that health, wallet inspection, and signing make no additional RPC calls after bootstrap.

Before fault injection, the current run’s load Job completes paid creation and redirect and emits the safe readiness marker. Its configuration prechecks the app and signer. This proves useful initial work without publishing a wallet, signature, payment envelope, or transaction identifier.

During a paid iteration, the load client calls /sign-permit2 before /shorten. If the signer response is not 200, or its JSON lacks the required authorization fields, the client marks signing, settlement, and end-to-end success as false and returns. It does not send an application creation request for that iteration.

Consequently, a zero increase in app /shorten errors can coexist with unsuccessful paid traffic. The same is true for a low payment replay counter: no creation request may have reached the replay boundary. These metrics are evidence about application isolation under the supplied workload, not automatic proof of sustained payment service.

The load source also offers an optional independent redirect probe. If configured with an existing short URL, it can keep exercising resolution while signing is unavailable. Its independent_redirect_ok_rate provides a different observation from paid end-to-end success. The scenario cannot assume the probe was enabled in every historical run; check the actual run configuration and summaries.

Apply one fault and inspect the six checks

Section titled “Apply one fault and inspect the six checks”

The signer fault selector requires staging namespace plus stable signer workload labels. Its mode: one and 60s duration bound the requested fault. The deployment separates the key-bearing process from the app and load generator.

The scorer requires an observed signer readiness minimum of zero and a final readiness value of one. It also requires application restart increase, app /shorten 5xx increase, and application payment replay increase each below 0.5. Release identity validation is the sixth check.

Unlike PostgreSQL and Redis checks, signer checks contain no meaningful-traffic rule for the fault window. The three low-counter queries also include or vector(0). It would therefore be incorrect to claim the signer card independently proves continuous paid traffic or that every missing error series causes failure. Pre-injection paid readiness and strict measurement validity are useful protections, but they do not add a missing per-window throughput contract.

For a broader operational assessment, compare target readiness with sign_success_rate, end_to_end_success_rate, signing duration, app request volume, and any independently configured redirect success. Those describe separate client and server events. The repository’s gate thresholds remain unchanged; this explanation makes their implemented scope clear.

When a replacement signer process starts, bootstrap runs again. It can recognize an existing sufficient allowance without submitting another approval. A failed chain verification prevents approval attempts, and failed slots stay unavailable. Bootstrap is idempotent for successful slots, but there is no HTTP retry-bootstrap endpoint or continuous background bootstrap loop in the service.

The retained historical signer scorecard records passing checks in its 1 October window and binds the same candidate as that gate’s other cards. It observes a target outage followed by final readiness, zero observed application restarts, zero observed /shorten 5xx, and zero observed application replay increments. It does not add a fault-window paid-success measurement that its schema did not collect.

Bootstrap tests cover allowance reuse, partial bootstrap, chain verification failure, and missing configuration without opening RPC. Permit2 tests recover signatures and verify binding to payment terms. Signer scoring tests exercise its exact rules. These are credential-free mechanism checks, separate from the historical live record.

Continue with the paid request, application design, or measurements. Isolation protects component responsibility; measuring useful work is what tells you how much service survived that isolation boundary.

Maintained by Satyam Agnihotri · DevOps & Cloud Engineer