Dependencies and fault boundaries
A dependency’s importance is determined by what the request needs it to do. The URL shortener cannot create a durable mapping without PostgreSQL. It can resolve a known mapping without Redis if PostgreSQL remains available. The paid load client needs the signer to obtain new authorizations, while the application accepts those authorizations without calling the signer itself.
This classification helps define the expected behavior during a fault. Asking every dependency to disappear without any visible consequence would contradict the application’s actual correctness requirements.
Essential storage and optional acceleration
Section titled “Essential storage and optional acceleration”PostgreSQL owns the mappings and payment audit association. Both new creation and existing-destination lookup use it. A redirect requires it after a cache miss, although a valid cached destination can be returned without a database query. That cache hit does not make PostgreSQL optional for the application overall: /ready still checks the database on every call.
Redis stores url:<code> with a configured expiry, defaulting to 3600 seconds. Creation does not write Redis; the first database-backed redirect populates it. The cache is a performance optimization, and failures in its get/set boundary are converted into CacheUnavailable so routes can use PostgreSQL. A missing entry and an unavailable cache both increment url_shortener_cache_misses_total.
There is a lifecycle condition. Before the process has ever passed /ready, Redis must respond to a bounded ping. After that first successful readiness, the route stops checking Redis and continues to require PostgreSQL. Consequently, “Redis is optional” describes steady-state redirect service; it does not describe a cold application startup. The health page explains the trade-off.
Payment dependencies live on two sides of the request
Section titled “Payment dependencies live on two sides of the request”The signer belongs to the client authorization path. It keeps payer key material, verifies its configured chain and bootstraps Permit2 allowance at startup, then creates signatures locally. Request-time signing uses a random Permit2 nonce and performs no RPC calls. An RPC outage after successful bootstrap therefore has a different scope from an unavailable signer pod.
The facilitator belongs to the application’s payment path. The application sends server-owned terms and the submitted envelope to /verify, then /settle. A payment-enabled new creation cannot return 201 unless settlement validates and persistence succeeds. The application does not hold payer keys or submit settlement transactions directly.
| Dependency or owner | Caller | Needed for | Defined failure boundary |
|---|---|---|---|
| PostgreSQL | Application | Durable creation, destination reuse, uncached resolution, readiness. | Controlled database 503; cached redirects may still work. |
| Redis | Application | Initial readiness and accelerated redirects. | Initial readiness can fail; established process falls back to PostgreSQL. |
| Signer | Paid client/load generator | New Permit2 authorizations. | Paid iterations can stop before reaching /shorten; app is not its direct caller. |
| RPC provider | Signer bootstrap | Chain verification, balance/allowance checks, optional approval. | Startup or wallet slot may remain degraded; hot-path signing does not query RPC. |
| Facilitator | Application | Verification and settlement of a new paid creation. | Transport/HTTP/JSON availability errors become 503; rejected or structurally invalid object outcomes currently become 402. |
| Prometheus | Gate scorer | Measurements for the verdict. | Insufficient usable evidence prevents a justified pass. |
A fault boundary is also an ownership boundary
Section titled “A fault boundary is also an ownership boundary”The three retained PodChaos definitions select one staging primary database, Redis master, or signer pod using namespace and workload labels. They do not target the application, every dependency instance in the cluster, or production-like resources. The narrow selector helps attribute the result to one hypothesis.
This is bounded blast radius in a shared lab, rather than separate infrastructure for every environment. A dependency outage may increase work elsewhere: Redis loss sends more reads to PostgreSQL, and retries or failed paid iterations change the client traffic mix. The Redis workflow explains why capacity and latency must be measured along with successful fallback.
Use the boundary to avoid the wrong diagnosis
Section titled “Use the boundary to avoid the wrong diagnosis”If paid creations stop while redirects continue, inspect signer/client outcomes before concluding the application has isolated every dependency. If /livez returns 200 while /ready returns 503, the process may be behaving exactly as designed. If cache misses increase, look for target outage measurements before labeling that increase a Redis incident.
The implemented boundaries are in application wrappers and routes, payment adapter, signer service, and load client. Compare them with application design and bounded experiments to connect request correctness with release verification.
Maintained by Satyam Agnihotri · DevOps & Cloud Engineer