Healthy is not resilient
A healthy deployment tells you that the current conditions support service. Resilience describes a behavior across changing conditions: a fault occurs, the system degrades according to a defined contract, and recovery becomes observable. A green readiness check before a fault cannot establish what happens during it.
Resilience Gate therefore treats health as a prerequisite for an experiment rather than its verdict. The application and selected dependencies must be ready, useful traffic must reach the correct candidate, and the runner must own a bounded run. Only then can a dependency failure produce an attributable answer.
Three hypotheses, three different answers
Section titled “Three hypotheses, three different answers”The application’s dependencies support different hypotheses. When PostgreSQL fails, durable URL creation and uncached resolution cannot continue correctly. The appropriate behavior is a controlled availability failure and eventual readiness recovery, while process liveness stays independent of storage. When Redis fails after first readiness, redirects should fall back to PostgreSQL with bounded latency. When the signer fails, the application should remain isolated from that service, but the paid client’s ability to obtain fresh signatures is affected.
These hypotheses cannot share a generic “no errors” requirement. A database fault is allowed to cause expected 503 responses. A cache fault tests successful fallback and the cost of additional database work. A signer fault may cause client iterations to stop before any /shorten request. An unchanged application error counter in that case says little about paid throughput.
The three failure workflows show the concrete distinctions. The failure model gives the broader decision categories: failed, blocked, unavailable, and not collected are all different from a pass.
Build the conclusion from several observations
Section titled “Build the conclusion from several observations”A sound experiment needs evidence for several separate propositions:
| Proposition | Relevant observation | Why another signal cannot replace it |
|---|---|---|
| The intended fault reached its target. | Applied workflow state plus target readiness falling in the window. | A scheduled manifest does not prove injection. |
| Requests exercised useful behavior. | Run-scoped paid creation/redirect marker and appropriate traffic measurements. | Probe responses and idle counters do not exercise creation or redirect. |
| The application kept its contract. | Expected status behavior, fallback and latency checks, and restart measurements. | Process liveness alone does not prove successful service. |
| The dependency recovered. | Final readiness samples within the bounded window. | Seeing a single outage sample establishes failure, not recovery. |
| The run ended safely. | Run-scoped cleanup and absence checks, enclosing Job result. | A passing scorecard cannot establish cleanup. |
Each measurement also needs a trustworthy collection path. Empty or stale telemetry can make a seemingly good value meaningless. The scorer validates response shapes and sample quality before applying the configured check. Some expressions deliberately use or vector(0) to represent an absent error series as zero, so even strict validity does not mean every absent raw series causes failure. Read trustworthy measurements before interpreting a quiet graph.
Bind the behavior to the candidate
Section titled “Bind the behavior to the candidate”A resilience conclusion belongs to the source, rendered manifests, deployed image, supporting runtime identities, and run that produced it. A later image with the same friendly tag inherits none of that evidence. Similarly, a historical Redis pass cannot establish the status of a fresh release simply because the same workflow filename is still present.
The public historical records retain explicit windows and immutable identity. The later release’s verification report distinguishes its own deployment and observations. The evidence interpretation guide explains how to connect those records without merging their dates or claims.
What this design deliberately leaves open
Section titled “What this design deliberately leaves open”The gate is a finite contract for an owned testnet lab. Its sequential 60s pod-failure scenarios are useful controlled probes of the application and verification system. They do not establish resilience to concurrent dependency loss, a persistent-volume disaster, long outages, every facilitator failure, or region failure. The lab also shares cluster resources across environment namespaces.
The value of the design is the ability to explain why one bounded result supports one release decision, and where additional evidence is required. That is stronger than treating every available metric as proof of universal resilience. Inspect orchestration, scoring, and gate tests for the implemented decision boundaries.
Maintained by Satyam Agnihotri · DevOps & Cloud Engineer