Inside the gate
The gate is one bounded Kubernetes Job that coordinates load, fault injection, scoring, and cleanup for an identified candidate. Its main protection against a misleading pass is refusing to treat missing work, missing faults, missing measurements, or an unfinished cleanup as successful verification.
Enlarge diagram: Gate lifecycle and failure exits · Version-controlled diagram source
Preconditions precede fault injection. Once objects exist, every normal exit and handled signal enters cleanup; scoring and cleanup have distinct effects on the final Job outcome.
Validate the candidate and the time budget
Section titled “Validate the candidate and the time budget”The AnalysisTemplate supplies RELEASE_REVISION from Freight’s source/chart commit and RELEASE_DIGEST from its selected application image. The runner rejects an empty or multiline revision, a digest other than lowercase sha256: plus 64 hexadecimal characters, an invalid run ID, unreadable workflow, or unsupported timeout settings.
This happens before acquiring a Lease or starting traffic. The runner’s initial RUN_ID comes from the gate Pod UID. After workflow creation, the generated workflow name becomes the scorecards’ durable run identity. The load Job retains the initial run ID in its exact name and labels.
Budgets reserve time for the complete lifecycle, not merely the fault nodes:
| Boundary | Current default | Validation or purpose |
|---|---|---|
| Load startup | 75 seconds | Must not exceed 90 seconds; each status/log probe is separately bounded. |
| Workflow observation | 660 seconds | Cannot exceed the workflow’s hard deadline. |
| Each scorer | 105 seconds | Must cover up to six 15-second Prometheus requests. |
| Optional annotation | 15 seconds | A diagnostic action that cannot change the verdict. |
| Cleanup | 120 seconds | One shared deadline across exact deletions and readback. |
| Enclosing Job | 1320 seconds | Must cover startup, workflow, all scorers, annotation, cleanup, and reserve. |
| Lease | 1470 seconds | Must outlast the Job budget by at least 150 seconds. |
The minimum runner budget at current defaults is 1260 seconds; the Job allows 1320. It has no retry (backoffLimit: 0) and allows 120 seconds for Pod termination grace. Cleanup budget tests check these contracts.
Acquire the right to run
Section titled “Acquire the right to run”The runner atomically creates a Lease in resilience-gate. If it exists, it reads holder, renewal time, duration, and resource version. A non-expired or incomplete Lease blocks the run. An expired Lease may be replaced only using its observed resource version, so a concurrent change prevents an unconditional takeover.
This is a fixed-duration lock; the script does not continually renew it. The duration is chosen to cover the bounded Job and cleanup buffer. It serializes runners following this contract; it is not a cluster-wide ban on unrelated experiments.
The runner then lists labeled gate workflows in staging and refuses to continue if an earlier one remains. It does not sweep the old object away. That leftover is an investigation boundary rather than permission to inject another fault.
Confirm a ready target and useful paid work
Section titled “Confirm a ready target and useful paid work”An EndpointSlice query requires at least one ready address for the configured target Service. A missing ready target stops before the load Job or workflow exists. This tests target availability, not each dependency independently.
The runner creates loadgen-<RUN_ID> from the suspended staging CronJob, labels it, and waits for evidence. A merely active Job is insufficient. The exact Job must emit RESILIENCE_GATE_PAID_TRAFFIC_READY v1; load-generator source emits it only after paid creation returns 201 and the subsequent redirect returns 302. The runner then rechecks Job status so a just-terminated generator cannot pass the startup handoff.
Each status and log probe recomputes its remaining allowance against the startup deadline. Failed, completed, or unconfirmed Jobs stop startup. No raw payment body, signature, or header is echoed by the marker filter. Orchestrator tests cover missing targets, absent markers, terminal load, and the final workflow/load race.
This is a startup proof. During chaos, polling establishes that the Job remains active; it does not continuously establish successful paid requests. The signer outage can prevent fresh signing and cause the client to return before calling the application. That limitation is part of interpreting the signer score, not a reason to invent traffic continuity.
Create the serial workflow and observe actual application
Section titled “Create the serial workflow and observe actual application”The workflow runs baseline 90 seconds, PostgreSQL fault 60 and recovery 120, Redis fault 60 and recovery 60, then signer fault 60 and recovery 60. Each PodChaos action uses mode: one and staging namespace selectors. Its parent deadline is 660 seconds.
The runner polls the exact workflow’s conditions while confirming the load Job remains active. Failed=True fails the observation; Accomplished=True is accepted only after the final load-status check. An observation timeout also fails.
Before each conditions check, the runner reads workflow-node references to typed PodChaos children and preserves the earliest successful Apply event for each dependency. It deliberately does not use WorkflowNode startTime, which represents node scheduling. Capturing events while the workflow runs survives terminal child-resource removal. The historical failed predecessor and fault-retention test show why this distinction matters.
Score what can be scored, retaining failures
Section titled “Score what can be scored, retaining failures”After the workflow observation phase, the runner collects available injection timestamps again and attempts each experiment scorer. It does this even when workflow observation failed, allowing available diagnostics to remain visible. A missing successful Apply timestamp fails that experiment and skips its invocation. A timeout or failed scorer adds another failure.
Prometheus queries occur here. There is no Prometheus preflight before chaos creation. Every scorer uses a 60-second fault duration and 60-second scored settle period, plus 20 seconds of range lead. PostgreSQL’s longer workflow recovery pause does not enlarge its score window. Metrics and scoring provides exact checks and the deliberate or vector(0) exceptions.
Each command emits structured JSON and writes a scorecard in /results. That directory is an ephemeral Job volume, so long-term evidence requires collection or recovery from the emitted sanitized log. A scorecard is not automatically a durable public artifact.
Annotate, report, then finish cleanup
Section titled “Annotate, report, then finish cleanup”Optional Grafana annotation follows scoring. Missing credentials, missing helper, transport error, or annotation timeout are diagnostic failures only. The main function then logs GATE VERDICT: PASS or FAIL and returns its result.
The exit trap still runs. It deletes the exact workflow and its labeled PodChaos children, the named load Job, and its Lease, then checks absence or empty lists within the shared cleanup deadline. Cleanup failure forces nonzero exit, including after a PASS log. Thus a PASS line is an intermediate verdict, not standalone proof that the Kubernetes Job succeeded.
One implementation limit deserves precision: verify_absent treats any failing single-object kubectl get as absence rather than distinguishing NotFound from other read errors. Other cleanup calls can fail the cleanup path, but the helper itself does not establish error-independent absence. Independent post-run readback and retained cleanup receipts strengthen the recorded conclusion.
Handled INT and TERM signals use the trap; abrupt process termination cannot be claimed to run shell cleanup. The Job/Lease budgets reduce normal deadline overlap, and a residual workflow blocks the next cooperating run. Continue with score to promotion for the enclosing AnalysisRun and eligibility decision.
Maintained by Satyam Agnihotri · DevOps & Cloud Engineer