Historical failure and recovery
Historical evidence is valuable because it shows the gate refusing an inadequate claim. It also makes a correction reviewable: a new passing run does not erase the failed predecessor. These records concern 1–2 October 2026 candidates and runtimes, separately from the fresh October 3 release.
There are three evidence levels here: selected public Kargo-managed bundles, narrower public manual boundary records, and historical campaign summaries whose detailed bundles remain private. The available artifact determines what a reader can independently inspect.
A runner failure followed by fresh re-verification
Section titled “A runner failure followed by fresh re-verification”The public failed predecessor is bound to AnalysisRun staging.01m3w75z7m23m5jjefxxwg05e9.5664b01 and gate Job 7caf8b2e-c494-42a3-bcbc-48833f9573cb.chaos-verdict.1. The earlier runner could not obtain the Workflow timestamps required to score the three experiments. It recorded failure and entered run-scoped cleanup. This is an evidence/orchestration failure, not a passing fault experiment.
The fresh re-verification used AnalysisRun staging.01m3x2psvxxngj38j0gnkch5rv.5664b01, gate Job e1f39b47-8433-49bd-a05c-d0f469750568.chaos-verdict.1, and Workflow chaos-gate-2gfdx. Source f156b076ce090594529ffe735b9da561e6e9a2f8, rendered revision 3781f8a5206c660485d0ef2a214663d104eaded5, and application digest sha256:a5beda836336de1677e27c3c2ff09a1ab27d9366040cd985c347d700de6f476b remained the recorded candidate identity; the gate-runner digest changed to the corrected runtime.
All three new scorecards passed: PostgreSQL had six checks, Redis seven, and signer six. The log recorded cleanup and PASS. Public metadata, logs, and scorecards make this pair more directly inspectable than a summary alone. It demonstrates corrected verification for the same staging candidate, not a new app release or prod promotion.
An earlier public passing scorecard
Section titled “An earlier public passing scorecard”The 1 October staging pass belongs to AnalysisRun staging.01m3w93zc6ddewce35bzy8d2t4.5664b01 and Workflow chaos-gate-fkh9p. It binds the same earlier application digest, not October 3’s ab88d89… release.
Its public Redis scorecard records injection at 2026-10-01T17:49:47Z, a 60-second fault, and 60-second settling period. Non-probe traffic was 293.71428571428567 against > 50; the miss/fallback counter increased by 98.28571428571428; maximum scored redirect p95 was approximately 0.32 seconds against < 1.2; app restarts were zero. The readiness checks separately observed Redis unavailable and recovered.
The fractional counters are Prometheus increase extrapolations, not literal fractional requests. The cache counter alone is not an outage detector: normal misses also increment it. The score’s fault interpretation comes from the combined outage/recovery, traffic, latency, and restart checks in the bounded window. Redis disappears explains that conjunction.
A degraded pipeline candidate was refused before chaos
Section titled “A degraded pipeline candidate was refused before chaos”The historical full campaign summary names regression-blocked-20261002-050844, source 86d6f17f20a53de34c2d81949016fdab1fd82d16, and rendered revision c6515a622eeae3850e749cae1833801827d3c915. It traveled through the real Kargo staging path. The runner acquired the Lease and created the load Job but could not prove paid traffic within 75 seconds, so it failed before fault injection and cleaned the Job.
That release outcome was correctly fail. The negative test succeeded by preventing a vacuous chaos pass. It does not demonstrate a scored outage, and its private bundle is not available for public recomputation.
The corrected recovery-20261002-052031 used source 0a98c08ccab5fa7b05efcecf877c714e5c89f025 and rendered revision bc388eaff8ff294e1f59b3c417a18c6b96dd5d39. It completed all three faults, recovery scores, and cleanup. The later historical prod smoke used Freight 89773faa71018568b4eef88317856536430a58b3, rendered prod revision 4d7ffb5a423cb80f49e5d84e97d7619813482439, and the earlier a5beda… app digest. Those successful observations do not transfer to a different Freight.
Manual boundary records have narrower meaning
Section titled “Manual boundary records have narrower meaning”| Public record | What was demonstrated | What was not reached |
|---|---|---|
| No target | Direct runner blocked an intentionally nonexistent Service before a vacuous fault; Lease cleanup was logged. | Workflow, traffic, fault, scoring, Kargo verification |
| Missing load source | Direct runner created a Workflow, failed on an intentionally absent CronJob, and entered cleanup. | Normal load startup, paid traffic, fault, score |
| Unreachable telemetry | Scorer-only test against http://127.0.0.1:1 emitted five transport-error check failures and a failing verdict. |
Shared Prometheus outage, Workflow/load/fault, promotion |
| One-second deadline | Direct runner failed before normal baseline/fault phases and exact-object cleanup was verified. | Default timeout, natural Chaos Mesh timeout, useful paid traffic, scored recovery |
Manual direct records are not Freight promotion or release approval. Their strength is showing a supplied boundary failure closes the relevant path. Source tests cover additional fixtures, but a fixture is not another retained live campaign.
The verification report, canonical evidence index, and manifest preserve chronology and availability. Compare their records with evidence interpretation before using a historical screenshot to support a new claim.
Maintained by Satyam Agnihotri · DevOps & Cloud Engineer