Troubleshoot the failed layer
A gate failure is actionable only when its layer is clear. Missing target endpoints, absent useful traffic, unproven injection, bad telemetry, a breached threshold, and failed cleanup all block a release for different reasons. Preserve the exact run’s artifacts before retrying; a fresh green dashboard cannot explain an earlier failed window.
Use the read-only verifier first and retain sanitized status, logs, scorecards, identities, and cleanup observations. Never read Secret data to explain a workload status, publish payment headers, broaden load or fault bounds, or patch controller-owned branches to make a result look successful.
Follow the earliest failed responsibility
Section titled “Follow the earliest failed responsibility”| Layer | Observation to inspect | Interpretation and next step |
|---|---|---|
| Context and prerequisites | Exact context, configured namespaces, Stage/templates, Ready secret-store signals | A mismatch or absent contract prevents a meaningful request. Fix reviewed configuration before proceeding. |
| Candidate and reconciliation | Freight digest/source, rendered environment commit, Promotion, Argo CD sync/health | A rendered commit is not deployed merely because it exists. Resolve controller/repository access or workload readiness separately. |
| Target | Service and Ready EndpointSlice address before injection | No target means the experiment would be vacuous. Do not count an absent pod as a successfully applied fault. |
| Load source and startup | Suspended source CronJob, exact run Job/Pod, setup checks, traffic preflight | Source creation failure differs from startup timeout or failed paid traffic. Inspect signer/app/facilitator layers with sanitized logs. |
| Workflow and injection | Workflow/WorkflowNode conditions and applied start/end timestamps | Scheduled intent does not prove the fault reached its target. Missing timestamps prevent trustworthy scoring. |
| Telemetry | Query transport, result shape, timestamps, sample count, finite observations | Unavailable or invalid evidence cannot justify passing the associated check. Preserve the error reason. |
| Score | Expression, exact window, aggregation, operator, threshold, observed value | A real finite threshold breach is a behavior failure; distinguish it from evidence failure. |
| Cleanup | Exact Workflow, WorkflowNode/PodChaos, run load Job, source suspension, Lease release | Passing experiment scores can still yield a failed enclosing Job when cleanup is incomplete. |
| Downstream policy | Successful verification attached to Freight, manual promotion request | A healthy app or approved Freight does not establish normal upstream verification. |
Investigate in this order because later absence may be a consequence of an earlier stop. The no-target boundary record intentionally contains no scored fault; searching for a missing scorecard there does not reveal a second independent failure.
Useful traffic is its own check
Section titled “Useful traffic is its own check”The gate waits for paid creation traffic before fault injection. A degraded app or signer can make a load Job exist without accomplishing useful work. The historical failed candidate reached the real Kargo path but failed the 75-second paid-traffic startup requirement before chaos. Its correct release outcome was failure; the negative-test objective succeeded.
During signer loss, the app can remain healthy while the client cannot obtain a signature for a new paid request. Read signer availability, attempted/completed paid iterations, and replay/settlement metrics alongside app errors. Likewise, ordinary cache misses are not proof of Redis failure. Correlate errors and dependency readiness with the actual injection window.
Distinguish no data from zero
Section titled “Distinguish no data from zero”The scorer records reason, evidence_error, and sample_count for each check. A transport failure, empty series, stale timestamp, insufficient range coverage, and non-finite value are different evidence defects. The scorer’s deliberately chosen or vector(0) expressions allow zero when no matching error series exists. They do not imply all absent series pass, nor that every missing series always fails.
Read measurement semantics and the metrics reference before reinterpreting a query. Changing thresholds, adding zero defaults, or widening windows would change the release contract and is outside a documentation fix.
Read the enclosing outcome
Section titled “Read the enclosing outcome”The runner gathers applied timestamps, scores each bounded experiment, and runs cleanup. The scorecards explain measurement decisions; the gate Job exit status includes orchestration and cleanup. Kargo then records the AnalysisRun result. Preserve all three layers. A PASS line copied from one scorecard cannot override a later cleanup error, and request acceptance cannot override a failed AnalysisRun.
The retained timestamp failure and re-verification illustrate why this separation matters: the earlier runner could not obtain required workflow timestamps and failed closed; the corrected runner produced a fresh successful run without editing the old record.
The primary source is orchestrate.sh and score_experiment.py. Relevant tests include target/orchestration, cleanup, and no-data behavior. After diagnosis, choose fresh verification and reviewed collection, with new identities and a new run ID.
Maintained by Satyam Agnihotri · DevOps & Cloud Engineer