Skip to content

Follow one release

The release journey begins with a candidate, not a deployment. The question is whether a particular application digest, paired with a particular chart-source revision, has earned normal eligibility to move into the next environment. Publication, reconciliation, verification, and promotion are separate handoffs with separate owners.

Release identities and promotion handoffs

Enlarge diagram: Release identities and promotion handoffs · Version-controlled diagram source

The chart source and application image meet in Freight. Each Stage produces its own rendered commit; verification grants downstream eligibility while staging and prod-like moves remain explicit.

1. Validate and publish the application bytes

Section titled “1. Validate and publish the application bytes”

The application workflow validates pull requests and builds an image without publishing. Its trusted-main publish job depends on that validation. A push or manual dispatch on main can receive short-lived Google credentials through GitHub OIDC and publish to the configured Artifact Registry repository.

Docker’s build action supplies the emitted digest. The workflow signs repository@digest with keyless Cosign and verifies the signature against the exact GitHub workflow identity and issuer. It then uploads a verified identity JSON record. The source commit, digest, signature subject, workflow, and CI run have distinct fields because they answer different provenance questions.

A failure here can prevent a verified identity record even after an image was pushed. The Warehouse does not independently verify Cosign or consume the CI identity record; the cluster has no signature admission enforcement in this repository. The artifact trust concept describes that implemented limitation. Signature success also says nothing about dependency failure behavior.

2. Discover a candidate and pair it with chart source

Section titled “2. Discover a candidate and pair it with chart source”

The Warehouse discovers image tags matching ^sha-[a-f0-9]{7,40}$, selecting by newest build. Its Git subscription watches main changes under helm/url-shortener. Kargo’s Freight combines the selected image digest with the selected chart-source revision.

This is not an assertion that the two source commits are identical. An application release can remain the same while chart or observability work changes. The Warehouse excludes env/* rendered output branches from source discovery, preventing a promotion result from becoming a new chart input merely because it was committed.

Progress can stop on registry identity, access, or Git retrieval. CI publishing an image does not prove Kargo’s separate registry-reader identity can discover it. Warehouse tests check source selection and immutable image use.

The Project permits automatic dev promotion. Dev accepts Freight directly from the Warehouse. Its promotion template checks out the exact Freight chart-source commit into a source directory and env/dev into an output directory. It changes the source checkout’s image values to the selected repository and digest, clears prior output, renders Helm for the target namespace, and writes a new output commit.

Kargo pushes that commit and asks Argo CD to synchronize the exact result. The ApplicationSet continues tracking env/dev; it does not replace the configured branch with a floating source checkout. Argo CD reconciles desired state into Kubernetes, including its automated self-healing and pruning behavior.

Git credentials, chart rendering, push access, missing materialized Secrets, image pulls, scheduling, or application startup can block this handoff. A successful render proves manifests were produced. A reconciled healthy Application proves a narrower runtime state. Neither establishes chaos verification.

4. Verify dev before normal staging eligibility

Section titled “4. Verify dev before normal staging eligibility”

Dev’s service-health analysis queries /ready and expects JSON status ready. It is configured for three measurements at 15-second intervals, zero allowed failures, and two consecutive successes. This verifies the service’s readiness contract in the environment; it does not inject faults or repeat a paid smoke.

Successful upstream verification gives Freight its normal staging eligibility. Staging does not automatically promote merely because that eligibility exists. An operator explicitly requests the move after reviewing the candidate, target, and testnet budget. A manual Freight approval is a distinct override and must not be presented as evidence that dev verification passed.

The staging contract tests check the upstream source and manual Project policy. The existing promotion contract explains the operator boundary.

5. Render staging, then run both health and the gate

Section titled “5. Render staging, then run both health and the gate”

The staging Stage performs the same source-to-render pattern into env/staging, using staging values and the exact selected digest. Argo CD reconciles the resulting commit. Staging verification references both service-health and chaos-gate.

Kargo passes the Freight source/chart revision as release-revision and the application image digest as release-digest. The gate runner is a separately digest-pinned image in the AnalysisTemplate; its orchestration and scoring code do not come from a mutable runtime ConfigMap. That identity matters because verifier changes can change outcomes while application bytes remain constant.

The gate checks identities and budgets, acquires its Lease, checks ready target endpoints, and starts an exact run-scoped load Job. It requires a marker emitted after a successful paid create-and-redirect pair while the Job remains active. Only then does it create the serial PostgreSQL, Redis, and signer fault workflow. It preserves successful fault Apply timestamps, evaluates the query windows, and exits through cleanup. Inside the gate explains those transitions and early failures.

Prometheus scoring happens after the workflow observation phase. No implemented Prometheus preflight runs before fault injection. A telemetry defect can therefore block verification after an experiment has begun, and the cleanup path must still run.

Three passing scorecards are required but not sufficient. Workflow failure, terminated traffic, missing Apply evidence, scoring timeout, or cleanup failure can make the gate Job fail. Grafana annotations are optional diagnostics and cannot repair or invalidate a verdict. The Kargo AnalysisRun includes service health and the Job-backed chaos metric; a passing screenshot of one layer does not prove the complete enclosing result.

A successful staging verification makes the Freight normally eligible for prod. It does not perform an automatic prod promotion. The distinction protects the release narrative: “eligible” describes the evidence boundary; “promoted” describes an explicit consequential action.

7. Request prod-like promotion and verify its own target

Section titled “7. Request prod-like promotion and verify its own target”

The prod Stage accepts normal Freight from staging and remains manual. It renders env/prod, asks Argo CD to reconcile the output commit, and runs prod-post-deploy-health. That analysis checks both /ready and /livez after deployment.

This is a production-like Radius testnet environment. The smoke does not run paid load or repeat staging chaos. It establishes the prod-like target’s readiness and process liveness, with the staging verification available upstream. The environment shares the cluster with dev and staging; three application replicas do not turn single-instance PostgreSQL, Redis, or lab observability into highly available infrastructure.

A recorded journey, with each identity intact

Section titled “A recorded journey, with each identity intact”

The 3 October 2026 report records Freight d3b4380d40e87e244162da80aae9eb90503be15d for the fresh v1.0.0 alignment. Application release source was 3b70835335462f1b6f9b1dd17ab20d1108724f89; Freight chart-source was 0a98c08ccab5fa7b05efcecf877c714e5c89f025; application digest was sha256:ab88d89c20ccf1a79b9f47f90aa417ac0f54d7852c39b444b7fb788975fdd6a1.

Staging reconciled 64d2ddedb1493e0c59aef9ecad0ad0a2ff4196cf. Its analysis ran from 14:37:29Z to 14:46:28Z and passed all three fault scorecards. Prod later reconciled 7d334d5a05e747cabff59d28de75e73534afe578; readiness/liveness analysis ran from 15:34:59Z to 15:35:29Z. These are recorded historical observations, not the website’s current live state.

Detailed fresh bundles are privately retained outside Git. Public readers can inspect the report and reviewed screenshots, plus selected earlier scorecards under public evidence. The public summary must not imply the private archive is directly accessible. Continue with score to promotion to assess what each green layer establishes.

Maintained by Satyam Agnihotri · DevOps & Cloud Engineer