Skip to content

Delivery and platform design

This source-derived design describes how Resilience Gate turns selected artifacts into environment state and verification eligibility. The design covers the existing platform, not the documentation website. Its deployed scope is one owned GCP/Radius testnet lab; prod is a production-like profile within that scope.

The delivery subsystem must preserve the selected application bytes, make desired-state changes attributable, keep promotion separate from reconciliation, and require environment-appropriate verification. These requirements are inferred from the implementation and existing promotion and artifact contracts. They are not claims of a production-certified supply chain.

The principal quality attributes are traceability, repeatable rendering, bounded authority, and explicit consequential actions. Registry retention, Git availability, cloud IAM correctness, and live controller operation remain dependencies. The rationale for keeping source and output separate is source-derived analysis: it makes the transformation independently inspectable and avoids discovering rendered output as a new source candidate.

Component Input Output / responsibility
GitHub application CI Trusted source revision and portable registry configuration Validated build, pushed digest, verified Cosign identity record
Kargo Warehouse Constrained application tags and chart-source changes on main Freight carrying image digest and Git source revision
Kargo Project/Stages Freight, promotion policy, environment values Selected candidate, rendered output commit, Argo CD sync request, verification context
Kargo Git credential ExternalSecret-backed token Authenticated output-branch writes
Argo CD root Application Bootstrap path on main Reconciliation of platform resources
Argo CD ApplicationSet Three environment descriptors and env/* branches Digest-pinned workload reconciliation
AnalysisTemplates Service URL and, for staging, release identity Readiness/liveness measurements and Job-backed gate outcome
Kubernetes Rendered objects and external materialized Secrets Pods, Services, persistent dependencies, execution state

Delivery ownership diagram

Enlarge diagram: Delivery ownership diagram · Version-controlled diagram source

A rendered commit is the interface between selection and reconciliation. Verification consumes the resulting target; it does not replace either controller’s responsibility.

The ApplicationSet uses an exact authorized-stage annotation for each target Application and tracks env/<environment>. Its automated prune/self-heal loop belongs to reconciliation. The root Application follows main for bootstrap content, a separate input path from workload output branches.

The Warehouse subscription uses newest application builds with sha-only tags and chart changes under helm/url-shortener. Each Stage template clones source at the Freight Git commit and output at its environment branch. It modifies image repository/digest in the source checkout, renders Helm with base and environment values for the configured Kubernetes version, commits and pushes the output, then requests synchronization to that exact commit.

Application build source, chart-source revision, digest, render commit, verifier digest, and run identity are separate records. They must not collapse into one “release SHA.” The render is not a mutable tag deployment: the chart’s runtime reference uses the selected digest. The gate image is separately pinned and must be reviewed into the AnalysisTemplate after a verifier publication.

Cosign verification runs in image publication CI. Neither Warehouse selection nor Kubernetes admission in this repository independently checks the signature or requires CI’s identity record. A pushed but unsuccessfully signed image can remain discoverable. This is an implemented trust limitation, not something this design silently remedies.

Stage Normal Freight source Promotion policy Verification
Dev Warehouse directly Automatic allowed /ready service health
Staging Dev upstream Manual /ready plus bounded chaos Job
Prod-like Staging upstream Manual Post-deploy /ready and /livez

The Project owns the policy. Upstream verification supplies normal eligibility, not an automatic downstream action. A manual Freight approval is a separate override. Successful staging therefore means the selected candidate met the verification boundary; prod still needs an explicit promotion, a new output commit, target reconciliation, and its smoke result.

The prod-like smoke does not repeat paid load or chaos. Its AnalysisTemplate requires readiness and liveness at its own target. The Stage is testnet-only and shares one cluster with dev/staging. Different replica counts and namespaces improve the experiment’s organization without establishing regional isolation or dependency high availability.

The GitOps bootstrap phase registers the Project, credential reference, policies, analyses, Stages, AppProject, and ApplicationSet before enabling Warehouse discovery. It checks rendered public configuration and the configured target context. It does not execute promotions or experiments.

The identity flow separates CI publisher, Kargo registry reader, External Secrets cloud reader, Git credential, and gate Kubernetes service account. Secret values are external; environment render success does not establish remote value availability. Missing or unusable materialization can block workload startup and readiness after GitOps succeeds structurally.

Each handoff is independently observable. Discovery can fail before candidate selection. Rendering or pushing can fail before desired state exists. Argo CD can fail to reconcile a pushed commit. Health or chaos can fail after deployment. Cleanup can fail after metric checks pass. Operators should diagnose the first unsupported transition instead of patching live workloads and describing them as an intact release path.

There is no single transaction spanning image registry, Git branches, Kubernetes reconciliation, and verification. A failure can leave an uploaded image, rendered commit, or deployed unverified candidate behind. Recorded source/render/runtime identities make that partial progress explainable. They do not automatically roll back every subsystem.

The GitOps tests, staging tests, prod tests, and CI tests check configuration and offline contracts. The verification report separately records live owned-testnet observations. Those evidence classes should remain distinct.

The design deliberately avoids adding a database or live dashboard to the documentation site, and website CI does not drive this platform. The platform’s lab lifecycle and chaos runbook govern real operations with explicit context and budget checks.

A valid release claim links Freight, application source/digest, chart-source, rendered revision, relevant runtime digests, analysis window, scorecards, and cleanup. Historical campaigns cannot prove later candidates. The fresh October 3 report is a recorded snapshot with detailed bundles privately retained, not a permanently current deployment assertion. Continue with gate design and integrated system design for the behavioral boundary layered onto delivery.

Maintained by Satyam Agnihotri · DevOps & Cloud Engineer