A connected system tour
The system has two connected journeys. A request goes through the application to create or resolve a URL. A release goes through publication, rendering, deployment, and verification to establish whether that application’s candidate may move forward. Keeping the journeys separate makes it easier to locate a failure: an HTTP problem can arise in the application, its dependencies, or the load client; a release problem can arise before any fault reaches the workload.
Enlarge diagram: Application, data, signing, and verification responsibilities connected at their runtime boundaries · Version-controlled diagram source
Follow the request through its owners
Section titled “Follow the request through its owners”POST /shorten validates an absolute HTTP or HTTPS destination, strips surrounding whitespace, and rejects embedded user credentials. It first asks PostgreSQL whether the destination already exists. An existing mapping returns 200 and its original code without creating a second mapping or settling another payment. A new mapping receives a random code and returns 201 after a successful insertion.
PostgreSQL is the source of truth. Its table constrains the code and destination to be unique, and, for paid creations, constrains a non-null settlement identifier to be unique. The application makes up to three insertion attempts with conflict-safe SQL. This bounds allocation work while accounting for code collisions and concurrent duplicate destinations.
GET /{code} tries Redis first. A cached mapping produces 302 immediately. An absent cache value or cache exception falls through to PostgreSQL. A found database mapping is written into Redis on a best-effort basis before the redirect; a cache write failure does not cancel the redirect. An unknown code returns 404, while an unavailable database returns 503 when the route needs storage.
Payments add another path for new destinations. The application issues an x402 challenge, the client asks the signer for a Permit2 authorization, and the application sends that signed envelope with server-owned requirements to facilitator /verify and /settle. The signer is a client dependency; the application does not call it when handling /shorten. The request walkthrough explains that separation and the post-settlement persistence gap.
Follow a candidate through its owners
Section titled “Follow a candidate through its owners”CI builds and publishes images and performs Cosign signing and verification. Kargo discovers Freight, selects a candidate, and renders environment manifests using its digest. The rendered Git revision and image digest identify different things: one identifies deployment output, the other image bytes. Argo CD reconciles the output into the environment.
Only development is configured for automatic promotion. Staging and the production-like testnet stage require explicit promotion. The normal production-like selection path accepts Freight verified in staging; manually approving Freight is a separate override. It is not additional proof that staging verification succeeded. Current stage resources do not independently enforce CI signatures, so digest identity and artifact trust must also be kept distinct.
Staging’s gate acquires run ownership, verifies prerequisites and useful paid traffic, starts sequential PostgreSQL, Redis, and signer pod-failure experiments, and invokes the scorer over bounded windows. Target selectors narrow faults to staging dependencies. Prometheus measurements, scorecard metadata, enclosing workflow state, and cleanup jointly determine the result. A favorable per-experiment score alone is insufficient if the enclosing job or cleanup fails.
Notice the health boundary
Section titled “Notice the health boundary”The application’s /livez checks only that the process can respond. /ready always checks PostgreSQL and requires Redis until the process has achieved its first successful readiness. After that first success, Redis failure can exercise fallback without making an otherwise usable application unready. A replacement application process must pass the initial Redis check again.
Kubernetes startup and readiness probes call /ready; liveness calls /livez. Local Docker health checks call /, which reports that the service is configured and does not test the database. These are different contracts. A Docker container can be reported healthy while dependency readiness remains unsuccessful. The health and recovery explanation shows how those boundaries affect restart and recovery interpretation.
Inspect the evidence after the mechanism
Section titled “Inspect the evidence after the mechanism”The repository retains public sanitized historical gate logs and scorecards, plus a reviewed screenshot gallery and later release verification report. Read the candidate, run, and time window before comparing an observed value with a threshold. Do not carry a historical pass forward to a new candidate, or treat a historical payment dashboard as the final release’s payment receipt.
The remaining lab limits are architectural. Namespaces share a cluster; lab-sized persistence and observability are not a backup or availability design. Faults are sequential and bounded rather than a survey of every possible combination. These choices make observations attributable and understandable while leaving larger operational questions for future work.
Continue with dependencies and fault boundaries or the application design. Verify the tour against application code, load generation, deployment probes, and the platform architecture.
Maintained by Satyam Agnihotri · DevOps & Cloud Engineer