Skip to content

Bounded load testing

A resilience experiment needs work to protect. Quiet error counters may simply mean that no request reached the app. The maintained load runner distinguishes an unpaid dev baseline from paid staging traffic, caps both, and keeps the source CronJob suspended so a reviewed one-off run cannot become an unattended spending loop.

This guide is an operator procedure for an explicitly owned testnet lab. Plan examples make no network call. Execution creates a Job, and staging iterations can settle testnet tokens. A documentation review or local test run does not authorize executing them.

Scenario Target and purpose Enforced bounds
baseline Unpaid GET / against dev; normal service traffic and telemetry Arrival rate only, one VU, at most 90 seconds, at most 60 iterations/minute, at most 180-second wait; Job deadline 150 seconds.
staging (default) Paid URL creation and redirect checks against staging At most three VUs and ten minutes; testnet network eip155:72344.

The baseline reuses the suspended staging loadgen CronJob’s digest-pinned k6 template but overrides traffic mode and target. Its Job runs in the source namespace while requests target dev. It never calls the signer or a payment endpoint. This implemented baseline path is broader than the older staging-only runbook; use the checked-in script as the command contract.

Terminal window
./scripts/run-loadgen.sh --plan --scenario baseline \
--duration 90s --vus 1 --arrival-rate 30 --arrival-time-unit 1m

For paid staging traffic:

Terminal window
./scripts/run-loadgen.sh --plan --scenario staging \
--profile arrival-rate --duration 10m --vus 3 \
--arrival-rate 60 --arrival-time-unit 1m

An arrival-rate profile schedules attempts at the requested rate; a closed-loop profile cycles a fixed number of users with a pause. Neither configuration proves that the requested rate was achieved. Inspect actual iteration counts, failed checks, latency, and payment results. A signer outage can prevent the load client from creating paid requests, so an app’s quiet 5xx series alone is insufficient evidence of continuous useful traffic.

Confirm the exact intended Kubernetes context, app health, and the suspended source CronJob. For staging, also establish signer readiness, Ready signer/loadgen ExternalSecrets, and funded Permit2-bootstrapped wallets within the reviewed testnet budget. A VU corresponds to one bounded signer wallet.

Private keys belong only to the signer Secret boundary. The loadgen Secret holds SERVICE_WALLET_ADDRESS, not payer keys. Never copy keys into a Job, shell environment, logs, artifacts, or Git. The runner checks the configured testnet network; do not broaden it or increase bounds to work around a failing run.

After a separately reviewed operation, execute the same bounded configuration by changing the mode:

Terminal window
./scripts/run-loadgen.sh --execute --scenario baseline \
--duration 90s --vus 1 --arrival-rate 30 --arrival-time-unit 1m

For staging, the equivalent --execute command is an explicit acknowledgement of paid testnet traffic. The runner checks readiness and source suspension before creating a uniquely named Job. It does not unsuspend or modify the CronJob. The k6 program checks app /livez and /ready, then signer /livez, /health, and /wallets for paid mode. Anonymous k6 usage reporting is disabled.

Paid mode creates unique destinations under example.invalid and checks redirects with redirects: 0; it does not follow the external destination. A reviewed REDIRECT_PROBE_URL can separately measure an existing short URL. Those two probes answer different traffic questions.

The in-container summary is /results/summary.json. The default local destination is /tmp/resilience-gate-loadgen/<job>-summary.json; --artifact-dir selects an absolute output path. The staging Job has a 24-hour completion TTL. Baseline copies its summary and records actual Pod start/finish timestamps, then deletes and verifies absence of its exact Job. Cleanup failure makes the operation fail.

The baseline scorer is separate:

Terminal window
python3 scripts/score-baseline.py \
--run-id <actual-run-id> --namespace url-shortener-dev \
--source-revision <actual-source-sha> \
--release-digest <actual-sha256-digest> \
--started-at <actual-pod-start-utc> --ended-at <actual-k6-finish-utc> \
--prom <credential-free-prometheus-url> --output <local-scorecard-path>

It requires a 60–150-second observation window, more than 20 GET / requests, fewer than 0.5 5xx responses, final-one-minute p95 below one second, ready PostgreSQL/Redis across the window, and immutable release identity. Prometheus increase can produce fractional extrapolated counts; the historical baseline’s 44.89 requests is not a literal fractional HTTP request.

A manual load result is operational evidence. It is not Kargo verification, applied chaos, or downstream eligibility. The release gate owns its own run-scoped load and paid-traffic preflight. If a Job fails, preserve available logs/summary, stop creating new Jobs, and use layered troubleshooting before retrying.

Inspect run-loadgen.sh, loadgen.js, baseline scorer, runner tests, and baseline scoring tests for the implemented boundary.

Maintained by Satyam Agnihotri · DevOps & Cloud Engineer