test-load — Measured concurrency, not a guess
Degree of freedom: MIXED — journey design [HIGH freedom]; env choice
and profile order [LOW freedom]. Never production without explicit
sign-off and rate caps.
Turn "I think it'll hold" into numbers: at N concurrent users, p95 is X, error
rate is Y, it falls over at Z.
audit-resilience tells you retries exist; this proves whether they save you
when 500 people arrive at once.
This skill vs neighbors
| Skill |
Owns |
| test-load (this) |
Concurrent traffic profiles + measured SLO |
audit-resilience |
Timeouts/retries/idempotency in code |
audit-infra-cost |
Consumes these capacity numbers to right-size |
audit-performance |
Single-user Web Vitals / bundle |
backend-db-performance |
Query/index fixes after this names the bottleneck |
plan-llm-cost-guardrails |
Model-token spend, not HTTP concurrency |
Never production without explicit sign-off and rate caps. Prefer prod-like
staging.
How to reason
- Observe — profile, VUs, p95/p99, error types, host/DB metrics
- Interpret — SLO miss vs capacity vs cascade (pool / limiter / memory)
- Classify — pass / degrade / breaking-point / invalid (unsigned prod)
- Severity — named failing resource + whether failure is graceful
Worked example
Observe: load profile 200 VU, p95 2.4s (SLO 500ms), 8% 529s; DB
remaining connections = 0.
Interpret: pool exhaustion, not the Node process.
Classify: SLO fail; breaking resource = DB pool.
Handoff: backend-db-performance (pool + queries); numbers → audit-infra-cost.
Phase 0 — Detect stack and set targets [HIGH freedom]
- Tool: k6 preferred (scriptable, CI-friendly); Artillery as fallback
- Target env + auth method
- Goal first (without it the test is noise):
- Expected peak concurrency (ads / launch)
- Acceptable p95/p99 and max error rate (SLO)
- Capacity (hold at peak) vs breaking point (where it fails)
Phase 1 — Model realistic load [HIGH freedom]
- Journeys, not one URL — signup → browse → action, with think-time
- Mix — weight by real traffic (mostly reads)
- Data variety — pool of users/inputs, not one cached row
- Auth — real login or pre-provisioned tokens (unauthenticated misses RLS cost)
Phase 2 — Graduated profiles [LOW freedom — smoke → load → stress → spike]
- Smoke — few VUs; script + baseline healthy
- Load — ramp to expected peak, soak several minutes (SLO pass/fail)
- Stress — past peak until degrade; name the failing resource
- Spike — instant jump (viral ad); does it recover?
- Optional soak — 30+ min for leaks / pool exhaustion
Phase 3 — Measure the right things [HIGH freedom]
- Latency percentiles (p50/p95/p99) — averages hide pain
- Error rate and types (timeout vs 5xx vs refused)
- Throughput (req/s) per level
- Bottleneck: DB pool, function concurrency, rate limiter, memory, cold start
- Correlate Supabase / Sentry / host metrics during the run
Commit the reusable script. Hand capacity numbers to audit-infra-cost.
Definition of Done
Self-critique before reporting [LOW freedom — do not skip]
- SLO stated first — a run without a target is noise
- Not unsigned prod — if only prod exists, stop
- Percentiles, not averages
- Breaking resource named — not "it got slow"
- Journeys, not one URL
Output format
- Design — SLO, journeys, env, profiles
- Results — profile | VUs | throughput | p50/p95/p99 | errors | vs SLO
- Breaking point — where, why, graceful vs cascade
- Handoff —
backend-db-performance, backend-patterns, audit-resilience, audit-infra-cost
Related
audit-resilience — code-level NFRs
audit-infra-cost — right-size from these numbers
backend-db-performance / backend-patterns — fix the named bottleneck
audit-performance — single-user frontend
1---2name: test-load3description: Design and run a k6/Artillery load profile that measures throughput, latency percentiles, error rate, and the breaking point under concurrent traffic. Use when "load test this", "will it handle launch traffic", or "find the breaking point". Resilience-by-reading-code → audit-resilience. Never hit prod unsigned.4license: MIT5---67# test-load — Measured concurrency, not a guess89**Degree of freedom: MIXED** — journey design `[HIGH freedom]`; env choice10and profile order `[LOW freedom]`. **Never production without explicit11sign-off and rate caps.**1213Turn "I think it'll hold" into numbers: at N concurrent users, p95 is X, error14rate is Y, it falls over at Z.1516**`audit-resilience` tells you retries exist; this proves whether they save you17when 500 people arrive at once.**1819## This skill vs neighbors2021| Skill | Owns |22|---|---|23| **test-load** (this) | Concurrent traffic profiles + measured SLO |24| `audit-resilience` | Timeouts/retries/idempotency *in code* |25| `audit-infra-cost` | Consumes these capacity numbers to right-size |26| `audit-performance` | Single-user Web Vitals / bundle |27| `backend-db-performance` | Query/index fixes after this names the bottleneck |28| `plan-llm-cost-guardrails` | Model-token spend, not HTTP concurrency |2930**Never production without explicit sign-off and rate caps.** Prefer prod-like31staging.3233## How to reason34351. **Observe** — profile, VUs, p95/p99, error types, host/DB metrics362. **Interpret** — SLO miss vs capacity vs cascade (pool / limiter / memory)373. **Classify** — pass / degrade / breaking-point / invalid (unsigned prod)384. **Severity** — named failing resource + whether failure is graceful3940## Worked example4142> **Observe:** load profile 200 VU, p95 2.4s (SLO 500ms), 8% 529s; DB43> `remaining connections = 0`.44> **Interpret:** pool exhaustion, not the Node process.45> **Classify:** SLO fail; breaking resource = DB pool.46> **Handoff:** `backend-db-performance` (pool + queries); numbers → `audit-infra-cost`.4748---4950## Phase 0 — Detect stack and set targets [HIGH freedom]5152- Tool: **k6** preferred (scriptable, CI-friendly); Artillery as fallback53- Target env + auth method54- **Goal first** (without it the test is noise):55 - Expected peak concurrency (ads / launch)56 - Acceptable p95/p99 and max error rate (SLO)57 - Capacity (hold at peak) vs breaking point (where it fails)5859---6061## Phase 1 — Model realistic load [HIGH freedom]6263- **Journeys, not one URL** — signup → browse → action, with think-time64- **Mix** — weight by real traffic (mostly reads)65- **Data variety** — pool of users/inputs, not one cached row66- **Auth** — real login or pre-provisioned tokens (unauthenticated misses RLS cost)6768---6970## Phase 2 — Graduated profiles [LOW freedom — smoke → load → stress → spike]71721. **Smoke** — few VUs; script + baseline healthy732. **Load** — ramp to expected peak, soak several minutes (SLO pass/fail)743. **Stress** — past peak until degrade; name the failing resource754. **Spike** — instant jump (viral ad); does it recover?765. Optional **soak** — 30+ min for leaks / pool exhaustion7778---7980## Phase 3 — Measure the right things [HIGH freedom]8182- Latency **percentiles** (p50/p95/p99) — averages hide pain83- Error rate **and types** (timeout vs 5xx vs refused)84- Throughput (req/s) per level85- Bottleneck: DB pool, function concurrency, rate limiter, memory, cold start86- Correlate Supabase / Sentry / host metrics during the run8788Commit the reusable script. Hand capacity numbers to `audit-infra-cost`.8990---9192## Definition of Done9394- [ ] Tool + env chosen; prod excluded or signed + capped95- [ ] SLO stated before running96- [ ] Journeys weighted, think-time, varied data, real auth97- [ ] Smoke → load → stress → spike run98- [ ] Percentiles, errors, throughput captured99- [ ] Breaking point + failing resource named100- [ ] SLO verdict + fix handoff101- [ ] Script committed102103## Self-critique before reporting [LOW freedom — do not skip]1041051. **SLO stated first** — a run without a target is noise1062. **Not unsigned prod** — if only prod exists, stop1073. **Percentiles, not averages**1084. **Breaking resource named** — not "it got slow"1095. **Journeys, not one URL**110111## Output format1121131. **Design** — SLO, journeys, env, profiles1142. **Results** — profile | VUs | throughput | p50/p95/p99 | errors | vs SLO1153. **Breaking point** — where, why, graceful vs cascade1164. **Handoff** — `backend-db-performance`, `backend-patterns`, `audit-resilience`, `audit-infra-cost`117118## Related119120- `audit-resilience` — code-level NFRs121- `audit-infra-cost` — right-size from these numbers122- `backend-db-performance` / `backend-patterns` — fix the named bottleneck123- `audit-performance` — single-user frontend