Backend Testing
Use this skill as a packet-first backend testing router.
The job is not to dump boilerplate for every framework. The job is to:
- classify the request into the right backend test packet,
- pick the smallest credible mix of test layers,
- make dependency realism and data control explicit,
- split local vs PR vs slower lanes honestly,
- route policy, contract-shape, and auth-implementation work away when they are the real task.
Read these when needed:
- references/intake-packets-and-route-outs.md
- references/test-layer-matrix.md
- references/stability-checklist.md
When to use this skill
- Add or repair backend coverage for APIs, services, repositories, workers, integrations, or auth flows
- Decide whether a backend change needs unit, integration, contract/API, or narrow smoke coverage
- Design fixture, factory, seed/reset, auth bootstrap, or environment-control strategy
- Decide when to use mocks, fakes, containers, or real dependencies
- Stabilize flaky backend suites, especially CI-only failures and local-vs-CI drift
- Review whether a backend suite is too broad, too slow, too mock-heavy, or missing a key layer
When not to use this skill
- The main task is org-wide test policy, gate design, release evidence, or company-wide QA philosophy → use
testing-strategies
- The main task is API contract shape, versioning, or schema design before tests can be scoped honestly → use
api-design
- The main task is implementing auth/session/provider behavior rather than testing it → use
authentication-setup
- The main task is frontend/browser testing or UI workflow coverage
- There is no concrete backend behavior or regression target yet; in that case define the missing behavior packet first instead of pretending the test plan is settled
Instructions
Step 1: Classify the request into one packet
Choose the single best entry packet before giving advice.
Packets
coverage-plan — which layers to add for a concrete backend change
fixture-and-reset-plan — how to seed, isolate, reset, or bootstrap data/auth state
contract-and-api-checks — how to protect response/event/schema compatibility once the interface already exists
flake-stabilization — how to stabilize CI-only or intermittent backend failures
execution-lane-split — how to divide local-fast, PR, nightly, and release-only backend checks
If the request mixes several concerns, name the primary packet and one secondary concern.
Step 2: Frame the backend surface and risk
Capture the smallest useful context:
- surface: endpoint, service, repository, worker, queue consumer, auth flow, integration, or migration
- highest-risk behaviors: validation, permissions, persistence, retries, idempotency, ordering, serialization, side effects, compatibility
- existing coverage already present
- external dependencies involved: DB, cache, queue, email, payment, third-party API, identity provider, filesystem
- runtime/language stack
- where the evidence must hold: local loop, PR CI, scheduled CI, release smoke
If the request is vague, choose the smallest regression slice worth protecting first.
Step 3: Choose the right test layers
Use the packet and risk to select the lightest credible layer mix.
Unit / service
Prefer when the main risk is branching logic, validation, orchestration, or pure-ish business rules.
Integration
Prefer when database behavior, framework wiring, middleware, transactions, queues, caches, or serialization matter.
Contract / API
Prefer when clients depend on response shapes, status codes, schemas, or events and the interface already exists.
Smoke / selective end-to-end
Prefer only when a narrow release-critical journey crosses several backend boundaries and lower layers would miss the core risk.
State what is in scope, what is out of scope, and why.
Step 4: Decide dependency realism on purpose
For each dependency, choose one of:
- mock / stub — expensive, unstable, or irrelevant to the behavior under test
- fake / simulator — behavior matters, but a lightweight substitute is enough
- containerized real dependency — queries, migrations, message semantics, or wire behavior matter enough that drift would hurt
- shared external environment — only when unavoidable; call out the fragility cost explicitly
Good defaults:
- prefer real DB behavior when repository, migration, transaction, or serialization behavior is central
- prefer mocks for outbound third-party APIs unless the integration contract itself is under test
- prefer a narrow containerized slice over a giant all-dependencies-in-PR setup
- do not claim fake and real dependencies are equivalent when production parity is the whole risk
Step 5: Define fixture, data, auth, and environment control
A backend suite becomes untrustworthy when state is vague.
Specify:
- fixture/factory strategy
- seed/reset/rollback plan
- auth/bootstrap helpers for users, roles, tenants, tokens, or sessions
- time/randomness/idempotency control where needed
- isolation rule: per test, per file, per suite, or per environment
- debugging signals to capture when failures happen
If the suite relies on ordering, leftovers, or sleeps, call that fragility out directly.
Step 6: Split the execution lanes
Treat local, PR, and slower lanes as different jobs.
Define:
- local-fast path — what developers should run repeatedly
- PR path — what must gate merges
- scheduled / nightly path — heavier breadth or expensive realism
- release / incident path — narrow confidence checks or regression ratchets when needed
If the suite is slow, split it. Do not pretend one giant authoritative path is practical everywhere.
Step 7: Produce one backend test packet
Return one concise packet, not a general essay.
Recommended packet shapes:
coverage-plan → coverage table + dependency strategy + exclusions
fixture-and-reset-plan → fixture/reset memo + auth/bootstrap notes
contract-and-api-checks → compatibility packet + consumer/provider scope + route-outs
flake-stabilization → flake memo with likely causes, isolation fixes, readiness checks, and debug signals
execution-lane-split → lane matrix with local/PR/scheduled/release responsibilities
Minimum packet contents:
- change surface and primary risk
- chosen packet and any secondary concern
- selected layers and why
- dependency realism decisions
- fixture/data/auth/environment control
- execution-lane split
- explicit route-outs when the request is partly owned elsewhere
Step 8: Verify scope boundaries before finalizing
Check:
- does the packet protect the real backend regression risk rather than generic coverage vanity?
- did you keep org-wide validation policy in
testing-strategies?
- did you route contract shape decisions to
api-design while keeping contract protection here only when the interface already exists?
- did you route auth implementation work to
authentication-setup?
- will a maintainer understand why a dependency is mocked, faked, containerized, or real?
Output format
## Backend Test Packet: [Surface or Change]
### Packet choice
- Primary packet: coverage-plan | fixture-and-reset-plan | contract-and-api-checks | flake-stabilization | execution-lane-split
- Secondary concern: optional
- Confidence: high | medium | low
### Change framing
- Surface: ...
- Main risks: ...
- Runtime: ...
- Existing coverage: ...
### Layer decisions
| Layer | In scope? | What it protects | Notes |
|------|-----------|------------------|-------|
| Unit / service | yes/no | ... | ... |
| Integration | yes/no | ... | ... |
| Contract / API | yes/no | ... | ... |
| Smoke / selective E2E | yes/no | ... | ... |
### Dependency realism
| Dependency | Strategy | Why |
|------------|----------|-----|
| Database / queue / cache | ... | ... |
| External API | ... | ... |
| Auth provider | ... | ... |
### Data and environment control
- Fixtures / factories: ...
- Seed / reset: ...
- Auth bootstrap: ...
- Isolation rule: ...
- Debug signals: ...
### Execution lanes
- Local-fast: ...
- PR CI: ...
- Scheduled / nightly: ...
- Release / incident: ...
### Route-outs
- `testing-strategies`: ...
- `api-design`: ...
- `authentication-setup`: ...
Examples
Example 1: auth-heavy API change
Input: “We added refresh-token rotation and new admin-only endpoints to our Express API. I need backend tests that catch auth failures, token replay issues, and DB persistence bugs without turning CI into a giant end-to-end suite.”
Good response shape:
- chooses
coverage-plan as the primary packet
- combines unit/service plus integration/API coverage instead of one giant E2E suite
- keeps real DB or containerized persistence where token/session behavior matters
- defines auth bootstrap helpers and reset strategy
- limits smoke coverage to a narrow release-critical path
Example 2: CI-only flake in a service suite
Input: “Our FastAPI tests pass locally but fail in CI around seeded Postgres state and background jobs. Give me a stabilization plan.”
Good response shape:
- chooses
flake-stabilization as the primary packet
- identifies seed/reset drift, readiness, async timing, or leftover state as likely causes
- recommends stronger isolation, readiness checks, and debugging signals instead of just retries
- separates local-fast and CI-authoritative behavior clearly
Example 3: contract protection after an API already exists
Input: “Our payment service and webhook consumers keep drifting on response fields. I do not need API redesign, I need backend tests that catch compatibility regressions.”
Good response shape:
- chooses
contract-and-api-checks as the primary packet
- keeps contract protection here because the interface already exists
- routes any schema redesign or versioning debate to
api-design
- recommends consumer/provider or schema-compatibility coverage rather than broader smoke inflation
Example 4: too-broad policy request
Input: “Design our overall engineering org testing strategy for frontend, backend, mobile, and QA.”
Good response shape:
- recognizes that the primary task belongs to
testing-strategies
- keeps any backend-specific advice scoped as a handoff only
- refuses to turn
backend-testing into a universal QA-governance skill
Best practices
- Start from the packet, not from the framework.
- Protect the real backend regression risk before chasing coverage percentages.
- Prefer layered backend coverage over giant brittle end-to-end suites.
- Make fixture, seed, and auth bootstrap strategy explicit; hidden state is where trust dies.
- Split local-fast, PR, scheduled, and release lanes intentionally.
- Use real dependencies when wire behavior matters, but keep expensive realism bounded.
- Treat flaky tests as a trust problem, not just an annoyance.
- Route policy, contract-shape, and auth-implementation ownership away instead of absorbing them.
References
1---2name: backend-testing3description: Turn backend test ambiguity into one practical backend test packet. Use when the user needs API/service/repository/auth-flow coverage design, fixture or seed/reset strategy, container-vs-mock dependency choices, contract/API compatibility checks, or flaky backend-suite stabilization across local and CI.4---567891011121314# Backend Testing1516Use this skill as a **packet-first backend testing router**.1718The job is not to dump boilerplate for every framework. The job is to:191. classify the request into the right backend test packet,202. pick the smallest credible mix of test layers,213. make dependency realism and data control explicit,224. split local vs PR vs slower lanes honestly,235. route policy, contract-shape, and auth-implementation work away when they are the real task.2425Read these when needed:26- references/intake-packets-and-route-outs.md27- references/test-layer-matrix.md28- references/stability-checklist.md2930## When to use this skill31- Add or repair backend coverage for APIs, services, repositories, workers, integrations, or auth flows32- Decide whether a backend change needs unit, integration, contract/API, or narrow smoke coverage33- Design fixture, factory, seed/reset, auth bootstrap, or environment-control strategy34- Decide when to use mocks, fakes, containers, or real dependencies35- Stabilize flaky backend suites, especially CI-only failures and local-vs-CI drift36- Review whether a backend suite is too broad, too slow, too mock-heavy, or missing a key layer3738## When not to use this skill39- The main task is org-wide test policy, gate design, release evidence, or company-wide QA philosophy → use `testing-strategies`40- The main task is API contract shape, versioning, or schema design before tests can be scoped honestly → use `api-design`41- The main task is implementing auth/session/provider behavior rather than testing it → use `authentication-setup`42- The main task is frontend/browser testing or UI workflow coverage43- There is no concrete backend behavior or regression target yet; in that case define the missing behavior packet first instead of pretending the test plan is settled4445## Instructions4647### Step 1: Classify the request into one packet48Choose the single best entry packet before giving advice.4950**Packets**51- `coverage-plan` — which layers to add for a concrete backend change52- `fixture-and-reset-plan` — how to seed, isolate, reset, or bootstrap data/auth state53- `contract-and-api-checks` — how to protect response/event/schema compatibility once the interface already exists54- `flake-stabilization` — how to stabilize CI-only or intermittent backend failures55- `execution-lane-split` — how to divide local-fast, PR, nightly, and release-only backend checks5657If the request mixes several concerns, name the **primary packet** and one secondary concern.5859### Step 2: Frame the backend surface and risk60Capture the smallest useful context:61- surface: endpoint, service, repository, worker, queue consumer, auth flow, integration, or migration62- highest-risk behaviors: validation, permissions, persistence, retries, idempotency, ordering, serialization, side effects, compatibility63- existing coverage already present64- external dependencies involved: DB, cache, queue, email, payment, third-party API, identity provider, filesystem65- runtime/language stack66- where the evidence must hold: local loop, PR CI, scheduled CI, release smoke6768If the request is vague, choose the smallest regression slice worth protecting first.6970### Step 3: Choose the right test layers71Use the packet and risk to select the lightest credible layer mix.7273#### Unit / service74Prefer when the main risk is branching logic, validation, orchestration, or pure-ish business rules.7576#### Integration77Prefer when database behavior, framework wiring, middleware, transactions, queues, caches, or serialization matter.7879#### Contract / API80Prefer when clients depend on response shapes, status codes, schemas, or events and the interface already exists.8182#### Smoke / selective end-to-end83Prefer only when a narrow release-critical journey crosses several backend boundaries and lower layers would miss the core risk.8485State what is **in scope**, what is **out of scope**, and why.8687### Step 4: Decide dependency realism on purpose88For each dependency, choose one of:89- **mock / stub** — expensive, unstable, or irrelevant to the behavior under test90- **fake / simulator** — behavior matters, but a lightweight substitute is enough91- **containerized real dependency** — queries, migrations, message semantics, or wire behavior matter enough that drift would hurt92- **shared external environment** — only when unavoidable; call out the fragility cost explicitly9394Good defaults:95- prefer real DB behavior when repository, migration, transaction, or serialization behavior is central96- prefer mocks for outbound third-party APIs unless the integration contract itself is under test97- prefer a narrow containerized slice over a giant all-dependencies-in-PR setup98- do not claim fake and real dependencies are equivalent when production parity is the whole risk99100### Step 5: Define fixture, data, auth, and environment control101A backend suite becomes untrustworthy when state is vague.102103Specify:104- fixture/factory strategy105- seed/reset/rollback plan106- auth/bootstrap helpers for users, roles, tenants, tokens, or sessions107- time/randomness/idempotency control where needed108- isolation rule: per test, per file, per suite, or per environment109- debugging signals to capture when failures happen110111If the suite relies on ordering, leftovers, or sleeps, call that fragility out directly.112113### Step 6: Split the execution lanes114Treat local, PR, and slower lanes as different jobs.115116Define:117- **local-fast path** — what developers should run repeatedly118- **PR path** — what must gate merges119- **scheduled / nightly path** — heavier breadth or expensive realism120- **release / incident path** — narrow confidence checks or regression ratchets when needed121122If the suite is slow, split it. Do not pretend one giant authoritative path is practical everywhere.123124### Step 7: Produce one backend test packet125Return one concise packet, not a general essay.126127Recommended packet shapes:128- `coverage-plan` → coverage table + dependency strategy + exclusions129- `fixture-and-reset-plan` → fixture/reset memo + auth/bootstrap notes130- `contract-and-api-checks` → compatibility packet + consumer/provider scope + route-outs131- `flake-stabilization` → flake memo with likely causes, isolation fixes, readiness checks, and debug signals132- `execution-lane-split` → lane matrix with local/PR/scheduled/release responsibilities133134Minimum packet contents:135- change surface and primary risk136- chosen packet and any secondary concern137- selected layers and why138- dependency realism decisions139- fixture/data/auth/environment control140- execution-lane split141- explicit route-outs when the request is partly owned elsewhere142143### Step 8: Verify scope boundaries before finalizing144Check:145- does the packet protect the real backend regression risk rather than generic coverage vanity?146- did you keep org-wide validation policy in `testing-strategies`?147- did you route contract *shape* decisions to `api-design` while keeping contract *protection* here only when the interface already exists?148- did you route auth implementation work to `authentication-setup`?149- will a maintainer understand why a dependency is mocked, faked, containerized, or real?150151## Output format152153```markdown154## Backend Test Packet: [Surface or Change]155156### Packet choice157- Primary packet: coverage-plan | fixture-and-reset-plan | contract-and-api-checks | flake-stabilization | execution-lane-split158- Secondary concern: optional159- Confidence: high | medium | low160161### Change framing162- Surface: ...163- Main risks: ...164- Runtime: ...165- Existing coverage: ...166167### Layer decisions168| Layer | In scope? | What it protects | Notes |169|------|-----------|------------------|-------|170| Unit / service | yes/no | ... | ... |171| Integration | yes/no | ... | ... |172| Contract / API | yes/no | ... | ... |173| Smoke / selective E2E | yes/no | ... | ... |174175### Dependency realism176| Dependency | Strategy | Why |177|------------|----------|-----|178| Database / queue / cache | ... | ... |179| External API | ... | ... |180| Auth provider | ... | ... |181182### Data and environment control183- Fixtures / factories: ...184- Seed / reset: ...185- Auth bootstrap: ...186- Isolation rule: ...187- Debug signals: ...188189### Execution lanes190- Local-fast: ...191- PR CI: ...192- Scheduled / nightly: ...193- Release / incident: ...194195### Route-outs196- `testing-strategies`: ...197- `api-design`: ...198- `authentication-setup`: ...199```200201## Examples202203### Example 1: auth-heavy API change204**Input:** “We added refresh-token rotation and new admin-only endpoints to our Express API. I need backend tests that catch auth failures, token replay issues, and DB persistence bugs without turning CI into a giant end-to-end suite.”205206**Good response shape:**207- chooses `coverage-plan` as the primary packet208- combines unit/service plus integration/API coverage instead of one giant E2E suite209- keeps real DB or containerized persistence where token/session behavior matters210- defines auth bootstrap helpers and reset strategy211- limits smoke coverage to a narrow release-critical path212213### Example 2: CI-only flake in a service suite214**Input:** “Our FastAPI tests pass locally but fail in CI around seeded Postgres state and background jobs. Give me a stabilization plan.”215216**Good response shape:**217- chooses `flake-stabilization` as the primary packet218- identifies seed/reset drift, readiness, async timing, or leftover state as likely causes219- recommends stronger isolation, readiness checks, and debugging signals instead of just retries220- separates local-fast and CI-authoritative behavior clearly221222### Example 3: contract protection after an API already exists223**Input:** “Our payment service and webhook consumers keep drifting on response fields. I do not need API redesign, I need backend tests that catch compatibility regressions.”224225**Good response shape:**226- chooses `contract-and-api-checks` as the primary packet227- keeps contract protection here because the interface already exists228- routes any schema redesign or versioning debate to `api-design`229- recommends consumer/provider or schema-compatibility coverage rather than broader smoke inflation230231### Example 4: too-broad policy request232**Input:** “Design our overall engineering org testing strategy for frontend, backend, mobile, and QA.”233234**Good response shape:**235- recognizes that the primary task belongs to `testing-strategies`236- keeps any backend-specific advice scoped as a handoff only237- refuses to turn `backend-testing` into a universal QA-governance skill238239## Best practices2401. Start from the packet, not from the framework.2412. Protect the real backend regression risk before chasing coverage percentages.2423. Prefer layered backend coverage over giant brittle end-to-end suites.2434. Make fixture, seed, and auth bootstrap strategy explicit; hidden state is where trust dies.2445. Split local-fast, PR, scheduled, and release lanes intentionally.2456. Use real dependencies when wire behavior matters, but keep expensive realism bounded.2467. Treat flaky tests as a trust problem, not just an annoyance.2478. Route policy, contract-shape, and auth-implementation ownership away instead of absorbing them.248249## References250- [The Practical Test Pyramid](https://martinfowler.com/articles/practical-test-pyramid.html)251- [pytest: Good Integration Practices](https://docs.pytest.org/en/stable/explanation/goodpractices.html)252- [Testcontainers](https://testcontainers.com/)253- [Pact Docs](https://docs.pact.io/)254- [GitHub Actions: Building and testing Python](https://docs.github.com/en/actions/use-cases-and-examples/building-and-testing/building-and-testing-python)