Test-Driven Development
TDD uses tests as a design tool. The test suite is a byproduct — the real output is clean design that emerges from writing tests first. Code is only written to satisfy a failing test.
Before Starting
Gather context before writing anything. A wrong assumption means wrong tests.
- Find existing tests. Match their framework and style exactly. Only ask about tooling if nothing exists yet.
- Read code this task integrates with. Discover interfaces rather than inventing them.
- Clarify behavior if vague. If the request is ambiguous ("add authentication"), ask for a concrete scenario: "What should happen when a user logs in with the wrong password?" One focused question beats a checklist.
The Cycle: Red -> Green -> Refactor
Every unit of work follows three steps. Never skip the third.
Red — Write a failing test
- Pick the next smallest piece of behavior
- Write a test that will pass once that behavior exists
- Name it for the behavior, not the function under test —
rejects_login_with_wrong_password, nottest_login_2 - Assert what you actually expect, not a placeholder. If you don't know the
exact value (e.g., a specific error message string), assert the contract
you care about (e.g., "raises
ValueError") rather than guess. Tweaking the assertion later to match whatever the implementation happens to emit is test-after, not test-first — see Pitfalls. - Run it — confirm it fails for the right reason (not a syntax error or import failure)
- If writing the test is hard, that's a design signal: the interface isn't clear yet
Green — Minimal code to pass
- Write only enough production code to make the failing test pass
- Inelegant or hardcoded is fine here — correctness first
- Run all tests — the new one passes and nothing else broke
Refactor — Clean up, behavior unchanged
- Improve names, remove duplication, simplify logic
- Refactor both production code and test code — for tests that means naming and clarity, not deduplication (see DAMP below)
- Run tests after every small change
- This step is mandatory. Skipping it turns TDD into messy code accumulation
Repeat for each new behavior.
Picking What to Test Next
Before writing any code, list the scenarios:
- What's the simplest behavior that must exist?
- What are the edge cases and boundaries?
- What should happen on invalid input?
Work simplest to most complex. Each test should force a small, specific generalization in the production code. As tests get more specific, the code gets more generic — if logically equivalent cases wouldn't pass without changes, the production code is still too specific.
For bug fixes: write a test that reproduces the bug first (it fails), fix it (it passes), refactor. The test is regression protection forever.
Test Structure: Arrange -> Act -> Assert
Arrange (Given): set up the object/state under test
Act (When): call the behavior being tested
Assert (Then): verify the result matches expectation
One behavior per test. Multiple assertions are fine if they verify the same behavior — split the test if they don't.
Keep Tests DAMP, Not DRY
Production code wants DRY. Test code wants DAMP — Descriptive And Meaningful Phrases. A test is a specification, so it has to be readable straight through without chasing helpers.
- Repeating the input shape in every test is fine, and usually better than a shared factory.
- Extract a helper only for noise that is not part of the specification (a connection, an auth token, a temp directory). Never for the values the assertion depends on.
- If knowing what a test verifies means jumping to
beforeEachormake_fixture(), the abstraction cost more than the duplication did.
Duplication across tests is acceptable whenever it keeps each test independently understandable.
Test Doubles
Default to real objects. Only introduce doubles when the real dependency is:
- Slow (database, network, filesystem, external API)
- Non-deterministic (current time, randomness, third-party service)
- Hard to set up to the required state
Then pick the least fake thing that works:
real > fake > stub > mock
highest confidence ..... lowest confidence
- real — the actual collaborator. Catches real bugs.
- fake — a simplified working implementation (in-memory repository, a clock you control).
- stub — canned return values, no behavior.
- mock — asserts which calls happened. Outermost boundary only: payment gateway, email, third-party API.
- spy — records calls for a later assertion. Same caution as a mock.
Each step to the right buys speed and pays for it in confidence. A mock-heavy suite goes green while production breaks.
Verify state, not interactions — the classical, or Detroit, approach:
# Good — asserts the outcome
tasks = list_tasks(sort_by="created_at", order="desc")
assert tasks[0].created_at > tasks[1].created_at
# Bad — asserts the mechanism
list_tasks(sort_by="created_at", order="desc")
assert db.query.called_with(containing("ORDER BY created_at DESC"))
The second test breaks when you swap the query builder, even though the behavior never changed. A suite that fails on refactors nobody broke is a suite people stop trusting.
Special Cases
Legacy code without tests
You can't safely refactor untested code. The approach:
- Write characterization tests that capture current behavior (even if buggy)
- Once covered, modify using normal Red-Green-Refactor
- For known bugs: write a test exposing the bug, then fix it
Spikes (unclear requirements)
TDD requires knowing what "correct" looks like. If you don't:
- Do a spike — exploratory code, no tests, to understand the problem
- Throw the spike away
- Write the real implementation test-first, informed by what you learned
Never let spike code become production code.
The Test Pyramid
/\
/E2E\ few, slow — critical user journeys only
/------\
/ Integr \ moderate — components work together
/----------\
/ Unit (TDD)\ many, fast, cheap — TDD's home
/--------------\
TDD unit tests are the foundation but not sufficient alone. Combine with integration and acceptance tests for full coverage.
Rationalizations
Each of these shows up exactly when the discipline would have paid off.
| Rationalization | Reality |
|---|---|
| "I'll write the test after the code works" | Test-after tests the implementation you just wrote, bugs included. The design pressure is already spent. |
| "This is too simple to break" | Simple code accumulates conditions. The test is what records the intended behavior. |
| "TDD is slower" | Slower for this change. Faster for every later change to the same code. |
| "I verified it by hand" | A manual check does not persist and does not run in CI. Tomorrow's edit breaks it silently. |
| "Tests are green, that's good enough" | Green that was never red proves nothing — you never watched the test fail. |
| "It's only a prototype" | Prototypes ship. This is where test debt starts. |
| "Let me re-run the suite to be sure" | A repeat run on unchanged code adds no information. Re-run after an edit, not for reassurance. |
Pitfalls
| Pitfall | Fix |
|---|---|
| Skipping refactor step | It's mandatory — the #1 TDD failure mode |
| Tests depending on each other | Each test must be fully independent |
| Testing implementation, not behavior | Test interfaces and outputs, not internal calls |
| Over-mocking | Default to real objects; mock only what's genuinely awkward |
| Writing code before a failing test | If there's no red, go back and write the test |
| Editing the test's expected value to match what the implementation produced | Tests specify the contract; implementation conforms to tests. If the assertion was wrong or over-specific, change it deliberately (back to red, then green again) — not silently while debugging to green. If the exact value was incidental, assert the looser contract (type, shape, key invariant) instead of pinning a string you guessed at. |
Before You Call It Done
- Every new behavior has a test that failed before the code existed
- Each red was observed, and it failed for the right reason — not an import or syntax error
- Every bug fix carries a reproduction test that failed before the fix
- The refactor step ran, on production code and test code
- Full suite run after the last edit
- Nothing skipped, disabled,
xfailed, or commented out to reach green - No assertion was loosened or edited to match implementation output
- Test names read as behavior descriptions
Report the command you ran and what it printed. "Tests should pass" is not a result.