# Testing Strategy

> Design a test suite for an open-source project that contributors can actually run and trust. Use when a project has no tests, when the suite is slow or flaky, when deciding what to test and at what level, when setting a coverage policy, or when contributors' PRs keep failing CI for unrelated reasons. Covers the test pyramid, fixtures, property-based and snapshot testing, cross-version matrices, flake elimination, and making tests runnable in one command. Also use for "how do I test this" and for reviewing test quality in a PR.

- Skill: `the-open-agent/testing-strategy` (Agent Skill)
- Install (CLI): `npx skillmds@latest add the-open-agent/testing-strategy`
- Raw SKILL.md: https://api.skillmd.com/api/skills/the-open-agent/testing-strategy/raw
- Safety review: pending
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Coding & Dev Tools
- Author: the-open-agent (https://skillmd.com/u/the-open-agent)
- Updated: 2026-09-17
- Page: https://skillmd.com/skills/the-open-agent/testing-strategy

---


# Testing Strategy

For an open-source project, the test suite has a second job beyond correctness: it is
the **contract that lets strangers change your code safely**. Without it, every
external PR is a risk you must personally verify, and you become the bottleneck.

## The one-command rule

A new contributor must be able to run the full suite in one command, from a fresh
clone, with no tribal knowledge:

```bash
make test          # or: npm test / pytest / cargo test / go test ./...
```

Requirements, all of them non-negotiable:

- **No manual database setup.** Use containers, or a file-backed engine in tests.
- **No secrets required.** Tests needing credentials are skipped by default with a
  clear message, and run only in CI with secrets present.
- **No network by default.** A suite that fails on a plane fails in a corporate proxy
  too, and looks like your bug.
- **Deterministic.** Seed randomness, freeze time, sort collections before comparing.
- **Fast enough to run before every commit.** Target under 60 seconds for the default
  suite; push the slow tests behind a flag.

If setup takes more than `git clone && make test`, contributions drop measurably. This
is the highest-leverage testing investment an OSS project can make.

## The pyramid, and where projects get it wrong

| Level | Share | Runtime | Tests |
|---|---|---|---|
| **Unit** | ~70% | ms | Pure logic, edge cases, error paths |
| **Integration** | ~20% | ~s | Component boundaries, real DB/filesystem |
| **End-to-end** | ~10% | ~10s+ | The two or three flows that must never break |

Two common failure shapes:

- **Ice cream cone** (mostly E2E) — slow, flaky, and when it fails it doesn't tell you
  where. Symptom: "just re-run CI" is normal team advice.
- **Hourglass** (units + E2E, no integration) — every component works, the assembly
  doesn't. Symptom: bugs are always at the seams.

Test the *public API*, not internals. Tests bound to private functions turn every
refactor into a test rewrite, which teaches contributors that tests are an obstacle.

## What to actually test

Prioritize by consequence-of-failure, not by ease:

1. **The documented behavior.** Every promise in the README should have a test. This
   catches the most damaging class of bug — the one that makes your docs a lie.
2. **Error paths.** Most projects test the happy path and ship broken error handling,
   which users hit precisely when they are already frustrated.
3. **Boundaries.** Empty, one, many, max. Zero, negative, off-by-one. Unicode,
   emoji, RTL text, embedded newlines. Paths with spaces. Windows line endings.
4. **Every bug you fix.** A regression test at fix time is the cheapest test you will
   ever write, and it names its own reason for existing.
5. **The public API shape itself** — snapshot the exported surface so accidental
   breaking changes fail CI (see `api-design`).

Do not test: framework behavior, third-party libraries, getters and setters, or
anything whose test is a restatement of the implementation.

## Techniques worth the investment

**Table-driven tests** — the highest density of coverage per line of test code, and
the easiest for contributors to extend when they find a new case.

```go
tests := []struct{ name, in, want string }{
    {"empty", "", ""},
    {"unicode", "café", "cafe"},
    {"emoji", "a👍b", "ab"},
}
for _, tt := range tests {
    t.Run(tt.name, func(t *testing.T) { /* ... */ })
}
```

**Property-based testing** (Hypothesis, fast-check, proptest, QuickCheck) — for
parsers, serializers, encoders, and anything with an inverse. One property replaces a
hundred examples and finds the case you would never have written:

```python
@given(st.text())
def test_roundtrip(s):
    assert decode(encode(s)) == s
```

**Snapshot tests** — good for CLI output, generated code, and rendered markup. Two
rules or they rot: review every snapshot diff as if it were source code, and never run
`--update-snapshots` without reading the diff. An auto-approved snapshot suite tests
nothing.

**Golden files** for compilers, formatters, and codegen: input file → expected output
file, both in the repo, regenerable with one flag.

**Mutation testing** (`mutmut`, `Stryker`, `cargo-mutants`) — the honest answer to "is
our coverage real". Expensive; run it quarterly or in a nightly job, not per-PR.

## Coverage policy

Coverage is a *diagnostic*, not a target. Enforcing 100% produces tests written to
satisfy the instrument.

A policy that works in practice:

- **Track coverage; do not gate on the absolute number.** Publish it, watch the trend.
- **Gate on coverage of the diff** — new code should come with tests. Codecov and
  similar report this natively; it is a much fairer bar for contributors than a global
  percentage they cannot influence.
- **Exempt generated code, `__main__` blocks, and platform-specific branches** in
  config, so the number means something.
- **Uncovered error paths are the signal that matters.** If your happy path is at 95%
  and error handling is at 20%, coverage tooling is telling you exactly what to do.

## Cross-version and cross-platform

Open-source code runs where you do not. Decide the support matrix explicitly, publish
it, and test it — an untested claim of Windows support is worse than no claim.

```yaml
strategy:
  fail-fast: false
  matrix:
    os: [ubuntu-latest, macos-latest, windows-latest]
    version: ['3.10', '3.13']          # oldest supported + newest
    include:
      - { os: ubuntu-latest, version: '3.11' }   # fill the middle on one OS only
```

`fail-fast: false` is important: you want to see *all* failing cells, not the first.
Full cartesian matrices waste CI minutes — test the corners exhaustively and the middle
on one platform. See `ci-pipelines`.

Windows deserves specific attention because it is where cross-platform bugs live: path
separators, `CRLF` vs `LF`, case-insensitive filesystems, `PATH` length limits, and
no `fork()`.

## Killing flakes

A flaky test is worse than no test: it trains everyone, including you, to ignore red
CI. Treat a flake as a P1 bug.

Root causes, in the order you should suspect them:

1. **Time** — `sleep()` as synchronization, tests that fail near midnight or DST
   boundaries, timeouts tuned to a fast laptop. Fix: inject a clock; poll for a
   condition with a generous timeout instead of sleeping a fixed duration.
2. **Order dependence** — shared global state, leaked singletons, a database not
   reset between tests. Fix: run with `--shuffle`/`-p` random ordering in CI so
   order dependence fails immediately rather than mysteriously.
3. **Concurrency** — real races in the code under test. These are the valuable ones;
   the test is telling you the truth.
4. **Network and external services** — never in unit tests; stub at the boundary.
5. **Resource contention** — ports, temp file names, CPU starvation on shared runners.
   Fix: bind to port 0 and read back the assigned port; use unique temp dirs.

Process: quarantine the flake into a separate job immediately so it stops blocking
contributors, open an issue with the failing logs, and fix or delete it within a
sprint. A permanently quarantined test is a deleted test with extra steps — be honest
and delete it.

## Contributor-facing test docs

In `CONTRIBUTING.md`, concretely:

```markdown
## Tests
make test            # full suite (~40s)
make test-unit       # fast, no containers
make test-e2e        # requires Docker
pytest tests/test_parser.py::test_unicode -x   # single test
```

Also state: what a new feature needs (tests + docs), what a bug fix needs (a
regression test), and what happens when CI is red for unrelated reasons — because a
contributor who cannot tell "my fault" from "not my fault" usually just leaves.

## Anti-patterns

- **Tests requiring undocumented local setup.** The contribution killer.
- **`sleep(2)` as synchronization.** Slow *and* flaky.
- **Asserting on log output** for behavior that has a return value.
- **One test asserting fifteen things.** When it fails you learn nothing.
- **Mocking the thing under test.** You are testing your mock.
- **Mocks so deep the test encodes the implementation.** Any refactor breaks it.
- **Skipped tests with no issue link.** Delete them or fix them.
- **`if (CI) skip()`.** The test is a lie in exactly the environment that matters.
- **Coverage gates that reject a one-line docs-adjacent fix.** You just taught a
  first-time contributor that this project is not worth the trouble.

