Testing Pyramid Reference
Target Agents
expert-testing - Primary: applies patterns during test creation and coverage analysis
manager-tdd - Secondary: applies during RED-GREEN-REFACTOR cycles
Test Pyramid Ratios
/ E2E \ 10% — Critical user journeys only
/----------\
/ Integration \ 20% — API endpoints, DB queries, service boundaries
/----------------\
/ Unit Tests \ 70% — Functions, hooks, utilities, pure logic
/--------------------\
| Level |
Speed |
Reliability |
Maintenance |
Coverage Target |
| Unit |
Fast (<100ms) |
High |
Low |
70% of tests |
| Integration |
Medium (1-5s) |
Medium |
Medium |
20% of tests |
| E2E |
Slow (10-60s) |
Lower |
High |
10% of tests |
Coverage Targets by Context
| Context |
Target |
Rationale |
| Critical business logic |
95%+ |
Revenue/security impact |
| API endpoints |
90%+ |
Contract compliance |
| Utility functions |
85%+ |
Reuse reliability |
| UI components |
80%+ |
Rendering correctness |
| Configuration/glue code |
60%+ |
Low complexity |
| Generated code |
0% |
Don't test generated code |
Test Pattern: AAA (Arrange-Act-Assert)
// Arrange: Set up test data and preconditions
input := CreateTestUser("test@example.com")
// Act: Execute the function under test
result, err := service.CreateUser(ctx, input)
// Assert: Verify the outcome
assert.NoError(t, err)
assert.Equal(t, "test@example.com", result.Email)
Unit Test Patterns
| Pattern |
When |
Example |
| Table-Driven |
Multiple input/output combinations |
Go: tests := []struct{...} |
| Mock/Stub |
External dependencies (DB, API) |
Interface injection, mock frameworks |
| Snapshot |
Complex output comparison |
Jest snapshots, golden files |
| Property-Based |
Mathematical properties |
quickcheck, hypothesis |
| Boundary Value |
Edge cases |
0, -1, MAX_INT, empty string, nil |
Integration Test Patterns
| Pattern |
When |
Example |
| Testcontainers |
Real DB needed |
Docker-based PostgreSQL for tests |
| HTTP Test Server |
API endpoint testing |
httptest.NewServer (Go), supertest (Node) |
| In-Memory DB |
Fast DB tests |
SQLite for development |
| Fixture Loading |
Consistent test data |
Factory functions, seed files |
What to Test vs What NOT to Test
ALWAYS Test
- Business logic and calculations
- Input validation and error handling
- Authentication and authorization flows
- Data transformations and mappings
- Edge cases and boundary conditions
- Race conditions (with -race flag in Go)
NEVER Test
- Framework internals (React rendering, Express routing)
- Third-party library behavior
- Simple getters/setters with no logic
- Private methods directly (test via public API)
- Generated code (protobuf, swagger)
- CSS styling and layout (use visual regression tools instead)
Test Quality Metrics
| Metric |
Target |
Tool |
| Line Coverage |
85%+ |
go test -cover, istanbul, coverage.py |
| Branch Coverage |
75%+ |
go test -covermode=count |
| Mutation Score |
70%+ |
go-mutesting, Stryker |
| Test Execution Time |
<2 min (unit), <10 min (all) |
CI timer |
| Flaky Test Rate |
<1% |
CI history analysis |
Test File Conventions
| Language |
Test File |
Location |
| Go |
*_test.go |
Same package |
| TypeScript |
*.test.ts / *.spec.ts |
__tests__/ or co-located |
| Python |
test_*.py |
tests/ directory |
| Java |
*Test.java |
src/test/ mirror |
| Rust |
#[cfg(test)] mod tests |
Same file or tests/ |
TDD RED-GREEN-REFACTOR Quick Reference
RED: Write a failing test that defines expected behavior
GREEN: Write minimal code to make the test pass
REFACTOR: Clean up while keeping tests green
Rules:
- Never write production code without a failing test
- Write the smallest test that fails
- Write the simplest code that passes
- Refactor only when all tests are green
- One assertion per test (when practical)
Common Rationalizations
| Rationalization |
Reality |
| "E2E tests cover everything, unit tests are redundant" |
E2E tests are slow and flaky. Unit tests provide fast, precise feedback. The pyramid exists because each level serves a different purpose. |
| "Integration tests are more realistic than unit tests" |
Realism comes at the cost of speed and isolation. A balanced pyramid gives both fast feedback and realistic validation. |
| "100% code coverage means the code is well tested" |
Coverage measures execution, not correctness. A test that executes code without meaningful assertions provides zero value. |
| "Mocking is bad, I prefer real dependencies" |
Real dependencies make tests slow and non-deterministic. Mock at boundaries, test business logic in isolation. |
| "This test is flaky, but it catches real bugs sometimes" |
Flaky tests erode trust in the entire suite. Fix the flakiness or quarantine the test with a tracking issue. |
DAMP over DRY: Test code should be descriptive and self-contained. A reader should understand the test without reading shared fixtures or helper methods.
Red Flags
- Test pyramid inverted: more E2E tests than unit tests
- Unit tests depend on external services (databases, APIs, file systems)
- Test assertions check implementation details instead of behavior
- No integration tests between unit and E2E layers
- Flaky test present without a quarantine label or tracking issue
Verification
1---2name: moai-ref-testing-pyramid3description: Test pyramid strategy, coverage targets, test patterns, and quality metrics reference. Agent-extending skill that amplifies expert-testing and manager-tdd expertise with production-grade testing patterns. NOT for: production code implementation, architecture design, DevOps, security audits.4---56# Testing Pyramid Reference78## Target Agents910- `expert-testing` - Primary: applies patterns during test creation and coverage analysis11- `manager-tdd` - Secondary: applies during RED-GREEN-REFACTOR cycles1213## Test Pyramid Ratios1415```16 / E2E \ 10% — Critical user journeys only17 /----------\18 / Integration \ 20% — API endpoints, DB queries, service boundaries19 /----------------\20 / Unit Tests \ 70% — Functions, hooks, utilities, pure logic21 /--------------------\22```2324| Level | Speed | Reliability | Maintenance | Coverage Target |25|-------|-------|-------------|-------------|-----------------|26| Unit | Fast (<100ms) | High | Low | 70% of tests |27| Integration | Medium (1-5s) | Medium | Medium | 20% of tests |28| E2E | Slow (10-60s) | Lower | High | 10% of tests |2930## Coverage Targets by Context3132| Context | Target | Rationale |33|---------|--------|-----------|34| Critical business logic | 95%+ | Revenue/security impact |35| API endpoints | 90%+ | Contract compliance |36| Utility functions | 85%+ | Reuse reliability |37| UI components | 80%+ | Rendering correctness |38| Configuration/glue code | 60%+ | Low complexity |39| Generated code | 0% | Don't test generated code |4041## Test Pattern: AAA (Arrange-Act-Assert)4243```44// Arrange: Set up test data and preconditions45input := CreateTestUser("test@example.com")4647// Act: Execute the function under test48result, err := service.CreateUser(ctx, input)4950// Assert: Verify the outcome51assert.NoError(t, err)52assert.Equal(t, "test@example.com", result.Email)53```5455## Unit Test Patterns5657| Pattern | When | Example |58|---------|------|---------|59| Table-Driven | Multiple input/output combinations | Go: `tests := []struct{...}` |60| Mock/Stub | External dependencies (DB, API) | Interface injection, mock frameworks |61| Snapshot | Complex output comparison | Jest snapshots, golden files |62| Property-Based | Mathematical properties | quickcheck, hypothesis |63| Boundary Value | Edge cases | 0, -1, MAX_INT, empty string, nil |6465## Integration Test Patterns6667| Pattern | When | Example |68|---------|------|---------|69| Testcontainers | Real DB needed | Docker-based PostgreSQL for tests |70| HTTP Test Server | API endpoint testing | httptest.NewServer (Go), supertest (Node) |71| In-Memory DB | Fast DB tests | SQLite for development |72| Fixture Loading | Consistent test data | Factory functions, seed files |7374## What to Test vs What NOT to Test7576### ALWAYS Test77- Business logic and calculations78- Input validation and error handling79- Authentication and authorization flows80- Data transformations and mappings81- Edge cases and boundary conditions82- Race conditions (with -race flag in Go)8384### NEVER Test85- Framework internals (React rendering, Express routing)86- Third-party library behavior87- Simple getters/setters with no logic88- Private methods directly (test via public API)89- Generated code (protobuf, swagger)90- CSS styling and layout (use visual regression tools instead)9192## Test Quality Metrics9394| Metric | Target | Tool |95|--------|--------|------|96| Line Coverage | 85%+ | go test -cover, istanbul, coverage.py |97| Branch Coverage | 75%+ | go test -covermode=count |98| Mutation Score | 70%+ | go-mutesting, Stryker |99| Test Execution Time | <2 min (unit), <10 min (all) | CI timer |100| Flaky Test Rate | <1% | CI history analysis |101102## Test File Conventions103104| Language | Test File | Location |105|----------|-----------|----------|106| Go | `*_test.go` | Same package |107| TypeScript | `*.test.ts` / `*.spec.ts` | `__tests__/` or co-located |108| Python | `test_*.py` | `tests/` directory |109| Java | `*Test.java` | `src/test/` mirror |110| Rust | `#[cfg(test)] mod tests` | Same file or `tests/` |111112## TDD RED-GREEN-REFACTOR Quick Reference113114```115RED: Write a failing test that defines expected behavior116GREEN: Write minimal code to make the test pass117REFACTOR: Clean up while keeping tests green118```119120Rules:121- Never write production code without a failing test122- Write the smallest test that fails123- Write the simplest code that passes124- Refactor only when all tests are green125- One assertion per test (when practical)126127<!-- moai:evolvable-start id="rationalizations" -->128## Common Rationalizations129130| Rationalization | Reality |131|---|---|132| "E2E tests cover everything, unit tests are redundant" | E2E tests are slow and flaky. Unit tests provide fast, precise feedback. The pyramid exists because each level serves a different purpose. |133| "Integration tests are more realistic than unit tests" | Realism comes at the cost of speed and isolation. A balanced pyramid gives both fast feedback and realistic validation. |134| "100% code coverage means the code is well tested" | Coverage measures execution, not correctness. A test that executes code without meaningful assertions provides zero value. |135| "Mocking is bad, I prefer real dependencies" | Real dependencies make tests slow and non-deterministic. Mock at boundaries, test business logic in isolation. |136| "This test is flaky, but it catches real bugs sometimes" | Flaky tests erode trust in the entire suite. Fix the flakiness or quarantine the test with a tracking issue. |137138**DAMP over DRY**: Test code should be descriptive and self-contained. A reader should understand the test without reading shared fixtures or helper methods.139140<!-- moai:evolvable-end -->141142<!-- moai:evolvable-start id="red-flags" -->143## Red Flags144145- Test pyramid inverted: more E2E tests than unit tests146- Unit tests depend on external services (databases, APIs, file systems)147- Test assertions check implementation details instead of behavior148- No integration tests between unit and E2E layers149- Flaky test present without a quarantine label or tracking issue150151<!-- moai:evolvable-end -->152153<!-- moai:evolvable-start id="verification" -->154## Verification155156- [ ] Test distribution follows the pyramid: unit > integration > E2E (show test counts per category)157- [ ] Unit tests run in under 30 seconds total158- [ ] Integration tests mock external dependencies at the boundary159- [ ] No flaky tests in the active suite (run 3x to verify stability)160- [ ] Test names describe behavior, not implementation (review naming convention)161- [ ] Coverage report shows meaningful assertions, not just line execution162163<!-- moai:evolvable-end -->