Testing Pyramid Reference
Target Agents
manager-develop - Primary: applies patterns during test creation and coverage analysis
manager-develop - Secondary: applies during RED-GREEN-REFACTOR cycles
Test Pyramid Ratios
/ E2E \ 10% — Critical user journeys only
/----------\
/ Integration \ 20% — API endpoints, DB queries, service boundaries
/----------------\
/ Unit Tests \ 70% — Functions, hooks, utilities, pure logic
/--------------------\
| Level |
Speed |
Reliability |
Maintenance |
Coverage Target |
| Unit |
Fast (<100ms) |
High |
Low |
70% of tests |
| Integration |
Medium (1-5s) |
Medium |
Medium |
20% of tests |
| E2E |
Slow (10-60s) |
Lower |
High |
10% of tests |
Coverage Targets by Context
| Context |
Target |
Rationale |
| Critical business logic |
95%+ |
Revenue/security impact |
| API endpoints |
90%+ |
Contract compliance |
| Utility functions |
85%+ |
Reuse reliability |
| UI components |
80%+ |
Rendering correctness |
| Configuration/glue code |
60%+ |
Low complexity |
| Generated code |
0% |
Don't test generated code |
Test Pattern: AAA (Arrange-Act-Assert)
// Arrange: Set up test data and preconditions
input := CreateTestUser("test@example.com")
// Act: Execute the function under test
result, err := service.CreateUser(ctx, input)
// Assert: Verify the outcome
assert.NoError(t, err)
assert.Equal(t, "test@example.com", result.Email)
Unit Test Patterns
| Pattern |
When |
Example |
| Table-Driven |
Multiple input/output combinations |
Go: tests := []struct{...} |
| Mock/Stub |
External dependencies (DB, API) |
Interface injection, mock frameworks |
| Snapshot |
Complex output comparison |
Jest snapshots, golden files |
| Property-Based |
Mathematical properties |
quickcheck, hypothesis |
| Boundary Value |
Edge cases |
0, -1, MAX_INT, empty string, nil |
Integration Test Patterns
| Pattern |
When |
Example |
| Testcontainers |
Real DB needed |
Docker-based PostgreSQL for tests |
| HTTP Test Server |
API endpoint testing |
httptest.NewServer (Go), supertest (Node) |
| In-Memory DB |
Fast DB tests |
SQLite for development |
| Fixture Loading |
Consistent test data |
Factory functions, seed files |
What to Test vs What NOT to Test
ALWAYS Test
- Business logic and calculations
- Input validation and error handling
- Authentication and authorization flows
- Data transformations and mappings
- Edge cases and boundary conditions
- Race conditions (with -race flag in Go)
NEVER Test
- Framework internals (React rendering, Express routing)
- Third-party library behavior
- Simple getters/setters with no logic
- Private methods directly (test via public API)
- Generated code (protobuf, swagger)
- CSS styling and layout (use visual regression tools instead)
Test Quality Metrics
| Metric |
Target |
Tool |
| Line Coverage |
85%+ |
go test -cover, istanbul, coverage.py |
| Branch Coverage |
75%+ |
go test -covermode=count |
| Mutation Score |
70%+ |
go-mutesting, Stryker |
| Test Execution Time |
<2 min (unit), <10 min (all) |
CI timer |
| Flaky Test Rate |
<1% |
CI history analysis |
Test File Conventions
| Language |
Test File |
Location |
| Go |
*_test.go |
Same package |
| TypeScript |
*.test.ts / *.spec.ts |
__tests__/ or co-located |
| Python |
test_*.py |
tests/ directory |
| Java |
*Test.java |
src/test/ mirror |
| Rust |
#[cfg(test)] mod tests |
Same file or tests/ |
TDD RED-GREEN-REFACTOR Quick Reference
RED: Write a failing test that defines expected behavior
GREEN: Write minimal code to make the test pass
REFACTOR: Clean up while keeping tests green
Rules:
- Never write production code without a failing test
- Write the smallest test that fails
- Write the simplest code that passes
- Refactor only when all tests are green
- One assertion per test (when practical)
Common Rationalizations
| Rationalization |
Reality |
| "E2E tests cover everything, unit tests are redundant" |
E2E tests are slow and flaky. Unit tests provide fast, precise feedback. The pyramid exists because each level serves a different purpose. |
| "Integration tests are more realistic than unit tests" |
Realism comes at the cost of speed and isolation. A balanced pyramid gives both fast feedback and realistic validation. |
| "100% code coverage means the code is well tested" |
Coverage measures execution, not correctness. A test that executes code without meaningful assertions provides zero value. |
| "Mocking is bad, I prefer real dependencies" |
Real dependencies make tests slow and non-deterministic. Mock at boundaries, test business logic in isolation. |
| "This test is flaky, but it catches real bugs sometimes" |
Flaky tests erode trust in the entire suite. Fix the flakiness or quarantine the test with a tracking issue. |
DAMP over DRY: Test code should be descriptive and self-contained. A reader should understand the test without reading shared fixtures or helper methods.
Red Flags
- Test pyramid inverted: more E2E tests than unit tests
- Unit tests depend on external services (databases, APIs, file systems)
- Test assertions check implementation details instead of behavior
- No integration tests between unit and E2E layers
- Flaky test present without a quarantine label or tracking issue
Verification
1---2name: moai-ref-testing-pyramid3description: Test pyramid strategy, coverage targets, test patterns, and quality metrics reference. Agent-extending skill that amplifies manager-develop test-creation and quality-validation work with production-grade testing patterns. NOT for: production code implementation, architecture design, DevOps, security audits.4---5
6# Testing Pyramid Reference
7
8## Target Agents
9
10- `manager-develop` - Primary: applies patterns during test creation and coverage analysis
11- `manager-develop` - Secondary: applies during RED-GREEN-REFACTOR cycles
12
13## Test Pyramid Ratios
14
15```
16 / E2E \ 10% — Critical user journeys only
17 /----------\
18 / Integration \ 20% — API endpoints, DB queries, service boundaries
19 /----------------\
20 / Unit Tests \ 70% — Functions, hooks, utilities, pure logic
21 /--------------------\
22```
23
24| Level | Speed | Reliability | Maintenance | Coverage Target |
25|-------|-------|-------------|-------------|-----------------|
26| Unit | Fast (<100ms) | High | Low | 70% of tests |
27| Integration | Medium (1-5s) | Medium | Medium | 20% of tests |
28| E2E | Slow (10-60s) | Lower | High | 10% of tests |
29
30## Coverage Targets by Context
31
32| Context | Target | Rationale |
33|---------|--------|-----------|
34| Critical business logic | 95%+ | Revenue/security impact |
35| API endpoints | 90%+ | Contract compliance |
36| Utility functions | 85%+ | Reuse reliability |
37| UI components | 80%+ | Rendering correctness |
38| Configuration/glue code | 60%+ | Low complexity |
39| Generated code | 0% | Don't test generated code |
40
41## Test Pattern: AAA (Arrange-Act-Assert)
42
43```
44// Arrange: Set up test data and preconditions
45input := CreateTestUser("test@example.com")
46
47// Act: Execute the function under test
48result, err := service.CreateUser(ctx, input)
49
50// Assert: Verify the outcome
51assert.NoError(t, err)
52assert.Equal(t, "test@example.com", result.Email)
53```
54
55## Unit Test Patterns
56
57| Pattern | When | Example |
58|---------|------|---------|
59| Table-Driven | Multiple input/output combinations | Go: `tests := []struct{...}` |
60| Mock/Stub | External dependencies (DB, API) | Interface injection, mock frameworks |
61| Snapshot | Complex output comparison | Jest snapshots, golden files |
62| Property-Based | Mathematical properties | quickcheck, hypothesis |
63| Boundary Value | Edge cases | 0, -1, MAX_INT, empty string, nil |
64
65## Integration Test Patterns
66
67| Pattern | When | Example |
68|---------|------|---------|
69| Testcontainers | Real DB needed | Docker-based PostgreSQL for tests |
70| HTTP Test Server | API endpoint testing | httptest.NewServer (Go), supertest (Node) |
71| In-Memory DB | Fast DB tests | SQLite for development |
72| Fixture Loading | Consistent test data | Factory functions, seed files |
73
74## What to Test vs What NOT to Test
75
76### ALWAYS Test
77- Business logic and calculations
78- Input validation and error handling
79- Authentication and authorization flows
80- Data transformations and mappings
81- Edge cases and boundary conditions
82- Race conditions (with -race flag in Go)
83
84### NEVER Test
85- Framework internals (React rendering, Express routing)
86- Third-party library behavior
87- Simple getters/setters with no logic
88- Private methods directly (test via public API)
89- Generated code (protobuf, swagger)
90- CSS styling and layout (use visual regression tools instead)
91
92## Test Quality Metrics
93
94| Metric | Target | Tool |
95|--------|--------|------|
96| Line Coverage | 85%+ | go test -cover, istanbul, coverage.py |
97| Branch Coverage | 75%+ | go test -covermode=count |
98| Mutation Score | 70%+ | go-mutesting, Stryker |
99| Test Execution Time | <2 min (unit), <10 min (all) | CI timer |
100| Flaky Test Rate | <1% | CI history analysis |
101
102## Test File Conventions
103
104| Language | Test File | Location |
105|----------|-----------|----------|
106| Go | `*_test.go` | Same package |
107| TypeScript | `*.test.ts` / `*.spec.ts` | `__tests__/` or co-located |
108| Python | `test_*.py` | `tests/` directory |
109| Java | `*Test.java` | `src/test/` mirror |
110| Rust | `#[cfg(test)] mod tests` | Same file or `tests/` |
111
112## TDD RED-GREEN-REFACTOR Quick Reference
113
114```
115RED: Write a failing test that defines expected behavior
116GREEN: Write minimal code to make the test pass
117REFACTOR: Clean up while keeping tests green
118```
119
120Rules:
121- Never write production code without a failing test
122- Write the smallest test that fails
123- Write the simplest code that passes
124- Refactor only when all tests are green
125- One assertion per test (when practical)
126
127<!-- moai:evolvable-start id="rationalizations" -->
128## Common Rationalizations
129
130| Rationalization | Reality |
131|---|---|
132| "E2E tests cover everything, unit tests are redundant" | E2E tests are slow and flaky. Unit tests provide fast, precise feedback. The pyramid exists because each level serves a different purpose. |
133| "Integration tests are more realistic than unit tests" | Realism comes at the cost of speed and isolation. A balanced pyramid gives both fast feedback and realistic validation. |
134| "100% code coverage means the code is well tested" | Coverage measures execution, not correctness. A test that executes code without meaningful assertions provides zero value. |
135| "Mocking is bad, I prefer real dependencies" | Real dependencies make tests slow and non-deterministic. Mock at boundaries, test business logic in isolation. |
136| "This test is flaky, but it catches real bugs sometimes" | Flaky tests erode trust in the entire suite. Fix the flakiness or quarantine the test with a tracking issue. |
137
138**DAMP over DRY**: Test code should be descriptive and self-contained. A reader should understand the test without reading shared fixtures or helper methods.
139
140<!-- moai:evolvable-end -->
141
142<!-- moai:evolvable-start id="red-flags" -->
143## Red Flags
144
145- Test pyramid inverted: more E2E tests than unit tests
146- Unit tests depend on external services (databases, APIs, file systems)
147- Test assertions check implementation details instead of behavior
148- No integration tests between unit and E2E layers
149- Flaky test present without a quarantine label or tracking issue
150
151<!-- moai:evolvable-end -->
152
153<!-- moai:evolvable-start id="verification" -->
154## Verification
155
156- [ ] Test distribution follows the pyramid: unit > integration > E2E (show test counts per category)
157- [ ] Unit tests run in under 30 seconds total
158- [ ] Integration tests mock external dependencies at the boundary
159- [ ] No flaky tests in the active suite (run 3x to verify stability)
160- [ ] Test names describe behavior, not implementation (review naming convention)
161- [ ] Coverage report shows meaningful assertions, not just line execution
162
163<!-- moai:evolvable-end -->