QA Testing Strategy (Dec 2025) — Quick Reference
Use this skill when the primary focus is how to test software effectively (risk-first, automation-first, observable systems) rather than how to implement features.
Core references: SLO/error budgets and troubleshooting patterns from the Google SRE Book (Service Level Objectives, Effective Troubleshooting); contract-driven API documentation via OpenAPI (OAS); and E2E ergonomics/practices via Playwright docs (Best Practices).
Core QA (Default)
Outcomes (Definition of Done)
- Strategy is risk-based: critical user journeys + likely failure modes are explicit.
- Test portfolio is layered: fastest checks catch most defects; slow checks are minimal and high-signal.
- CI is economical: fast pre-merge gates; heavy suites are scheduled or scoped.
- Failures are diagnosable: every failure yields actionable artifacts (logs/trace/screenshots/crash reports).
- Flakes are managed as reliability debt with an SLO and a deflake runbook.
Quality Model: Risk, Journeys, Failure Modes
- Model risk as
impact x likelihood x detectability per journey.
- Write failure modes per journey: auth/session, permissions, data integrity, dependency failure, latency, offline/degraded UX, concurrency/races.
- Define oracles per test: business rule oracle, contract/schema oracle, security oracle, accessibility oracle, performance oracle.
Test Portfolio (Modern Equivalent of the Pyramid)
- Prefer unit + component + contract + integration as the default safety net; keep E2E for thin, critical journeys.
- Add exploratory testing for discovery and usability; convert findings into automated checks when ROI is positive.
- Use the smallest scope that detects the bug class:
- Bug in business logic: unit/property-based.
- Bug in service wiring/data: integration/contract.
- Bug in cross-service compatibility: contract + a small number of integration scenarios.
- Bug in user journey: E2E (critical path only).
Shift-Left Gates (Pre-Merge by Default)
- Contracts: OpenAPI/AsyncAPI/JSON Schema validation where applicable (OpenAPI, AsyncAPI, JSON Schema).
- Static checks: lint, typecheck, dependency scanning, secret scanning.
- Fast tests: unit + key integration checks; avoid full E2E as a PR gate unless the product is E2E-only.
Coverage Model (Explicitly Separate)
- Code coverage answers: “What code executed?” (useful as a smoke signal).
- Risk coverage answers: “What user/business risk is reduced?” (the real target).
- REQUIRED: every critical journey has at least one automated check at the lowest effective layer.
CI Economics (Budgets and Levers)
- Budgets [Inference]:
- PR gate: p50 <= 10 min, p95 <= 20 min.
- Mainline health: >= 99% green builds per day.
- Levers:
- Parallelize by layer and shard long-running suites (Playwright supports sharding in CI: Sharding).
- Cache dependencies and test artifacts where your CI supports it.
- Run “full regression” on schedule (nightly) and “risk-scoped regression” on PRs.
Flake Management (SLO + Runbook)
- Define flake: test fails without product change and passes on rerun.
- Track flake rate as:
rerun_pass / (rerun_pass + rerun_fail) for a suite.
- SLO examples [Inference]:
- Suite flake rate <= 1% weekly.
- Time-to-deflake: p50 <= 2 business days, p95 <= 7 business days.
- REQUIRED: quarantine policy and deflake playbook:
templates/runbooks/template-flaky-test-triage-deflake-runbook.md.
- For rate-limited endpoints, run serially, reuse tokens, and add backoff for 429s; isolate 429 tests to avoid poisoning other suites.
Debugging Ergonomics (Make Failures Cheap)
- Always capture first failure context: request IDs, trace IDs, build URL, seed, test data IDs.
- Standardize artifacts per layer:
- Unit/integration: structured logs + minimal repro input.
- E2E: trace + screenshot/video on failure (Playwright tooling: Trace Viewer).
- Mobile:
xcresult bundles + screenshots + device logs.
Do / Avoid
Do:
- Write tests against stable contracts and user-visible behavior.
- Treat flaky tests as P1 reliability work; quarantine only with an owner and expiry.
- Make “how to debug this failure” part of every suite’s definition of done.
Avoid:
- “Everything E2E” as a default (slow, expensive, low-signal).
- Sleeps/time-based waits (prefer assertions and event-based waits).
- Using coverage % as the primary quality KPI (use risk coverage + defect escape rate).
When to Use This Skill
Invoke when users ask for:
- Test strategy for a new service or feature
- Unit testing with Jest or Vitest
- Integration testing with databases, APIs, external services
- E2E testing with Playwright or Cypress
- Performance and load testing with k6
- BDD with Cucumber and Gherkin
- API contract testing with Pact
- Visual regression testing
- Test automation CI/CD integration
- Test data management and fixtures
- Security and accessibility testing
- Test coverage analysis and improvement
- Flaky test diagnosis and fixes
- Mobile app testing (iOS/Android)
Quick Reference Table
| Test Type |
Framework |
Command |
When to Use |
| Unit Tests |
Vitest |
vitest run |
Pure functions, business logic (40-60% of tests) |
| Component Tests |
React Testing Library |
vitest --ui |
React components, user interactions (20-30%) |
| Integration Tests |
Supertest + Docker |
vitest run integration.test.ts |
API endpoints, database operations (15-25%) |
| E2E Tests |
Playwright |
playwright test |
Critical user journeys, cross-browser (5-10%) |
| Performance Tests |
k6 |
k6 run load-test.js |
Load testing, stress testing (nightly/pre-release) |
| API Contract Tests |
Pact |
pact test |
Microservices, consumer-provider contracts |
| Visual Regression |
Percy/Chromatic |
percy snapshot |
UI consistency, design system validation |
| Security Tests |
OWASP ZAP |
zap-baseline.py |
Vulnerability scanning (every PR) |
| Accessibility Tests |
axe-core |
vitest run a11y.test.ts |
WCAG compliance (every component) |
| Mutation Tests |
Stryker |
stryker run |
Test quality validation (weekly) |
Decision Tree: Test Strategy
Need to test: [Feature Type]
│
├─ Pure business logic?
│ └─ Unit tests (Jest/Vitest) — Fast, isolated, AAA pattern
│ ├─ Has dependencies? → Mock them
│ ├─ Complex calculations? → Property-based testing (fast-check)
│ └─ State machine? → State transition tests
│
├─ UI Component?
│ ├─ Isolated component?
│ │ └─ Component tests (React Testing Library)
│ │ ├─ User interactions → fireEvent/userEvent
│ │ └─ Accessibility → axe-core integration
│ │
│ └─ User journey?
│ └─ E2E tests (Playwright)
│ ├─ Critical path → Always test
│ ├─ Edge cases → Selective E2E
│ └─ Visual → Percy/Chromatic
│
├─ API Endpoint?
│ ├─ Single service?
│ │ └─ Integration tests (Supertest + test DB)
│ │ ├─ CRUD operations → Test all verbs
│ │ ├─ Auth/permissions → Test unauthorized paths
│ │ └─ Error handling → Test error responses
│ │
│ └─ Microservices?
│ └─ Contract tests (Pact) + integration tests
│ ├─ Consumer defines expectations
│ └─ Provider verifies contracts
│
├─ Performance-critical?
│ ├─ Load capacity?
│ │ └─ k6 load testing (ramp-up, stress, spike)
│ │
│ └─ Response time?
│ └─ k6 performance benchmarks (SLO validation)
│
└─ External dependency?
├─ Mock it (unit tests) → Use test doubles
└─ Real implementation (integration) → Docker containers (Testcontainers)
Decision Tree: Choosing Test Framework
What are you testing?
│
├─ JavaScript/TypeScript?
│ ├─ New project? → Vitest (faster, modern)
│ ├─ Existing Jest project? → Keep Jest
│ └─ Browser-specific? → Playwright component testing
│
├─ Python?
│ ├─ General testing? → pytest
│ ├─ Django? → pytest-django
│ └─ FastAPI? → pytest + httpx
│
├─ Go?
│ ├─ Unit tests? → testing package
│ ├─ Mocking? → gomock or testify
│ └─ Integration? → testcontainers-go
│
├─ Rust?
│ ├─ Unit tests? → Built-in #[test]
│ └─ Property-based? → proptest
│
└─ E2E (any language)?
├─ Web app? → Playwright (recommended)
├─ API only? → k6 or Postman/Newman
└─ Mobile? → Detox (RN), XCUITest (iOS), Espresso (Android)
Decision Tree: Flaky Test Diagnosis
Test is flaky?
│
├─ Timing-related?
│ ├─ Race condition? → Add proper waits (not sleep)
│ ├─ Animation? → Disable animations in test mode
│ └─ Network timeout? → Increase timeout, add retry
│
├─ Data-related?
│ ├─ Shared state? → Isolate test data
│ ├─ Random data? → Use seeded random
│ └─ Order-dependent? → Fix test isolation
│
├─ Environment-related?
│ ├─ CI-only failures? → Check resource constraints
│ ├─ Timezone issues? → Use UTC in tests
│ └─ Locale issues? → Set consistent locale
│
└─ External dependency?
├─ Third-party API? → Mock it
└─ Database? → Use test containers
Test Pyramid
/\
/ \
/ E2E \ 5-10% - Critical user journeys
/--------\ - Slow, expensive, high confidence
/Integration\ 15-25% - API, database, services
/--------------\ - Medium speed, good coverage
/ Unit \ 40-60% - Functions, components
/------------------\ - Fast, cheap, foundation
Target coverage by layer:
| Layer |
Coverage |
Speed |
Confidence |
| Unit |
80%+ |
~1000/sec |
Low (isolated) |
| Integration |
70%+ |
~10/sec |
Medium |
| E2E |
Critical paths |
~1/sec |
High |
Core Capabilities
Unit Testing
- Frameworks: Vitest, Jest, pytest, Go testing
- Patterns: AAA (Arrange-Act-Assert), Given-When-Then
- Mocking: Dependency injection, test doubles
- Coverage: Line, branch, function coverage
Integration Testing
- Database: Testcontainers, in-memory DBs
- API: Supertest, httpx, REST-assured
- Services: Docker Compose, localstack
- Fixtures: Factory patterns, seeders
E2E Testing
- Web: Playwright, Cypress
- Mobile: Detox, XCUITest, Espresso
- API: k6, Postman/Newman
- Patterns: Page Object Model, test locators
Performance Testing
- Load: k6, Locust, Gatling
- Profiling: Browser DevTools, Lighthouse
- Monitoring: Real User Monitoring (RUM)
- Benchmarks: Response time, throughput, error rate
Common Patterns
AAA Pattern (Arrange-Act-Assert)
describe('calculateDiscount', () => {
it('should apply 10% discount for orders over $100', () => {
// Arrange
const order = { total: 150, customerId: 'user-1' };
// Act
const result = calculateDiscount(order);
// Assert
expect(result.discount).toBe(15);
expect(result.finalTotal).toBe(135);
});
});
Page Object Model (E2E)
// pages/login.page.ts
class LoginPage {
async login(email: string, password: string) {
await this.page.fill('[data-testid="email"]', email);
await this.page.fill('[data-testid="password"]', password);
await this.page.click('[data-testid="submit"]');
}
async expectLoggedIn() {
await expect(this.page.locator('[data-testid="dashboard"]')).toBeVisible();
}
}
// tests/login.spec.ts
test('user can login with valid credentials', async ({ page }) => {
const loginPage = new LoginPage(page);
await loginPage.login('user@example.com', 'password');
await loginPage.expectLoggedIn();
});
Test Data Factory
// factories/user.factory.ts
export const createUser = (overrides = {}) => ({
id: faker.string.uuid(),
email: faker.internet.email(),
name: faker.person.fullName(),
createdAt: new Date(),
...overrides,
});
// Usage in tests
const admin = createUser({ role: 'admin' });
const guest = createUser({ role: 'guest', email: 'guest@test.com' });
CI/CD Integration
GitHub Actions Example
name: Test Suite
on: [push, pull_request]
jobs:
unit-tests:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
- run: npm ci
- run: npm run test:unit -- --coverage
- uses: codecov/codecov-action@v3
integration-tests:
runs-on: ubuntu-latest
services:
postgres:
image: postgres:15
env:
POSTGRES_PASSWORD: test
steps:
- uses: actions/checkout@v4
- run: npm ci
- run: npm run test:integration
e2e-tests:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- run: npm ci
- run: npx playwright install --with-deps
- run: npm run test:e2e
Quality Gates
| Gate |
Threshold |
Action on Failure |
| Unit test coverage |
80% |
Block merge |
| All tests pass |
100% |
Block merge |
| No new critical bugs |
0 |
Block merge |
| Performance regression |
<10% |
Warning |
| Security vulnerabilities |
0 critical |
Block deploy |
Anti-Patterns to Avoid
| Anti-Pattern |
Problem |
Solution |
| Testing implementation |
Breaks on refactor |
Test behavior, not internals |
| Shared mutable state |
Flaky tests |
Isolate test data |
| sleep() in tests |
Slow, unreliable |
Use proper waits/assertions |
| Testing everything E2E |
Slow, expensive |
Use test pyramid |
| No test data cleanup |
Test pollution |
Reset state between tests |
| Ignoring flaky tests |
False confidence |
Fix or quarantine immediately |
| Copy-paste tests |
Hard to maintain |
Use factories and helpers |
| Testing third-party code |
Wasted effort |
Trust libraries, test integration |
Optional: AI / Automation
Use AI assistance only as an accelerator for low-risk work; validate outputs with objective checks and evidence.
Do:
- Generate scaffolding (test file skeletons, fixtures) and then harden manually.
- Use AI to propose edge cases, then select based on your risk model and add explicit oracles.
- Use AI to summarize flaky-test clusters, but base actions on logs/traces and rerun evidence.
Avoid:
- Accepting generated assertions without validating the oracle (risk: confident nonsense).
- Letting AI “heal” tests by weakening assertions (risk: silent regressions).
Safety references (optional):
Navigation
Resources
- resources/operational-playbook.md — Testing pyramid guidance, BDD/test data patterns, CI gates, and anti-patterns
- resources/playwright-webapp-testing.md — Playwright decision tree, server lifecycle helper, and recon-first scripting pattern
- resources/comprehensive-testing-guide.md — Full testing methodology reference
- resources/test-automation-patterns.md — Automation patterns and best practices
- resources/shift-left-testing.md — Early testing strategies
Templates
- templates/test-strategy-template.md — Risk-based test strategy one-pager
- templates/template-test-case-design.md — Test case design (Given/When/Then + oracles)
- templates/runbooks/template-flaky-test-triage-deflake-runbook.md — Flake triage + deflake runbook
- templates/automation-pipeline-template.md — CI/CD automation pattern
- templates/unit/template-jest-vitest.md — Unit testing
- templates/integration/template-api-integration.md — Integration/API testing
- templates/e2e/template-playwright.md — Playwright E2E
- templates/bdd/template-cucumber-gherkin.md — BDD/Gherkin
- templates/performance/template-k6-load-testing.md — k6 performance
- templates/visual-regression/template-visual-testing.md — Visual regression
Data
- data/sources.json — Curated external references
Related Skills
1---2name: qa-testing-strategy3description: Risk-based quality engineering test strategy across unit, integration, contract, E2E, performance, and security testing with shift-left gates, flake control, CI economics, and observability-first debugging.4---56# QA Testing Strategy (Dec 2025) — Quick Reference78Use this skill when the primary focus is how to test software effectively (risk-first, automation-first, observable systems) rather than how to implement features.910---1112Core references: SLO/error budgets and troubleshooting patterns from the Google SRE Book ([Service Level Objectives](https://sre.google/sre-book/service-level-objectives/), [Effective Troubleshooting](https://sre.google/sre-book/effective-troubleshooting/)); contract-driven API documentation via OpenAPI ([OAS](https://spec.openapis.org/oas/latest.html)); and E2E ergonomics/practices via Playwright docs ([Best Practices](https://playwright.dev/docs/best-practices)).1314## Core QA (Default)1516### Outcomes (Definition of Done)1718- Strategy is risk-based: critical user journeys + likely failure modes are explicit.19- Test portfolio is layered: fastest checks catch most defects; slow checks are minimal and high-signal.20- CI is economical: fast pre-merge gates; heavy suites are scheduled or scoped.21- Failures are diagnosable: every failure yields actionable artifacts (logs/trace/screenshots/crash reports).22- Flakes are managed as reliability debt with an SLO and a deflake runbook.2324### Quality Model: Risk, Journeys, Failure Modes2526- Model risk as `impact x likelihood x detectability` per journey.27- Write failure modes per journey: auth/session, permissions, data integrity, dependency failure, latency, offline/degraded UX, concurrency/races.28- Define oracles per test: business rule oracle, contract/schema oracle, security oracle, accessibility oracle, performance oracle.2930### Test Portfolio (Modern Equivalent of the Pyramid)3132- Prefer unit + component + contract + integration as the default safety net; keep E2E for thin, critical journeys.33- Add exploratory testing for discovery and usability; convert findings into automated checks when ROI is positive.34- Use the smallest scope that detects the bug class:35 - Bug in business logic: unit/property-based.36 - Bug in service wiring/data: integration/contract.37 - Bug in cross-service compatibility: contract + a small number of integration scenarios.38 - Bug in user journey: E2E (critical path only).3940### Shift-Left Gates (Pre-Merge by Default)4142- Contracts: OpenAPI/AsyncAPI/JSON Schema validation where applicable ([OpenAPI](https://spec.openapis.org/oas/latest.html), [AsyncAPI](https://www.asyncapi.com/docs/reference/specification/v3.0.0), [JSON Schema](https://json-schema.org/)).43- Static checks: lint, typecheck, dependency scanning, secret scanning.44- Fast tests: unit + key integration checks; avoid full E2E as a PR gate unless the product is E2E-only.4546### Coverage Model (Explicitly Separate)4748- Code coverage answers: “What code executed?” (useful as a smoke signal).49- Risk coverage answers: “What user/business risk is reduced?” (the real target).50- REQUIRED: every critical journey has at least one automated check at the lowest effective layer.5152### CI Economics (Budgets and Levers)5354- Budgets [Inference]:55 - PR gate: p50 <= 10 min, p95 <= 20 min.56 - Mainline health: >= 99% green builds per day.57- Levers:58 - Parallelize by layer and shard long-running suites (Playwright supports sharding in CI: [Sharding](https://playwright.dev/docs/test-sharding)).59 - Cache dependencies and test artifacts where your CI supports it.60 - Run “full regression” on schedule (nightly) and “risk-scoped regression” on PRs.6162### Flake Management (SLO + Runbook)6364- Define flake: test fails without product change and passes on rerun.65- Track flake rate as: `rerun_pass / (rerun_pass + rerun_fail)` for a suite.66- SLO examples [Inference]:67 - Suite flake rate <= 1% weekly.68 - Time-to-deflake: p50 <= 2 business days, p95 <= 7 business days.69- REQUIRED: quarantine policy and deflake playbook: `templates/runbooks/template-flaky-test-triage-deflake-runbook.md`.70- For rate-limited endpoints, run serially, reuse tokens, and add backoff for 429s; isolate 429 tests to avoid poisoning other suites.7172### Debugging Ergonomics (Make Failures Cheap)7374- Always capture first failure context: request IDs, trace IDs, build URL, seed, test data IDs.75- Standardize artifacts per layer:76 - Unit/integration: structured logs + minimal repro input.77 - E2E: trace + screenshot/video on failure (Playwright tooling: [Trace Viewer](https://playwright.dev/docs/trace-viewer)).78 - Mobile: `xcresult` bundles + screenshots + device logs.7980### Do / Avoid8182Do:83- Write tests against stable contracts and user-visible behavior.84- Treat flaky tests as P1 reliability work; quarantine only with an owner and expiry.85- Make “how to debug this failure” part of every suite’s definition of done.8687Avoid:88- “Everything E2E” as a default (slow, expensive, low-signal).89- Sleeps/time-based waits (prefer assertions and event-based waits).90- Using coverage % as the primary quality KPI (use risk coverage + defect escape rate).9192## When to Use This Skill9394Invoke when users ask for:9596- Test strategy for a new service or feature97- Unit testing with Jest or Vitest98- Integration testing with databases, APIs, external services99- E2E testing with Playwright or Cypress100- Performance and load testing with k6101- BDD with Cucumber and Gherkin102- API contract testing with Pact103- Visual regression testing104- Test automation CI/CD integration105- Test data management and fixtures106- Security and accessibility testing107- Test coverage analysis and improvement108- Flaky test diagnosis and fixes109- Mobile app testing (iOS/Android)110111---112113## Quick Reference Table114115| Test Type | Framework | Command | When to Use |116|-----------|-----------|---------|-------------|117| Unit Tests | Vitest | `vitest run` | Pure functions, business logic (40-60% of tests) |118| Component Tests | React Testing Library | `vitest --ui` | React components, user interactions (20-30%) |119| Integration Tests | Supertest + Docker | `vitest run integration.test.ts` | API endpoints, database operations (15-25%) |120| E2E Tests | Playwright | `playwright test` | Critical user journeys, cross-browser (5-10%) |121| Performance Tests | k6 | `k6 run load-test.js` | Load testing, stress testing (nightly/pre-release) |122| API Contract Tests | Pact | `pact test` | Microservices, consumer-provider contracts |123| Visual Regression | Percy/Chromatic | `percy snapshot` | UI consistency, design system validation |124| Security Tests | OWASP ZAP | `zap-baseline.py` | Vulnerability scanning (every PR) |125| Accessibility Tests | axe-core | `vitest run a11y.test.ts` | WCAG compliance (every component) |126| Mutation Tests | Stryker | `stryker run` | Test quality validation (weekly) |127128---129130## Decision Tree: Test Strategy131132```text133Need to test: [Feature Type]134 │135 ├─ Pure business logic?136 │ └─ Unit tests (Jest/Vitest) — Fast, isolated, AAA pattern137 │ ├─ Has dependencies? → Mock them138 │ ├─ Complex calculations? → Property-based testing (fast-check)139 │ └─ State machine? → State transition tests140 │141 ├─ UI Component?142 │ ├─ Isolated component?143 │ │ └─ Component tests (React Testing Library)144 │ │ ├─ User interactions → fireEvent/userEvent145 │ │ └─ Accessibility → axe-core integration146 │ │147 │ └─ User journey?148 │ └─ E2E tests (Playwright)149 │ ├─ Critical path → Always test150 │ ├─ Edge cases → Selective E2E151 │ └─ Visual → Percy/Chromatic152 │153 ├─ API Endpoint?154 │ ├─ Single service?155 │ │ └─ Integration tests (Supertest + test DB)156 │ │ ├─ CRUD operations → Test all verbs157 │ │ ├─ Auth/permissions → Test unauthorized paths158 │ │ └─ Error handling → Test error responses159 │ │160 │ └─ Microservices?161 │ └─ Contract tests (Pact) + integration tests162 │ ├─ Consumer defines expectations163 │ └─ Provider verifies contracts164 │165 ├─ Performance-critical?166 │ ├─ Load capacity?167 │ │ └─ k6 load testing (ramp-up, stress, spike)168 │ │169 │ └─ Response time?170 │ └─ k6 performance benchmarks (SLO validation)171 │172 └─ External dependency?173 ├─ Mock it (unit tests) → Use test doubles174 └─ Real implementation (integration) → Docker containers (Testcontainers)175```176177## Decision Tree: Choosing Test Framework178179```text180What are you testing?181 │182 ├─ JavaScript/TypeScript?183 │ ├─ New project? → Vitest (faster, modern)184 │ ├─ Existing Jest project? → Keep Jest185 │ └─ Browser-specific? → Playwright component testing186 │187 ├─ Python?188 │ ├─ General testing? → pytest189 │ ├─ Django? → pytest-django190 │ └─ FastAPI? → pytest + httpx191 │192 ├─ Go?193 │ ├─ Unit tests? → testing package194 │ ├─ Mocking? → gomock or testify195 │ └─ Integration? → testcontainers-go196 │197 ├─ Rust?198 │ ├─ Unit tests? → Built-in #[test]199 │ └─ Property-based? → proptest200 │201 └─ E2E (any language)?202 ├─ Web app? → Playwright (recommended)203 ├─ API only? → k6 or Postman/Newman204 └─ Mobile? → Detox (RN), XCUITest (iOS), Espresso (Android)205```206207## Decision Tree: Flaky Test Diagnosis208209```text210Test is flaky?211 │212 ├─ Timing-related?213 │ ├─ Race condition? → Add proper waits (not sleep)214 │ ├─ Animation? → Disable animations in test mode215 │ └─ Network timeout? → Increase timeout, add retry216 │217 ├─ Data-related?218 │ ├─ Shared state? → Isolate test data219 │ ├─ Random data? → Use seeded random220 │ └─ Order-dependent? → Fix test isolation221 │222 ├─ Environment-related?223 │ ├─ CI-only failures? → Check resource constraints224 │ ├─ Timezone issues? → Use UTC in tests225 │ └─ Locale issues? → Set consistent locale226 │227 └─ External dependency?228 ├─ Third-party API? → Mock it229 └─ Database? → Use test containers230```231232---233234## Test Pyramid235236```text237 /\238 / \239 / E2E \ 5-10% - Critical user journeys240 /--------\ - Slow, expensive, high confidence241 /Integration\ 15-25% - API, database, services242 /--------------\ - Medium speed, good coverage243 / Unit \ 40-60% - Functions, components244 /------------------\ - Fast, cheap, foundation245```246247**Target coverage by layer:**248249| Layer | Coverage | Speed | Confidence |250|-------|----------|-------|------------|251| Unit | 80%+ | ~1000/sec | Low (isolated) |252| Integration | 70%+ | ~10/sec | Medium |253| E2E | Critical paths | ~1/sec | High |254255---256257## Core Capabilities258259### Unit Testing260261- **Frameworks**: Vitest, Jest, pytest, Go testing262- **Patterns**: AAA (Arrange-Act-Assert), Given-When-Then263- **Mocking**: Dependency injection, test doubles264- **Coverage**: Line, branch, function coverage265266### Integration Testing267268- **Database**: Testcontainers, in-memory DBs269- **API**: Supertest, httpx, REST-assured270- **Services**: Docker Compose, localstack271- **Fixtures**: Factory patterns, seeders272273### E2E Testing274275- **Web**: Playwright, Cypress276- **Mobile**: Detox, XCUITest, Espresso277- **API**: k6, Postman/Newman278- **Patterns**: Page Object Model, test locators279280### Performance Testing281282- **Load**: k6, Locust, Gatling283- **Profiling**: Browser DevTools, Lighthouse284- **Monitoring**: Real User Monitoring (RUM)285- **Benchmarks**: Response time, throughput, error rate286287---288289## Common Patterns290291### AAA Pattern (Arrange-Act-Assert)292293```javascript294describe('calculateDiscount', () => {295 it('should apply 10% discount for orders over $100', () => {296 // Arrange297 const order = { total: 150, customerId: 'user-1' };298299 // Act300 const result = calculateDiscount(order);301302 // Assert303 expect(result.discount).toBe(15);304 expect(result.finalTotal).toBe(135);305 });306});307```308309### Page Object Model (E2E)310311```typescript312// pages/login.page.ts313class LoginPage {314 async login(email: string, password: string) {315 await this.page.fill('[data-testid="email"]', email);316 await this.page.fill('[data-testid="password"]', password);317 await this.page.click('[data-testid="submit"]');318 }319320 async expectLoggedIn() {321 await expect(this.page.locator('[data-testid="dashboard"]')).toBeVisible();322 }323}324325// tests/login.spec.ts326test('user can login with valid credentials', async ({ page }) => {327 const loginPage = new LoginPage(page);328 await loginPage.login('user@example.com', 'password');329 await loginPage.expectLoggedIn();330});331```332333### Test Data Factory334335```typescript336// factories/user.factory.ts337export const createUser = (overrides = {}) => ({338 id: faker.string.uuid(),339 email: faker.internet.email(),340 name: faker.person.fullName(),341 createdAt: new Date(),342 ...overrides,343});344345// Usage in tests346const admin = createUser({ role: 'admin' });347const guest = createUser({ role: 'guest', email: 'guest@test.com' });348```349350---351352## CI/CD Integration353354### GitHub Actions Example355356```yaml357name: Test Suite358on: [push, pull_request]359360jobs:361 unit-tests:362 runs-on: ubuntu-latest363 steps:364 - uses: actions/checkout@v4365 - uses: actions/setup-node@v4366 - run: npm ci367 - run: npm run test:unit -- --coverage368 - uses: codecov/codecov-action@v3369370 integration-tests:371 runs-on: ubuntu-latest372 services:373 postgres:374 image: postgres:15375 env:376 POSTGRES_PASSWORD: test377 steps:378 - uses: actions/checkout@v4379 - run: npm ci380 - run: npm run test:integration381382 e2e-tests:383 runs-on: ubuntu-latest384 steps:385 - uses: actions/checkout@v4386 - run: npm ci387 - run: npx playwright install --with-deps388 - run: npm run test:e2e389```390391### Quality Gates392393| Gate | Threshold | Action on Failure |394|------|-----------|-------------------|395| Unit test coverage | 80% | Block merge |396| All tests pass | 100% | Block merge |397| No new critical bugs | 0 | Block merge |398| Performance regression | <10% | Warning |399| Security vulnerabilities | 0 critical | Block deploy |400401---402403## Anti-Patterns to Avoid404405| Anti-Pattern | Problem | Solution |406|--------------|---------|----------|407| Testing implementation | Breaks on refactor | Test behavior, not internals |408| Shared mutable state | Flaky tests | Isolate test data |409| sleep() in tests | Slow, unreliable | Use proper waits/assertions |410| Testing everything E2E | Slow, expensive | Use test pyramid |411| No test data cleanup | Test pollution | Reset state between tests |412| Ignoring flaky tests | False confidence | Fix or quarantine immediately |413| Copy-paste tests | Hard to maintain | Use factories and helpers |414| Testing third-party code | Wasted effort | Trust libraries, test integration |415416---417418## Optional: AI / Automation419420Use AI assistance only as an accelerator for low-risk work; validate outputs with objective checks and evidence.421422Do:423- Generate scaffolding (test file skeletons, fixtures) and then harden manually.424- Use AI to propose edge cases, then select based on your risk model and add explicit oracles.425- Use AI to summarize flaky-test clusters, but base actions on logs/traces and rerun evidence.426427Avoid:428- Accepting generated assertions without validating the oracle (risk: confident nonsense).429- Letting AI “heal” tests by weakening assertions (risk: silent regressions).430431Safety references (optional):432- OWASP Top 10 for LLM Applications: https://owasp.org/www-project-top-10-for-large-language-model-applications/433- NIST AI Risk Management Framework: https://www.nist.gov/itl/ai-risk-management-framework434435---436437## Navigation438439### Resources440441- [resources/operational-playbook.md](resources/operational-playbook.md) — Testing pyramid guidance, BDD/test data patterns, CI gates, and anti-patterns442- [resources/playwright-webapp-testing.md](resources/playwright-webapp-testing.md) — Playwright decision tree, server lifecycle helper, and recon-first scripting pattern443- [resources/comprehensive-testing-guide.md](resources/comprehensive-testing-guide.md) — Full testing methodology reference444- [resources/test-automation-patterns.md](resources/test-automation-patterns.md) — Automation patterns and best practices445- [resources/shift-left-testing.md](resources/shift-left-testing.md) — Early testing strategies446447### Templates448449- [templates/test-strategy-template.md](templates/test-strategy-template.md) — Risk-based test strategy one-pager450- [templates/template-test-case-design.md](templates/template-test-case-design.md) — Test case design (Given/When/Then + oracles)451- [templates/runbooks/template-flaky-test-triage-deflake-runbook.md](templates/runbooks/template-flaky-test-triage-deflake-runbook.md) — Flake triage + deflake runbook452- [templates/automation-pipeline-template.md](templates/automation-pipeline-template.md) — CI/CD automation pattern453- [templates/unit/template-jest-vitest.md](templates/unit/template-jest-vitest.md) — Unit testing454- [templates/integration/template-api-integration.md](templates/integration/template-api-integration.md) — Integration/API testing455- [templates/e2e/template-playwright.md](templates/e2e/template-playwright.md) — Playwright E2E456- [templates/bdd/template-cucumber-gherkin.md](templates/bdd/template-cucumber-gherkin.md) — BDD/Gherkin457- [templates/performance/template-k6-load-testing.md](templates/performance/template-k6-load-testing.md) — k6 performance458- [templates/visual-regression/template-visual-testing.md](templates/visual-regression/template-visual-testing.md) — Visual regression459460### Data461462- [data/sources.json](data/sources.json) — Curated external references463464---465466## Related Skills467468- [../software-backend/SKILL.md](../software-backend/SKILL.md) — API design and backend patterns to test469- [../software-frontend/SKILL.md](../software-frontend/SKILL.md) — Frontend components and UI patterns470- [../ops-devops-platform/SKILL.md](../ops-devops-platform/SKILL.md) — CI/CD pipelines and infrastructure471- [../qa-debugging/SKILL.md](../qa-debugging/SKILL.md) — Debugging failing tests472- [../software-security-appsec/SKILL.md](../software-security-appsec/SKILL.md) — Security testing patterns