Post-Development Verification
Automated full-stack quality verification after development. Real execution by default -- mock is the last resort. Deliverability is judged by external calls (HTTP requests, CLI invocations, browser interactions), not by internal function calls passing in isolation.
Core Philosophy: Real Execution First
Default realism level is L2: internal services run for real, only uncontrollable external dependencies (third-party APIs, paid services) may be mocked.
Downgrade signals (auto-detected from user intent):
- "快速验证" / "只测逻辑" / "mock就行" -> L0
- Pure function / utility library with no I/O -> L0
- Local environment cannot start service -> L1 (report reason)
- "生产级别" / "全面测试" / "验收测试" -> L3
Realism level definitions:
| Level |
Description |
Mock Ratio |
| L0 |
All dependencies mocked |
100% |
| L1 |
Core service real, databases mocked |
<=50% |
| L2 |
Internal services real, external deps mocked |
<=30% |
| L3 |
All services real (sandbox/test accounts) |
0% |
Workflow
Execute phases sequentially. Each phase produces required artifacts for the next.
Phase 0: Environment Awareness
v
Phase 1: Test Design (with anti-pattern scan)
v
Phase 2: Execution & Evaluation
v
Phase 3: Feedback & Fix Loop (if gates fail)
v
Phase 4: Validation & Output
Safety Requirements
This skill starts services, runs migrations, and makes network calls. Before execution:
- Run in isolation -- use a test/sandbox environment, never production systems
- Use test credentials only -- test accounts, test API keys, environment-variable tokens; never supply production secrets
- Phase 0 first -- review the Environment Report before allowing execution phases (Phase 2+) to understand what will be accessed
- Scoped side effects -- all service starts, DB migrations, test data seeding, and cleanup are limited to the test environment
- User control -- the user can downgrade realism level (e.g., "快速验证" for L0) to reduce scope at any time
Phase 0: Environment Awareness
Gather project context and determine feasibility before designing tests.
0.1 Project Analysis
Identify:
- Language & Framework: from config files (package.json, pyproject.toml, go.mod, etc.)
- Project type: monorepo / microservice / full-stack / library / CLI tool
- Monorepo scope (if applicable): identify which packages/services are affected by the current change
- Test runner: detect existing test framework (pytest, vitest, jest, go test, etc.)
- Coverage tool: detect coverage support (pytest-cov, vitest --coverage, nyc, etc.)
- Build/start commands: from scripts, Docker configs, Makefiles
0.2 Dependency Mapping
For each dependency service, classify:
- Controllable (self-hosted, has test env) -> must run real at L2+
- Uncontrollable (third-party, paid, no sandbox) -> acceptable to mock
0.3 Realism Level Decision
- Check user intent signals -> override default if found
- Check test target type -> pure functions auto-downgrade to L0
- Assess local environment feasibility -> downgrade if services can't start
- Record decision with rationale in the environment report
0.4 Environment Feasibility Check
Verify:
- Required tools installed (runtime, package manager, Docker if needed)
- Dependency services available (databases, caches, message queues)
- Target service can start locally
- Environment consistency: Docker/config matches production; environment variables and config files are complete; dependency service versions align with deployment target
Output: Environment Report -- language, framework, test runner, realism level, service availability, any blockers, consistency gaps.
Phase 1: Test Design
Design test scenarios systematically using the test taxonomy. Do NOT write test code before completing analysis.
1.1 Pre-Analysis (Anti Leap-to-Code)
Before writing any test code, complete:
- Code structure analysis: control flow, function signatures, module dependencies
- Dependency analysis: external services, databases, API contracts involved
- Constraint analysis: business rules, data invariants, input/output contracts
If test code references modules/imports not in the actual codebase -> hallucinated dependency. Remove and use actual project references.
1.2 Scenario Identification
Map each requirement to test scenarios:
- Functional scenarios (one per acceptance criterion minimum)
- Integration scenarios (cross-module, data flow, API contract)
- Error/exception scenarios (all error handling branches)
- End-to-end business flows (complete user journeys that traverse multiple steps/services -- e.g., "register -> verify email -> login -> place order -> pay -> receive confirmation"; "admin creates user -> assigns role -> user logs in with correct permissions"). Each identified critical business flow MUST have at least one E2E test validated through external calls.
Anti-pattern check: If >80% of scenarios are happy path -> Happy Path Obsession detected. Add error, boundary, and exception scenarios until >=30% target error scenarios.
1.3 Apply Test Taxonomy
For each scenario, apply the appropriate taxonomy category. Load detailed guidance from references/test-taxonomy.md when needed.
| Category |
Purpose |
Example |
| MFT (Minimum Functionality) |
Verify each decision branch/leaf node works |
Each code path returns correct result |
| INV (Invariance) |
Same logical request -> same result, different phrasings |
"show data" = "display info" = "list records" |
| DIR (Directional Expectation) |
Vary one input -> predict output direction |
Larger input -> larger output (monotonic) |
Coverage rule: Each of the 3 categories MUST have >=1 test. Taxonomy coverage = 100% is a hard gate.
1.4 Boundary Value Coverage
For every input parameter, identify and cover:
- Numeric: min-1, min, min+1, typical, max-1, max, max+1
- String/Collection: empty, single item, typical, max, exceeds max, special chars, Unicode
- Optional: not provided (default), explicit null/undefined, explicit empty
1.5 Real vs Mock Classification
Classify each scenario based on realism level:
- L2+: Internal service interactions -> must run real
- Any level: Uncontrollable external deps -> acceptable to mock (use realistic stubs)
- Track: test realism ratio = real tests / total tests >= 70% (hard gate at L2)
1.6 Visible/Hidden Test Separation
After designing all scenarios:
- Select 30-40% as hidden tests (prioritize edge cases, adversarial inputs, error scenarios)
- Hidden tests are NOT exposed during fix loop -- reserved for Phase 4 validation
- Record the split in the test plan
1.7 Anti-Pattern Scan
Run the pre-execution checklist. Load detailed guidance from references/anti-patterns.md when needed.
| Anti-Pattern |
Check |
Action |
| Happy Path Obsession |
>80% scenarios are normal flow |
Add error/boundary tests |
| Weak Assertions |
assert(result != null), assert(status == 200) without body |
Replace with specific value checks |
| Leap-to-Code |
Test code written before structure analysis |
Redo analysis first |
| Hallucinated Dependencies |
References non-existent modules/imports |
Replace with actual references |
| Missing Traceability |
Generic test names (test_1, test_func) |
Rename to describe specific behavior |
Rule: Each test name MUST describe the specific behavior being tested (e.g., test_submit_empty_form_returns_422). Each test MUST link to the requirement it validates.
Fix all detected anti-patterns before proceeding to Phase 2.
Phase 1 Output
- Test plan with all scenarios (tagged: MFT/INV/DIR, real/mock, visible/hidden)
- Boundary value coverage matrix
- Anti-pattern scan results (all clear)
Phase 2: Execution & Evaluation
2.1 Environment Preparation
Start services in dependency order with health checks:
for each service in dependency_order:
start service
wait_for_health_check (port/ping/readiness endpoint)
if health_check fails:
report blocker, downgrade realism level
Prepare test data using project's existing seed/migration mechanisms when available.
System boot validation -- before running any tests, verify the system itself is deliverable:
- The target service starts successfully from a clean state (no cached artifacts)
- All health/readiness endpoints return healthy
- Startup logs contain no ERROR-level entries (WARNs are logged for review)
- Required configuration and environment variables are complete
- If any boot validation fails -> this is a delivery blocker, report immediately
2.2 Test Execution
Run tests with coverage enabled. The test suite MUST include an E2E layer validated through external calls (HTTP requests to running services, CLI invocations, or browser interactions -- not internal function imports). Load execution templates from references/real-e2e-templates.md when designing this layer.
- Detect and use the project's native test runner and coverage tool
- Execute all visible tests (hidden tests reserved)
- E2E layer: exercise all identified critical business flows through external interfaces. If the project exposes an HTTP API, at minimum send real HTTP requests to each endpoint. If CLI, invoke real commands. If UI, drive real browser interactions.
- Capture: pass/fail per test, assertion output, error messages, coverage data
2.3 Metrics Collection
Compute all 4 layers of metrics. Load detailed definitions from references/metrics.md when needed.
Design Quality (computed after test design, before execution):
| Metric |
Formula |
Threshold |
| Scenario Coverage |
covered requirements / total requirements |
MUST = 100% |
| Taxonomy Coverage |
categories with >=1 test / 3 |
MUST = 100% |
| Boundary Value Coverage |
covered boundary points / total identified |
SHOULD >= 90% |
| Data Feature Coverage |
covered data dimensions / total identified |
SHOULD >= 85% |
Execution Quality (computed after test run):
| Metric |
Formula |
Threshold |
| Pass Rate |
passed tests / total tests |
SHOULD >= 95% |
| Code Coverage |
statements covered / total statements |
SHOULD >= 80% |
| Assertion Density |
total assertions / total tests |
SHOULD >= 2.0 |
| Weak Assertion Ratio |
weak assertions / total assertions |
SHOULD <= 10% |
| Test Realism Ratio |
real tests / total tests |
MUST >= 70% |
Delivery Quality (computed from test results):
| Metric |
Formula |
Threshold |
| Expectation Match Rate |
fully matching tests / total tests (core: MUST 100%) |
Core: MUST=100%, Overall: SHOULD>=95% |
| Boundary Handling Rate |
passing boundary tests / total boundary tests |
SHOULD >= 90% |
| Regression Safety |
still-passing tests / previously-passing tests |
MUST = 100% |
| Business Flow Coverage |
E2E-verified business flows / total identified business flows |
MUST = 100% |
Iteration Efficiency (computed during fix loop):
| Metric |
Formula |
Threshold |
| Fix Convergence Rate |
newly passing / previous failures |
<20% for 2 rounds -> STOP |
| Fix Introduction Rate |
newly failing / total fix attempts |
>30% -> STOP |
Phase 3: Feedback & Fix Loop
Triggered when hard gate metrics are not met.
3.1 Generate Feedback Report
Structure the report as JSON. Load schema from references/feedback-schema.md when needed.
Key sections:
- round: current iteration number
- project_context: language, framework, test runner, realism level
- metrics: all 4 layers with pass/fail per metric
- gate_result:
"pass" | "fix_and_retry" | "stop_and_report"
- failures_grouped: cluster by root cause, each with affected tests + fix direction
- fix_history: timeline of each round's metrics before/after, fix description, newly passing/failing
- hidden_tests: total count, run status, results (null during fix loop, populated in Phase 4)
- anti_patterns_detected: names of anti-patterns found in current test suite
- next_action: what the Agent should do next
3.2 Fix Prioritization
Address the largest failure cluster first (most affected tests) -- highest probability of improving overall pass rate.
3.3 Convergence Control -- Triple Exit Conditions
| Condition |
Threshold |
Action |
| Max iterations |
5 rounds |
Stop, report current state |
| No convergence |
<20% convergence rate for 2 consecutive rounds |
Stop, suggest fundamental issue |
| Regression |
>30% fix introduction rate in any round |
Stop, suggest wrong fix approach |
3.4 Fix Loop Cycle
while round <= 5 AND not converged_stop AND not regression_stop:
read feedback report -> identify largest failure cluster
apply targeted fix to code/tests
re-run full test suite (or failed-only if convergence is high)
compute new metrics
generate new feedback report
check exit conditions
record iteration in fix history
Phase 3 Output
- Fix history timeline (round, metrics before/after, fix description, newly passing/failing)
- Final feedback report with gate_result
Phase 4: Validation & Output
4.1 Hidden Test Validation
Run ALL hidden tests (not exposed during fix loop):
- If hidden tests pass -> validates that fixes are genuine, not overfitted
- If hidden tests fail -> report failures, do NOT re-enter fix loop
4.2 Quality Report
Generate final quality report:
- Verdict:
"PASS" (all hard gates met) or "FAIL" (hard gates not met) with specific reasons
- All 4 layers of metrics with pass/fail status
- Fix history timeline
- Hidden test validation results
- Anti-pattern scan summary
- Recommendations for improvement
4.3 Reusable Test Scripts
Output test scripts that can run independently in CI/CD:
- Follow project's native test format and directory conventions
- Include run instructions: prerequisites, service startup, how to run, expected output
- Include environment documentation: services needed, config required, ports
Quality Gate Quick Reference
Hard Gates (MUST pass -- blocks delivery)
| Metric |
Threshold |
| Scenario Coverage |
= 100% |
| Taxonomy Coverage |
= 100% |
| Test Realism Ratio |
>= 70% |
| Expectation Match (core) |
= 100% |
| Regression Safety |
= 100% |
| Business Flow Coverage |
= 100% |
Soft Targets (SHOULD pass -- reported, does not block unless configured)
| Metric |
Threshold |
| Boundary Value Coverage |
>= 90% |
| Data Feature Coverage |
>= 85% |
| Pass Rate |
>= 95% |
| Code Coverage |
>= 80% |
| Assertion Density |
>= 2.0 |
| Weak Assertion Ratio |
<= 10% |
| Boundary Handling Rate |
>= 90% |
References
Load these files as needed during the workflow:
references/metrics.md -- Complete definitions for all 14 metrics: calculation formulas, threshold rationale, and what it means when a metric is not met. Load when computing or interpreting metrics.
references/test-taxonomy.md -- Detailed guidance for MFT/INV/DIR test categories with pseudocode examples, plus systematic boundary value analysis methods. Load during Phase 1 test design.
references/feedback-schema.md -- Full JSON Schema for feedback reports with a populated example. Load when generating feedback reports in Phase 3.
references/anti-patterns.md -- Detailed detection methods and fix strategies for all 5 anti-patterns: Happy Path Obsession, Weak Assertions, Leap-to-Code, Hallucinated Dependencies, Missing Traceability. Load during Phase 1 anti-pattern scan or Phase 3 feedback.
references/real-e2e-templates.md -- Environment preparation scripts and real E2E test templates for HTTP API, CLI tools, and Browser automation. Load during Phase 2 environment preparation and test execution.
1---2name: post-dev-verification3description: Post-development full-stack verification skill. Automatically triggered after Agent completes a development task. Executes production-level validation (unit + integration + E2E) with real-execution-first philosophy. Use when: (1) Development task is complete and needs verification, (2) User says "run tests", "verify", "validate", "quality check", (3) User says "交付", "验证", "跑测试", "质量检查", "验收", (4) Before creating PRs or merging code, (5) After implementing features, bug fixes, or refactoring, (6) User asks "does this work?", "can we ship this?", "is this ready?". Covers: test design (MFT/INV/DIR taxonomy), quality metrics (4 layers, 15 metrics), feedback-driven fix loop, anti-pattern detection, visible/hidden test separation, reusable test script generation.4---56# Post-Development Verification78Automated full-stack quality verification after development. **Real execution by default** -- mock is the last resort. **Deliverability is judged by external calls** (HTTP requests, CLI invocations, browser interactions), not by internal function calls passing in isolation.910## Core Philosophy: Real Execution First1112Default realism level is **L2**: internal services run for real, only uncontrollable external dependencies (third-party APIs, paid services) may be mocked.1314**Downgrade signals** (auto-detected from user intent):15- "快速验证" / "只测逻辑" / "mock就行" -> L016- Pure function / utility library with no I/O -> L017- Local environment cannot start service -> L1 (report reason)18- "生产级别" / "全面测试" / "验收测试" -> L31920**Realism level definitions:**21| Level | Description | Mock Ratio |22|-------|-------------|------------|23| L0 | All dependencies mocked | 100% |24| L1 | Core service real, databases mocked | <=50% |25| L2 | Internal services real, external deps mocked | <=30% |26| L3 | All services real (sandbox/test accounts) | 0% |2728## Workflow2930Execute phases sequentially. Each phase produces required artifacts for the next.3132```33Phase 0: Environment Awareness34 v35Phase 1: Test Design (with anti-pattern scan)36 v37Phase 2: Execution & Evaluation38 v39Phase 3: Feedback & Fix Loop (if gates fail)40 v41Phase 4: Validation & Output42```4344---4546## Safety Requirements4748This skill starts services, runs migrations, and makes network calls. Before execution:49501. **Run in isolation** -- use a test/sandbox environment, never production systems512. **Use test credentials only** -- test accounts, test API keys, environment-variable tokens; never supply production secrets523. **Phase 0 first** -- review the Environment Report before allowing execution phases (Phase 2+) to understand what will be accessed534. **Scoped side effects** -- all service starts, DB migrations, test data seeding, and cleanup are limited to the test environment545. **User control** -- the user can downgrade realism level (e.g., "快速验证" for L0) to reduce scope at any time5556---5758## Phase 0: Environment Awareness5960Gather project context and determine feasibility before designing tests.6162### 0.1 Project Analysis6364Identify:65- **Language & Framework**: from config files (package.json, pyproject.toml, go.mod, etc.)66- **Project type**: monorepo / microservice / full-stack / library / CLI tool67- **Monorepo scope** (if applicable): identify which packages/services are affected by the current change68- **Test runner**: detect existing test framework (pytest, vitest, jest, go test, etc.)69- **Coverage tool**: detect coverage support (pytest-cov, vitest --coverage, nyc, etc.)70- **Build/start commands**: from scripts, Docker configs, Makefiles7172### 0.2 Dependency Mapping7374For each dependency service, classify:75- **Controllable** (self-hosted, has test env) -> must run real at L2+76- **Uncontrollable** (third-party, paid, no sandbox) -> acceptable to mock7778### 0.3 Realism Level Decision79801. Check user intent signals -> override default if found812. Check test target type -> pure functions auto-downgrade to L0823. Assess local environment feasibility -> downgrade if services can't start834. Record decision with rationale in the environment report8485### 0.4 Environment Feasibility Check8687Verify:88- Required tools installed (runtime, package manager, Docker if needed)89- Dependency services available (databases, caches, message queues)90- Target service can start locally91- **Environment consistency**: Docker/config matches production; environment variables and config files are complete; dependency service versions align with deployment target9293Output: **Environment Report** -- language, framework, test runner, realism level, service availability, any blockers, consistency gaps.9495---9697## Phase 1: Test Design9899Design test scenarios systematically using the test taxonomy. **Do NOT write test code before completing analysis.**100101### 1.1 Pre-Analysis (Anti Leap-to-Code)102103Before writing any test code, complete:104- **Code structure analysis**: control flow, function signatures, module dependencies105- **Dependency analysis**: external services, databases, API contracts involved106- **Constraint analysis**: business rules, data invariants, input/output contracts107108> If test code references modules/imports not in the actual codebase -> hallucinated dependency. Remove and use actual project references.109110### 1.2 Scenario Identification111112Map each requirement to test scenarios:113- Functional scenarios (one per acceptance criterion minimum)114- Integration scenarios (cross-module, data flow, API contract)115- Error/exception scenarios (all error handling branches)116- **End-to-end business flows** (complete user journeys that traverse multiple steps/services -- e.g., "register -> verify email -> login -> place order -> pay -> receive confirmation"; "admin creates user -> assigns role -> user logs in with correct permissions"). Each identified critical business flow MUST have at least one E2E test validated through external calls.117118**Anti-pattern check**: If >80% of scenarios are happy path -> **Happy Path Obsession detected**. Add error, boundary, and exception scenarios until >=30% target error scenarios.119120### 1.3 Apply Test Taxonomy121122For each scenario, apply the appropriate taxonomy category. Load detailed guidance from `references/test-taxonomy.md` when needed.123124| Category | Purpose | Example |125|----------|---------|---------|126| **MFT** (Minimum Functionality) | Verify each decision branch/leaf node works | Each code path returns correct result |127| **INV** (Invariance) | Same logical request -> same result, different phrasings | "show data" = "display info" = "list records" |128| **DIR** (Directional Expectation) | Vary one input -> predict output direction | Larger input -> larger output (monotonic) |129130**Coverage rule**: Each of the 3 categories MUST have >=1 test. Taxonomy coverage = 100% is a **hard gate**.131132### 1.4 Boundary Value Coverage133134For every input parameter, identify and cover:135- **Numeric**: min-1, min, min+1, typical, max-1, max, max+1136- **String/Collection**: empty, single item, typical, max, exceeds max, special chars, Unicode137- **Optional**: not provided (default), explicit null/undefined, explicit empty138139### 1.5 Real vs Mock Classification140141Classify each scenario based on realism level:142- L2+: Internal service interactions -> **must run real**143- Any level: Uncontrollable external deps -> **acceptable to mock** (use realistic stubs)144- Track: test realism ratio = real tests / total tests >= 70% (hard gate at L2)145146### 1.6 Visible/Hidden Test Separation147148After designing all scenarios:149- Select **30-40%** as **hidden tests** (prioritize edge cases, adversarial inputs, error scenarios)150- Hidden tests are NOT exposed during fix loop -- reserved for Phase 4 validation151- Record the split in the test plan152153### 1.7 Anti-Pattern Scan154155Run the pre-execution checklist. Load detailed guidance from `references/anti-patterns.md` when needed.156157| Anti-Pattern | Check | Action |158|-------------|-------|--------|159| Happy Path Obsession | >80% scenarios are normal flow | Add error/boundary tests |160| Weak Assertions | `assert(result != null)`, `assert(status == 200)` without body | Replace with specific value checks |161| Leap-to-Code | Test code written before structure analysis | Redo analysis first |162| Hallucinated Dependencies | References non-existent modules/imports | Replace with actual references |163| Missing Traceability | Generic test names (`test_1`, `test_func`) | Rename to describe specific behavior |164165**Rule**: Each test name MUST describe the specific behavior being tested (e.g., `test_submit_empty_form_returns_422`). Each test MUST link to the requirement it validates.166167**Fix all detected anti-patterns before proceeding to Phase 2.**168169### Phase 1 Output170171- Test plan with all scenarios (tagged: MFT/INV/DIR, real/mock, visible/hidden)172- Boundary value coverage matrix173- Anti-pattern scan results (all clear)174175---176177## Phase 2: Execution & Evaluation178179### 2.1 Environment Preparation180181Start services in dependency order with health checks:182```183for each service in dependency_order:184 start service185 wait_for_health_check (port/ping/readiness endpoint)186 if health_check fails:187 report blocker, downgrade realism level188```189190Prepare test data using project's existing seed/migration mechanisms when available.191192**System boot validation** -- before running any tests, verify the system itself is deliverable:193- The target service starts successfully from a clean state (no cached artifacts)194- All health/readiness endpoints return healthy195- Startup logs contain no ERROR-level entries (WARNs are logged for review)196- Required configuration and environment variables are complete197- If any boot validation fails -> this is a delivery blocker, report immediately198199### 2.2 Test Execution200201Run tests with coverage enabled. **The test suite MUST include an E2E layer validated through external calls** (HTTP requests to running services, CLI invocations, or browser interactions -- not internal function imports). Load execution templates from `references/real-e2e-templates.md` when designing this layer.202203- Detect and use the project's native test runner and coverage tool204- Execute all visible tests (hidden tests reserved)205- **E2E layer**: exercise all identified critical business flows through external interfaces. If the project exposes an HTTP API, at minimum send real HTTP requests to each endpoint. If CLI, invoke real commands. If UI, drive real browser interactions.206- Capture: pass/fail per test, assertion output, error messages, coverage data207208### 2.3 Metrics Collection209210Compute all 4 layers of metrics. Load detailed definitions from `references/metrics.md` when needed.211212**Design Quality** (computed after test design, before execution):213| Metric | Formula | Threshold |214|--------|---------|-----------|215| Scenario Coverage | covered requirements / total requirements | MUST = 100% |216| Taxonomy Coverage | categories with >=1 test / 3 | MUST = 100% |217| Boundary Value Coverage | covered boundary points / total identified | SHOULD >= 90% |218| Data Feature Coverage | covered data dimensions / total identified | SHOULD >= 85% |219220**Execution Quality** (computed after test run):221| Metric | Formula | Threshold |222|--------|---------|-----------|223| Pass Rate | passed tests / total tests | SHOULD >= 95% |224| Code Coverage | statements covered / total statements | SHOULD >= 80% |225| Assertion Density | total assertions / total tests | SHOULD >= 2.0 |226| Weak Assertion Ratio | weak assertions / total assertions | SHOULD <= 10% |227| Test Realism Ratio | real tests / total tests | MUST >= 70% |228229**Delivery Quality** (computed from test results):230| Metric | Formula | Threshold |231|--------|---------|-----------|232| Expectation Match Rate | fully matching tests / total tests (core: MUST 100%) | Core: MUST=100%, Overall: SHOULD>=95% |233| Boundary Handling Rate | passing boundary tests / total boundary tests | SHOULD >= 90% |234| Regression Safety | still-passing tests / previously-passing tests | MUST = 100% |235| Business Flow Coverage | E2E-verified business flows / total identified business flows | MUST = 100% |236237**Iteration Efficiency** (computed during fix loop):238| Metric | Formula | Threshold |239|--------|---------|-----------|240| Fix Convergence Rate | newly passing / previous failures | <20% for 2 rounds -> STOP |241| Fix Introduction Rate | newly failing / total fix attempts | >30% -> STOP |242243---244245## Phase 3: Feedback & Fix Loop246247Triggered when **hard gate metrics** are not met.248249### 3.1 Generate Feedback Report250251Structure the report as JSON. Load schema from `references/feedback-schema.md` when needed.252253Key sections:254- **round**: current iteration number255- **project_context**: language, framework, test runner, realism level256- **metrics**: all 4 layers with pass/fail per metric257- **gate_result**: `"pass"` | `"fix_and_retry"` | `"stop_and_report"`258- **failures_grouped**: cluster by root cause, each with affected tests + fix direction259- **fix_history**: timeline of each round's metrics before/after, fix description, newly passing/failing260- **hidden_tests**: total count, run status, results (null during fix loop, populated in Phase 4)261- **anti_patterns_detected**: names of anti-patterns found in current test suite262- **next_action**: what the Agent should do next263264### 3.2 Fix Prioritization265266Address the **largest failure cluster first** (most affected tests) -- highest probability of improving overall pass rate.267268### 3.3 Convergence Control -- Triple Exit Conditions269270| Condition | Threshold | Action |271|-----------|-----------|--------|272| Max iterations | 5 rounds | Stop, report current state |273| No convergence | <20% convergence rate for 2 consecutive rounds | Stop, suggest fundamental issue |274| Regression | >30% fix introduction rate in any round | Stop, suggest wrong fix approach |275276### 3.4 Fix Loop Cycle277278```279while round <= 5 AND not converged_stop AND not regression_stop:280 read feedback report -> identify largest failure cluster281 apply targeted fix to code/tests282 re-run full test suite (or failed-only if convergence is high)283 compute new metrics284 generate new feedback report285 check exit conditions286 record iteration in fix history287```288289### Phase 3 Output290291- Fix history timeline (round, metrics before/after, fix description, newly passing/failing)292- Final feedback report with gate_result293294---295296## Phase 4: Validation & Output297298### 4.1 Hidden Test Validation299300Run ALL hidden tests (not exposed during fix loop):301- If hidden tests pass -> validates that fixes are genuine, not overfitted302- If hidden tests fail -> report failures, do NOT re-enter fix loop303304### 4.2 Quality Report305306Generate final quality report:307- **Verdict**: `"PASS"` (all hard gates met) or `"FAIL"` (hard gates not met) with specific reasons308- All 4 layers of metrics with pass/fail status309- Fix history timeline310- Hidden test validation results311- Anti-pattern scan summary312- Recommendations for improvement313314### 4.3 Reusable Test Scripts315316Output test scripts that can run independently in CI/CD:317- Follow project's native test format and directory conventions318- Include run instructions: prerequisites, service startup, how to run, expected output319- Include environment documentation: services needed, config required, ports320321---322323## Quality Gate Quick Reference324325### Hard Gates (MUST pass -- blocks delivery)326327| Metric | Threshold |328|--------|-----------|329| Scenario Coverage | = 100% |330| Taxonomy Coverage | = 100% |331| Test Realism Ratio | >= 70% |332| Expectation Match (core) | = 100% |333| Regression Safety | = 100% |334| Business Flow Coverage | = 100% |335336### Soft Targets (SHOULD pass -- reported, does not block unless configured)337338| Metric | Threshold |339|--------|-----------|340| Boundary Value Coverage | >= 90% |341| Data Feature Coverage | >= 85% |342| Pass Rate | >= 95% |343| Code Coverage | >= 80% |344| Assertion Density | >= 2.0 |345| Weak Assertion Ratio | <= 10% |346| Boundary Handling Rate | >= 90% |347348---349350## References351352Load these files as needed during the workflow:353354- **`references/metrics.md`** -- Complete definitions for all 14 metrics: calculation formulas, threshold rationale, and what it means when a metric is not met. Load when computing or interpreting metrics.355356- **`references/test-taxonomy.md`** -- Detailed guidance for MFT/INV/DIR test categories with pseudocode examples, plus systematic boundary value analysis methods. Load during Phase 1 test design.357358- **`references/feedback-schema.md`** -- Full JSON Schema for feedback reports with a populated example. Load when generating feedback reports in Phase 3.359360- **`references/anti-patterns.md`** -- Detailed detection methods and fix strategies for all 5 anti-patterns: Happy Path Obsession, Weak Assertions, Leap-to-Code, Hallucinated Dependencies, Missing Traceability. Load during Phase 1 anti-pattern scan or Phase 3 feedback.361362- **`references/real-e2e-templates.md`** -- Environment preparation scripts and real E2E test templates for HTTP API, CLI tools, and Browser automation. Load during Phase 2 environment preparation and test execution.