Test Heuristics
Practical test-suite review, design, triage, strategy, and pruning for any
testing layer a developer writes, runs, debugs, or maintains. Provenance and
grounding sources live in skill.json; this file is runtime routing only.
Core principle
A test exists to catch the bugs that ship in this code class — and to be
diagnosable when it fails. A test that passes regardless of bug presence,
breaks on legitimate refactor, or fails uninformatively is failing at its job.
Activation
- Bare invocation (
"use test-heuristics", "test review", "start"):
load references/activity-router.csv, show the activity menu, wait. No
file inspection, no network calls, no writes.
- Concrete invocation with both activity and layer inferable: skip to
step 3 of the workflow.
- Concrete invocation with ambiguous scope: ask one blocker question
identifying activity or layer; do not inspect private systems first.
Workflow
- Pick activity. Load
references/activity-router.csv. Match the
prompt to: triage, review, author, strategize, prune.
Ambiguous → ask once.
- Pick layer. Load
references/activities/<activity>.csv. Match to
one or more layers, or all for cross-layer treatment (review
fan-out, strategize integrative pass). Ambiguous → ask with the
menu.
- Load grounded context. Load the files in the chosen row's
playbook column — most rows reference one playbook; cross-layer
rows list many — plus the listed core_refs. For review/all,
skip: each spawned layer agent loads its own.
- Identify the target persona from
references/core/personas.md.
- Handle purpose (
spec, regression, characterization,
exploration, gate). For review/author/triage/prune, ask
which applies (multiple can); heuristics for that purpose apply
first. For strategize, skip — the strategy template covers all
purposes via its purpose-by-purpose table.
- Spawn sub-agents in parallel (default for
review and prune).
Single-layer: one lens per agent. review/all: one layer per agent;
each runs the three lenses sequentially inside itself.
strategize/all: single integrative pass — no fan-out. See
references/subagent-dispatch.md. Fall back to sequential lenses
only if the host has no delegation primitive.
- Apply the playbook. Use heuristics tagged for this activity. For
review, score 0–10 using references/core/score-rubric.md. For
author, name the good-shaped pattern. For triage, rank hypotheses
before fixes. For strategize, produce a per-layer investment
recommendation. For prune, produce a deletion list with rationale.
Synthesize sub-agent findings here.
- Apply severity from
references/core/severity-rubric.md (0–4)
and tag failure modes from references/core/failure-modes.md on
every finding.
- Emit output per the default template in the activity router row.
Modes
- Guided Draft (default): one optionized question at a time, 3–4 likely
choices plus a freeform path.
- Autopilot: proceed from available context; state assumptions when the
task is clear and low-risk.
- Grill Me: open-ended questions, one at a time, when audience,
constraints, or trade-offs materially change the result.
Output requirements
Every output includes:
- Target persona.
- Layer(s) and purpose(s).
- Activity-specific load-bearing section per the template.
- Failure mode(s) tagged on every finding.
- Verification — how to prove the change worked.
Subagent dispatch
Independent perspectives catch issues a single pass misses. Default
for review and prune. Preferred for author when comparing
approaches. Optional for triage when ranking hypotheses. Skip for
tiny edits, deterministic single-step checks, or tasks requiring
secrets or production access.
Spawn three sub-agents — one per lens: intent reader, refactor
adversary, bug-shape hunter — and run them in parallel. Some
hosts do not auto-dispatch; instruct the main agent explicitly
("spawn three agents," "delegate this in parallel"). Load
references/subagent-dispatch.md for per-lens prompts, the dispatch
template, and the synthesis step. The three lenses each produce a
finding list; the synthesizing pass deduplicates, preserves
disagreements as open questions, and emits the template-shaped
output. Fall back to sequential only when the host has no delegation
primitive — the discipline of switching lens between passes matters
more than the parallelism.
Reference map
references/activity-router.csv — level-1 router (activity).
references/activities/<activity>.csv — level-2 router (layer) per activity.
references/layers/<layer>.md — layer-specific playbooks (one per layer
listed in the activity CSVs).
references/subagent-dispatch.md — three-lens prompts and synthesis.
references/core/severity-rubric.md — 0–4 severity scale.
references/core/score-rubric.md — 0–10 test-quality scale.
references/core/personas.md — target persona list.
references/core/failure-modes.md — six-mode test failure taxonomy.
references/core/oracles.md — SFDIPOT and FEW HICCUPPS oracles for
exploratory work and the bug-shape hunter lens.
templates/*.md — five activity-specific output templates.
evals/activation-cases.md — activation and behavioral cases (positive
and negative).
evals/run-static-checks.sh — structural and schema gates run in CI.
evals/trigger-evals.json — queries for the description-optimization loop.
skill.json — provenance, grounding sources, version, status.
1---2name: test-heuristics3description: Use when reviewing, designing, triaging, or rationalizing test suites — unit, integration, e2e/UI, exploratory, property-based, contract, snapshot, mutation, or performance tests. Trigger for flakiness triage, false-pass risk, brittleness on refactor, suite pruning, test-pyramid/trophy decisions, exploratory charters, and snapshot review. Routes by activity (triage / review / author / strategize / prune) and layer.4license: MIT5---6
7# Test Heuristics
8
9Practical test-suite review, design, triage, strategy, and pruning for any
10testing layer a developer writes, runs, debugs, or maintains. Provenance and
11grounding sources live in `skill.json`; this file is runtime routing only.
12
13## Core principle
14
15**A test exists to catch the bugs that ship in this code class — and to be
16diagnosable when it fails.** A test that passes regardless of bug presence,
17breaks on legitimate refactor, or fails uninformatively is failing at its job.
18
19## Activation
20
21- **Bare invocation** (`"use test-heuristics"`, `"test review"`, `"start"`):
22 load `references/activity-router.csv`, show the activity menu, wait. No
23 file inspection, no network calls, no writes.
24- **Concrete invocation** with both activity and layer inferable: skip to
25 step 3 of the workflow.
26- **Concrete invocation with ambiguous scope**: ask one blocker question
27 identifying activity or layer; do not inspect private systems first.
28
29## Workflow
30
311. **Pick activity.** Load `references/activity-router.csv`. Match the
32 prompt to: `triage`, `review`, `author`, `strategize`, `prune`.
33 Ambiguous → ask once.
342. **Pick layer.** Load `references/activities/<activity>.csv`. Match to
35 one or more layers, or `all` for cross-layer treatment (`review`
36 fan-out, `strategize` integrative pass). Ambiguous → ask with the
37 menu.
383. **Load grounded context.** Load the files in the chosen row's
39 `playbook` column — most rows reference one playbook; cross-layer
40 rows list many — plus the listed `core_refs`. For `review/all`,
41 skip: each spawned layer agent loads its own.
424. **Identify the target persona** from `references/core/personas.md`.
435. **Handle purpose** (`spec`, `regression`, `characterization`,
44 `exploration`, `gate`). For `review`/`author`/`triage`/`prune`, ask
45 which applies (multiple can); heuristics for that purpose apply
46 first. For `strategize`, skip — the strategy template covers all
47 purposes via its purpose-by-purpose table.
486. **Spawn sub-agents in parallel** (default for `review` and `prune`).
49 Single-layer: one lens per agent. `review/all`: one layer per agent;
50 each runs the three lenses sequentially inside itself.
51 `strategize/all`: single integrative pass — no fan-out. See
52 `references/subagent-dispatch.md`. Fall back to sequential lenses
53 only if the host has no delegation primitive.
547. **Apply the playbook.** Use heuristics tagged for this activity. For
55 `review`, score 0–10 using `references/core/score-rubric.md`. For
56 `author`, name the good-shaped pattern. For `triage`, rank hypotheses
57 before fixes. For `strategize`, produce a per-layer investment
58 recommendation. For `prune`, produce a deletion list with rationale.
59 Synthesize sub-agent findings here.
608. **Apply severity** from `references/core/severity-rubric.md` (0–4)
61 and tag failure modes from `references/core/failure-modes.md` on
62 every finding.
639. **Emit output** per the default template in the activity router row.
64
65## Modes
66
67- **Guided Draft (default):** one optionized question at a time, 3–4 likely
68 choices plus a freeform path.
69- **Autopilot:** proceed from available context; state assumptions when the
70 task is clear and low-risk.
71- **Grill Me:** open-ended questions, one at a time, when audience,
72 constraints, or trade-offs materially change the result.
73
74## Output requirements
75
76Every output includes:
77
78- Target persona.
79- Layer(s) and purpose(s).
80- Activity-specific load-bearing section per the template.
81- Failure mode(s) tagged on every finding.
82- Verification — how to prove the change worked.
83
84## Subagent dispatch
85
86Independent perspectives catch issues a single pass misses. **Default
87for `review` and `prune`.** Preferred for `author` when comparing
88approaches. Optional for `triage` when ranking hypotheses. Skip for
89tiny edits, deterministic single-step checks, or tasks requiring
90secrets or production access.
91
92Spawn three sub-agents — one per lens: **intent reader**, **refactor
93adversary**, **bug-shape hunter** — and run them in parallel. Some
94hosts do not auto-dispatch; instruct the main agent explicitly
95("spawn three agents," "delegate this in parallel"). Load
96`references/subagent-dispatch.md` for per-lens prompts, the dispatch
97template, and the synthesis step. The three lenses each produce a
98finding list; the synthesizing pass deduplicates, preserves
99disagreements as open questions, and emits the template-shaped
100output. Fall back to sequential only when the host has no delegation
101primitive — the discipline of switching lens between passes matters
102more than the parallelism.
103
104## Reference map
105
106- `references/activity-router.csv` — level-1 router (activity).
107- `references/activities/<activity>.csv` — level-2 router (layer) per activity.
108- `references/layers/<layer>.md` — layer-specific playbooks (one per layer
109 listed in the activity CSVs).
110- `references/subagent-dispatch.md` — three-lens prompts and synthesis.
111- `references/core/severity-rubric.md` — 0–4 severity scale.
112- `references/core/score-rubric.md` — 0–10 test-quality scale.
113- `references/core/personas.md` — target persona list.
114- `references/core/failure-modes.md` — six-mode test failure taxonomy.
115- `references/core/oracles.md` — SFDIPOT and FEW HICCUPPS oracles for
116 exploratory work and the bug-shape hunter lens.
117- `templates/*.md` — five activity-specific output templates.
118- `evals/activation-cases.md` — activation and behavioral cases (positive
119 and negative).
120- `evals/run-static-checks.sh` — structural and schema gates run in CI.
121- `evals/trigger-evals.json` — queries for the description-optimization loop.
122- `skill.json` — provenance, grounding sources, version, status.