Hunt
Requires: the sibling
protocolsskill (shared protocol masters); usesskills.config.jsonwhen present. Missing protocols → tell the user to install the full supermodo package.
Docs names come from config. Every
docs/…path below is the DEFAULT. Resolve folder and file names fromskills.config.json→docs.layout(defaults when unset) before reading or writing — a path typed from memory writes a second tree beside the real one. See../protocols/references/docs-convention.md.
Project rules. Read
.supermodo/rules/hunt.mdif present, plus any.supermodo/rules/INDEX.mdrows naminghunt— that file IS this project's hunt process and replaces the defaults below wherever they overlap. Contract:../protocols/references/rules.md. Never in that file, so never switchable off: finders stay blind todocs/, no unverified finding reaches the report, no evidence means no finding, report-only — never a code change.
Systematic bug hunting: automated scans → parallel blind finders → gap sweep → adversarial verification (Claude skeptics × Codex cross-check × docs adjudication) → open questions answered by the user (transport per config) → verified report at the location the repo's docs contract dictates → teardown (ask, then clean up). Report-only — never code changes.
Never escalates on its own. Findings go to tdd --debug to be fixed. The
bug-council skill is explicit-invocation only and is never chained into
from here — not for one finding, and certainly not for a list of them. If a
confirmed finding later resists an actual fix attempt, hunt may SUGGEST
convening the council in one line and wait for the user's yes.
The pipeline is find blind, judge informed. Finders never read docs/ —
a finder that knows "this is documented as intentional" stops reporting real
bugs hiding behind stale docs. Judges (verify phase) read everything and must
cite evidence to kill or resolve a finding. No unverified finding reaches the
report.
Invocation
/hunt <path> # focused: auto-detect relevant layers
/hunt <path> --<layer> # focused: specific layer only
/hunt --<layer> # project-wide: sweep for one layer
/hunt . # full project: all layers
/hunt --diff [base] # only files changed vs base (default: merge-base with main)
Layer flags
| Flag | Finders dispatched | Reference files |
|---|---|---|
--bugs |
semantic, async, error-handling, structure, comparison | semantic.md, async.md, error-handling.md, structure.md |
--data |
data-integrity, type-safety | data-integrity.md, type-safety.md |
--perf |
performance | perf.md |
--security |
security | security.md |
--frontend |
frontend (code-level) | frontend.md |
--browser |
browser (runtime) | browser.md |
No --layer given → auto-detect:
- Utility/backend code →
--bugs --perf - Data pipeline code →
--bugs --data --perf - React/frontend code →
--bugs --frontend --perf --browser - API handlers →
--bugs --security --perf - Full project → all layers
--diff→ detect per changed file, union the results
Core Principles
Evidence Over Opinion
Every finding must include: file path and line number, code snippet or screenshot, why it's a problem (impact), and a suggested fix. No evidence = not a finding. This applies to verification verdicts too: a skeptic kills a finding only with a citation (code line, type guard, or doc quote).
Finders Are Blind
Finder dispatch prompts must NOT include docs/ content, paths to design
docs, or "this is intentional" context. Uncertain findings get
"question": true — the verify phase answers them from docs, the finder
never self-censors.
Kind — defect or improvement
Every finding carries kind: defect when something behaves wrongly
today, improvement when nothing is broken but the code is worse for it
(missing tests, dead exports, structural risk, style). Severity does not imply
kind — the anchors below rank consequence, and both kinds appear at medium
and low. Consumers classify work from this field
(../protocols/references/promotion.md) and may not infer it from the title.
Severity — calibrated anchors
Claiming a severity means naming the concrete consequence at that level. Can't name it → drop one level.
- Critical: data loss, security breach reachable in production, a wrong domain calculation reaching a result users trust (e.g., a correction factor applied twice)
- High: silent failure hiding real errors (
catch → return []), race condition with a plausible trigger, O(n²) on a large-input hot path - Medium: divergent sibling implementations, missing tests on a money path, structural issue that will breed bugs
- Low: dead exports, naming collisions, minor type looseness, style
Cross-Layer Notes
Even in single-layer mode, an obvious critical bug from another layer gets flagged as informational — "out of scope but noted". A perf audit isn't blind to correctness.
Phase 1: Scope & Detect
- Resolve targets: for
--diff,git diff --name-only <base>...HEADplus uncommitted changes; otherwise the given path(s) - Determine code type per the auto-detect table (imports React/Hono JSX → frontend+browser; queries a SQL/columnar store → data; HTTP routes → security; always bugs+perf)
- List files in scope with line counts
- Assign the run-stamp
YYYYMMDDHHmmss— used later for finding ids
Phase 2: Automated Scans
Objective signals, gathered once, included in every finder dispatch.
Run the project's configured commands.test / commands.lint (argv arrays
from skills.config.json — the first use of each in a session needs the
user's approval per the config contract); fall back to the project's own
task runner when no config is present. Capture counts and failures.
# Tests + static checks — use the configured commands, e.g.:
# commands.test → fast suite
# commands.lint → format + lint + type-check
# Pattern greps
grep -rn 'catch.*return \[\]\|catch.*return null\|catch.*return 0' <target>
grep -rn '\blet\b\|: any\|as unknown as' <target>
# Repeated construction (5+ hits of one shape → flag prominently)
grep -rc '\.push({' <target>
# Duplicated computation (same math cluster in 2+ functions → flag)
grep -rn 'Math\.pow\|Math\.min\|Math\.max\|Math\.random\|Math\.floor' <target>
# Usage tracing: for each export, grep callers; flag zero-caller exports
grep -rn '^export' <target>
Adapt the grep patterns and static-check commands to the project's toolchain and language.
Phase 3: Finders — one parallel batch
Dispatch ALL finders for all active layers in ONE parallel batch via the Agent
tool — never inline, never sequential. Each finder reads ONE reference file
and checks only those patterns; dedup happens at merge, so overlap between
finders is cheap and missed coverage is not. While the batch runs, apply the
liveness protocol (../protocols/references/handoff.md, "Liveness"): check
each finder periodically for output growth; a stalled finder is killed and
retried once, a second stall drops it with the gap recorded in the report.
Dispatch table
Subagent types below are the generic defaults. When the project configures an
agent roster (agents.dir in skills.config.json), prefer a matching
specialist from that roster for a lane (e.g. a domain-data reviewer for
data-integrity, a UI reviewer for frontend/browser); with no roster, every
lane uses general-purpose. The "skill to invoke first" column is optional
polish — invoke it only if it's in the session's skill list, otherwise skip
it and rely on the reference file; never guess skill names.
| Finder | subagent_type | model | Skill to invoke first (optional) |
|---|---|---|---|
| semantic | general-purpose |
inherit | — |
| async | general-purpose |
inherit | — |
| error-handling | general-purpose |
inherit | — |
| structure | general-purpose |
sonnet |
a YAGNI/duplication-audit skill, if available, on top of structure.md |
| comparison | general-purpose |
inherit | — (no reference file; lens described below) |
| data-integrity | general-purpose (or a domain-data reviewer from agents.dir) |
inherit | — (read-only; its checklist + data-integrity.md) |
| type-safety | general-purpose |
sonnet |
— |
| perf | general-purpose |
inherit | — |
| security | general-purpose |
inherit | — |
| frontend | general-purpose (or a UI reviewer from agents.dir) |
inherit | a UI/UX audit skill, if available |
| browser | general-purpose (or a UI reviewer from agents.dir) |
inherit | a browser-automation skill (claude-in-chrome or equivalent); an accessibility skill for a11y items |
Comparison finder (part of --bugs): group sibling functions (similar
names, shared config types, same module) and hunt divergence — same formula
implemented differently, one sibling returns Result while another throws, one
respects a config field its twin hardcodes, different assumptions about shared
mutable state. These bugs live between functions; per-file finders miss them.
Feed it the Phase 2 math-operation grep locations.
Dual-model finders: run a Codex finder alongside the Claude one for five layers — semantic, async, data-integrity (where semantic blind spots cost most), plus structure and perf (single-finder lanes get out-sampled when only the loud layers are doubled: their long tail of duplication and constant-factor findings is a lottery draw one finder can't cover). One batched read-only CLI call per layer, in the same parallel wave:
codex exec -s read-only --json -o "$D/codex-<layer>.json" "
Hunt for <layer> bugs in these files: <file list>.
Checklist: <paste the layer's reference file content>.
Report ONLY a JSON array of findings:
[{\"severity\": \"...\", \"kind\": \"defect|improvement\", \"category\": \"...\",
\"file\": \"...\", \"line\": N,
\"title\": \"...\", \"evidence\": \"...\", \"impact\": \"...\", \"fix\": \"...\",
\"question\": false}]
Every finding needs file:line + evidence. No evidence = don't report it."
Codex and Claude findings merge identically in Phase 4. If codex --version
fails, skip the Codex finders and note "single-model hunt" in the report —
never silently degrade.
Finder dispatch prompt (every finder)
Finders don't inherit this conversation. Every dispatch prompt carries:
- The ABSOLUTE path of its ONE reference file, with the instruction to follow ONLY that checklist
- The target file list + Phase 2 automated results
- The skill invocation from the dispatch table (invoke FIRST, then apply the checklist) — if the skill isn't in the session's skill list, skip it and rely on the reference file; never guess skill names
- The finding format below — WITHOUT ids (ids are assigned at merge)
- The blindness rule: do not read
docs/; uncertain →"question": true - Never create or remove git worktrees; work read-only in the run's designated tree — the main tree, or the task worktree the orchestrator passed in (its path is in the dispatch prompt when worktree mode is on)
Fallbacks (user-level skill — environments differ): unregistered
subagent_type → general-purpose with the same prompt. Agents with restricted
tools can't invoke skills — the table already accounts for this.
Finding format
{
"severity": "critical|high|medium|low",
"kind": "defect|improvement",
"category": "Layer > Subcategory",
"file": "path/to/file.ts",
"line": 123,
"title": "Short description",
"evidence": "Code snippet or explanation",
"impact": "What breaks and for whom",
"fix": "Suggested approach",
"question": false
}
Browser finder (if --browser active)
Requires the app's dev server running — check its URL with
curl -s -o /dev/null -w "%{http_code}" <url> first; if not running, tell the
user. Follow references/browser.md.
Phase 4: Merge, Gap Sweep & Dedup
- Collect all finder outputs (Claude + Codex)
- Deduplicate: same file:line → keep the most specific finding; note when both models found it independently (that's corroboration — record it)
- Gap sweep — blind parallel finders converge on the loudest code;
mechanisms in quiet corners, and the polish tail of loud files, go
unclaimed. Dispatch ONE more finder (
general-purpose, inherit) carrying:- the deduped findings as a coverage map —
file:line — titleonly, never docs content (blindness holds: it sees findings, not docs) - the in-scope file list annotated with per-file finding counts
- the instruction: hunt where the map is thin — zero-finding files first, then the quiet corners of claimed files (duplication, hygiene, constant-factor perf that behavioral finders deprioritize). Report only mechanisms absent from the map. Same finding format, same blindness rule. Merge and dedup its output like any finder's.
- the deduped findings as a coverage map —
- Assign ids:
HNT-<run-stamp>-<seq>in severity order, sequential. Ids are final from here — verification annotates them, never renumbers - Sort by severity, then file path
Phase 5: Verify — no finding skips this
Read references/verification.md and follow it. Summary: every finding
(every severity) gets an adversarial Claude skeptic — one per finding for
small lists, subsystem clusters of 6-12 above ~25 — attacking it: not
reproducible / impossible by construction / documented-intentional /
severity inflated. Each skeptic's verdicts are persisted to a file the
moment they return. A batched Codex cross-check attacks the same list
independently. Questions get answered from docs/ with citations. Verdicts
merge mechanically; disputes are kept and shown, never silently resolved.
Open questions are answered live, not shipped. After both verify legs
return and the merge matrix leaves questions OPEN (docs silent on both sides),
ask the user — batched, max 4 questions per call — on the configured transport
(AskUserQuestion only when questions.perSkill.hunt/questions.transport = "tool"; default plain chat) — BEFORE
writing the report. First present each question per the mandatory format in
references/verification.md (which defers to ../protocols/references/questions.md): a
plain-words explanation (max 4 lines, no doc/id/phase references), a one-line
Claude suggestion, a one-line Codex adversarial counter. Each answer: (a) is
recorded via the librarian in the project's decisions/ convention so the
next hunt resolves it from docs, and (b) becomes the citation that re-resolves the
finding (Documented / Confirmed / Refuted per the answer). Only questions the
user explicitly defers ("skip" / "decide later") reach the report's Open
Questions section.
The point: finder output is inflated by construction (blind finders, overlap-tolerant dispatch). Verification is where precision comes from — skipping it ships the inflation to the user.
Phase 6: Report
- Resolve the report location and format from the project's own docs contract
before writing — never assume
docs/audits/. Many repos put audits somewhere specific with mandatory front-matter, a size cap, and a fixed template; writing to the wrong path or shape fails their docs checks.Read the project's docs router (
docs.entryfromskills.config.json, defaultdocs/README.md) first, and follow the convention it points at exactly: the required directory (oftendocs/work/<initiative>/or a standalonedocs/work/<scope>-hunt/), the filename (audit-YYYY-MM-DD-<scope>.mdis common), the front-matter keys, the finding-id scheme (e.g.F-NN), the per-finding disposition, and any size cap. Compact the report at birth — a bloated "full report" draft must never be committed anywhere, and never park hunt output in anarchive/folder (archives are for retired work, not fresh evidence).Ship the machine-readable findings as JSONL shards under
<report-stem>/findings/, allocated with the report so both take the same collision suffix (reports protocol, "Machine-readable findings") — a sharedfindings/folder is overwritten by the next hunt. Split by disposition so a fix agent loads only the actionable set:findings-confirmed-<severity>.jsonl, one per severity, omitting empty ones. Per line: id, F, severity, kind, category, file, line, title, evidence, impact, fix — severity matching the file name,file:linea real repo location (runtime-only observations setlocus: "runtime", never a fake line 0). AnOVERSTATEDfinding is KEPT: it ships at its DOWNGRADED severity withoverstated_from, never as a fourth disposition and never left in prose alone.findings-disputed.jsonl— adds both models' verdicts and evidence.findings-other.jsonl— addsverdict: resolved|documented|doc-drift|refuted+ resolution citation.
Each shard ≤100KB (overflow →
findings-<shard>-2.jsonl).Resolve the live location with the project's own docs tooling when it exists (a
docs:find/docs:checkcommand in config, or grep for prioraudit-*.md/ "Hunt Report"). Mirror the most recent existing hunt report's structure.Hunt ALWAYS writes its report at
.skills/supermodo/hunt/YYYY-MM-DD-<target-name>.mdwith its shards in.skills/supermodo/hunt/YYYY-MM-DD-<target-name>/findings/(reports protocol), then PUBLISHES it —node <skills>/reports/scripts/render.ts --report <that path>, naming the page in the final message (standalone runs only; insideflowthe orchestrator renders the run page) — never directly intodocs/, never a newdocs/audits/folder. When the project's docs contract dictates an audit location insidedocs/, that location is honored THROUGH librarian: hand the finished report over (invoke librarian, or flag it for its next pass) and let it place a copy or pointer at the contract's location. The contract decides WHERE the audit lives; librarian remains the only writer underdocs/.After librarian places anything, run the project's docs validator if it has one (
commands.docsCheck, or the librarian's bundleddocs-check.ts) and fix what it flags.
- If the project runs a findings ledger (a configured or documented
findings-filing tool), file every CONFIRMED critical and high finding to it:
dimension= category,blast_radius= impact,owner_hint= the agent best placed to fix (from the project'sagents.dirroster). Refuted / documented / disputed findings are never filed. Filing must be idempotent — re-running never double-files. Dedupe on the PAIR — this run's identity (its report stem) plus the finding id, never the id alone: ids are unique within a run, not globally (../protocols/references/reports.md), so two same-second hunts mint identical ones and an id-only ledger silently drops the second run's criticals. Projects without such a ledger skip this step. - Findings are not work items. The report and its shards are the record.
Never write to
BACKLOG.mdor create a triad; promotion intodocs/work/is../protocols/references/promotion.md, run by librarian when the user asks. Severity andkindare settled here and consumed there unchanged, so a finding missing either cannot be promoted at all. - Print one line:
Hunt complete: N confirmed (C crit, H high), N refuted, N documented, N disputed, N open questions. Report: <path>Then, when anything was confirmed:To turn findings into work: /supermodo:librarian --promote <path> [finding-id…]
Phase 7: Teardown
After the report ships, inventory what the hunt left running: background
subagents still alive, background Bash shells (Codex exec calls, log tails,
dev servers started for --browser), browser tabs opened by the browser finder.
Then ask the user (transport per questions.transport/perSkill.hunt) — one
question listing exactly what is still up — whether to tear it all down. On yes: stop background tasks
(TaskStop), kill lingering shells, close hunt-opened browser tabs. On no:
leave everything and list what stayed up so nothing lingers silently. Never
tear down without asking — the user may want a finder's transcript or a
running dev server.
Report format
# Hunt Report: <target>
Generated: YYYY-MM-DD HH:MM
## Summary
```supermodo:bars
{"title":"Confirmed findings by severity","unit":"findings","series":[
{"label":"critical","value":N,"state":"bad"},
{"label":"high","value":N,"state":"bad"},
{"label":"medium","value":N,"state":"warn"},
{"label":"low","value":N}]}
| Verdict | Critical | High | Medium | Low | Total |
|---|---|---|---|---|---|
| Confirmed | |||||
| Disputed | |||||
| Documented (intentional) | |||||
| Refuted | |||||
| Open questions: N |
Automated Results
Tests / lint / type-check / dead-export counts. "Single-model hunt" note if Codex was unavailable.
Confirmed
HNT--:
- File:
path:line - Category / Evidence / Impact / Fix
- Verification: <strongest surviving evidence; "corroborated by both models" when true>
Disputed
(verdicts disagreed — both arguments quoted verbatim, user decides)
Documented (intentional)
(finding + the doc citation that resolves it — informational, not filed)
Doc Drift
(code contradicts docs/ — either the code or the doc is wrong; user decides which)
Open Questions
(should be empty — docs-silent questions are asked to the user before the
report is written. Only questions the user explicitly deferred land here.
Answers live in the project's decisions/ convention — the next hunt's verify
phase resolves them automatically instead of re-asking)
Refuted (appendix, collapsed)
(what was checked and killed, with the killing citation — documents coverage)
Browser Findings (if --browser)
Console / network / accessibility / performance / memory / interactions
The bars block leads the report deliberately (`../protocols/references/reports.md`,
"Report bodies"): a reader opening the page sees the severity shape before any
prose, and eleven findings across four severities is one glance instead of a
paragraph. Omit an empty severity rather than drawing a zero bar. When a
confirmed finding's root cause runs through several modules, add a
`supermodo:graph` under it — nodes for the modules, `kind: "cycle"` on the
edge that closes the loop — instead of describing the chain in sentences.
Frontmatter: `skill: hunt`; `status` per the vocabulary in the reports
protocol (`ok` even when the hunt found plenty — `ok` means the hunt did its
job, `failed` means it could not); `summary` naming the counts a reader
decides on ("3 confirmed (1 critical), 2 disputed, 14 refuted"); `task` set to
the triad slug whenever the hunt was scoped to one; `findings` set to this
run's shard directory and `run_stamp` to the `YYYYMMDDHHmmss` stamp its ids
embed — the report's own name has no stamp in it, so without those two a
consumer cannot tell which shards are this report's, and ids are unique within
a run rather than globally (`../protocols/references/reports.md`); and every
deferred open question repeated in `questions` — that is what surfaces it in
the archive.
---
## Flow integration
When invoked by the `flow` orchestrator, hunt is **stage 3** (bug audit,
optional) running as a subagent with its own context:
- **Write the stage report** per `../protocols/references/reports.md` to
`.skills/supermodo/runs/<run-id>/03-hunt.md` — YAML frontmatter with
`skill: hunt`, `status` (`ok` | `failed` | `needs-input` | `skipped`),
`summary` (compressed outcome + the confirmed-finding counts), `drift_notes`
(docs that promise behavior the code doesn't match — DOC-DRIFT verdicts go
here), `decisions`, and `questions` (only when `status: needs-input`). The
full report + JSONL shards ship to the RUN DIRECTORY — shards under
`03-hunt/findings/`, beside the stage report, which the run id already makes
unique, and `findings` / `run_stamp` still go in the frontmatter so a
promotion can prove the association. In flow, nothing is written under
`docs/` at stage 3; if the
project's docs contract wants the audit placed in `docs/`, queue that
placement as a `decisions` entry for the stage-7 librarian pass.
- **Never mutate documentation — in ANY mode.** Standalone: record answers
and doc-worthy decisions in the hunt report (under `.skills/supermodo/`)
and hand them to librarian (invoke it, or flag them for its next pass) —
hunt never writes ADRs or any `docs/` file itself. In flow: emit them as
`decisions` / `drift_notes` in the stage report — the stage-7 librarian
pass persists them. Drift notes only; no doc writes.
- **Read prior stage reports** from the run directory for context — the
stage-2 `work` report scopes the audit to the changed feature.
- **Questions mid-flow** don't call AskUserQuestion: unresolved OPEN questions
go in the report's `questions` frontmatter with `status: needs-input`; the
orchestrator routes them and continues this subagent with the answers.
## Layer Reference Files
One finder reads ONE reference file and checks only its patterns. The full
table — file, patterns, focus — is in `references/layers.md`.