/audit Skill — Systematic Code Review
[CONFIG]
- Model: Opus for all code review. Sonnet for orchestration plumbing only.
- Unit of review: directory (subsystem), not individual function.
- Context slice: function source + 1-hop callers + 1-hop callees + checklist metadata.
- Checklist item fields:
name, kind ("function"/"global"/"macro"/"class"), line_start, line_end, signature, checked_by, metadata (visibility, params, return_type, attributes). The field is kind, not type. Source: core/inventory/extractors.CodeItem.
- Findings format: standard
findings.json (same as /scan, fed to /validate unchanged).
- Annotations: markdown per source file, structured metadata in HTML comments.
- Prerequisite:
/understand --map must have run first. If context-map.json is missing from the output directory (or project siblings), auto-run it before starting the review loop.
- Scoping:
--scope <dir> restricts gap selection to a subdirectory (e.g. ipc/, net/ipv4/). All annotations and coverage records still write to the project-level output dir, so successive scoped runs accumulate into one audit trail.
[EXEC] Execution Rules
- No env-var prefixes on commands. NEVER write
OUTPUT_DIR=... libexec/raptor-audit ... or VAR=val command. It breaks permission patterns. Capture OUTPUT_DIR from raptor-run-lifecycle start output and pass it via --out flags on every subsequent command.
- Read code before reasoning. Never describe or hypothesize about code you haven't read with the Read tool.
- LLM generates hypotheses; tools validate. Never directly classify code as vulnerable — that gets 37% accuracy. Form a hypothesis, generate a mechanical test, run it, evaluate the result.
- Pure self-critique is prohibited. All iteration MUST include tool feedback. "Review again" without running a tool is forbidden. Generate a new Semgrep rule, CodeQL query, or SMT check instead. Run all tool invocations through
libexec/raptor-audit sweep (or the Python API packages.semgrep.runner / packages.coccinelle.runner) — never call semgrep or spatch directly via Bash. This ensures results are logged to the audit trail automatically.
- Tool evidence is the verdict. If a tool confirms a hypothesis, the finding includes the tool's output as proof. If a tool refutes it, the hypothesis is discarded — no "but I still think..."
- Annotations record what was tested, not opinions. Each annotation lists the hypotheses tested, the tools run, and their results. A reviewer can re-run the generated rules to verify.
- One status per function. Call
libexec/raptor-audit record exactly once per reviewed function. Status is clean (no issues), dormant (real bug but currently unreachable/dead code), suspicious (concern but not confirmed), finding (tool-confirmed reachable vulnerability), or error (review blocked). Use dormant — not clean — when a function has a genuine bug that is unreachable today (dead code, no callers, commented out). A dormant bug becomes a finding when reachability changes.
- Findings need concrete evidence. A finding must cite: the vulnerable code (file:line), what assumption is violated, and the tool output that confirms it. No "this could be dangerous if..."
- Checker synthesis for patterns. When a confirmed finding suggests a repeatable pattern (e.g., "unchecked return used as index"), generate a codebase-wide Semgrep or Coccinelle rule and run it. One hypothesis → sweep the whole codebase.
- Run lifecycle. Start with
raptor-run-lifecycle start, end with raptor-run-lifecycle complete. On failure: raptor-run-lifecycle fail.
- Do NOT narrate gate compliance. Only show substantive work — the hypothesis, the tool, the result. No "I am now following EXEC rule 2..."
[GATES] Must-Pass Gates
- G1 [HYPOTHESIS-FIRST]: Every suspicion MUST be framed as a testable hypothesis before any finding is emitted. "X looks dangerous" is not a hypothesis. "If input Y reaches sink Z without check W, CWE-N applies" is.
- G2 [TOOL-GROUNDED]: Every finding MUST have at least one mechanical validation (Semgrep match, CodeQL path, Coccinelle hit, SMT sat result, or compilation test). Ungrounded findings are annotation-only (
suspicious), never finding.
- G3 [NO-SELF-CRITIQUE-LOOP]: Iteration without tool feedback is prohibited. If re-reviewing, generate a NEW tool invocation.
- G4 [EVIDENCE-IN-ANNOTATION]: The annotation body MUST include tool names and results, not just prose.
- G5 [READ-FIRST]: Code must be read with the Read tool before any hypothesis is formed about it.
- G6 [ASSUMPTION-TRUST]: For every function, identify what it trusts (inputs, return values, global state, caller guarantees) and ask what happens when each trust is violated.
- G7 [REACHABILITY]: Findings must be reachable. The
orchestrator chokepoint mechanically enforces this:
- Two-signal hard gate: zero static callers AND binary oracle
absent (full DWARF) → orchestrator refuses finding, forces dormant. The compiler deleted the function — the bug is real but unexploitable.
- Soft gate: zero static callers, no binary oracle data →
orchestrator requires --reach-via explaining how the function is reachable (callback, HTTP route, exported API, cross-language binding, dynamic dispatch). This prevents honeyslop (planted dead code with obvious bugs) from inflating finding counts.
- Entry points (from context-map) and functions with static callers bypass this gate automatically.
[STYLE]
- Status values in JSON: snake_case (
clean, dormant, suspicious, finding, error)
- Status in human output: Title Case (
Clean, Suspicious, Finding, Error)
- No red/green indicators (perspective-dependent)
- Annotations are markdown prose — no JSON in annotation bodies
- Findings in
findings.json use standard RAPTOR schema
[STRATEGIES]
Review strategies are selected per-function based on file paths, parameter types, and return types. Multiple strategies can apply to the same function.
| Strategy |
When |
Key questions |
| General |
Default for all code |
What does it trust? What happens when assumptions are violated? What's surprising? |
| Input handling |
Parsers, protocol handlers, decoders |
Input format/size assumptions? Length fields trusted before use? |
| Concurrency |
Lock APIs, mutexes, atomics |
Lock windows? Concurrent interleavings? Memory barriers? |
| Memory |
Allocators, refcounts, pools |
Ownership model? Symmetric refcounting? Cleanup on failure? |
| Auth/privilege |
Permission checks, ACLs, credentials |
Check bypass? Error path security? Unvalidated transitions? |
| Crypto |
Crypto APIs, key material, RNG |
Correct algorithm usage? Timing side channels? Key lifecycle? |
| Aliasing |
splice, zero-copy, scatterlist, sk_buff |
Alias assumptions? Who owns backing pages? Can another subsystem write through the alias? |
Strategy details and CVE exemplars are in .claude/skills/audit/review.md.
[CRITIQUE]
After reviewing a batch of functions, run the tool-grounded critique pass:
libexec/raptor-audit critique --out "$OUTPUT_DIR"
This mechanically identifies:
- Low sweep coverage: functions reviewed with <2 tool checks. Generate additional Semgrep/SMT/CodeQL tests for these.
- Mode 2 gaps: confirmed findings without codebase-wide rules. Write a generalized checker and run it via
rules save + rules run.
- Suspicious with untried tools: suspicious functions where not all tool types were attempted. Try the suggested tools.
The critique pass produces action items, not prose. Each item should result in a new sweep call, never just "review again."
[CONTEXT]
The context slice (assembled by raptor-audit context) includes:
- Function source lines with line numbers
- 1-hop callers and callees from the call graph
- Checklist metadata (signature, visibility, parameters, return type, attributes)
- Reachable sinks from
context-map.json
- Pre-computed trust surface questions (per-parameter, per-callee)
- Strategy exemplars: per-strategy CVE worked examples showing the reasoning chain that found the bug
- Flow traces: cross-function data flow paths from
/understand --trace that pass through this function
- Prior labeled attempts (if available) from the shared corpus
- Existing annotations (for re-review context)
The context is strategy-aware: a function taking (char *buf, size_t len) gets the input handling strategy exemplar (CVE-2023-0179) alongside the general exemplar. Functions in aliasing-relevant code get CVE-2026-31431 (CopyFail).
[REMIND]
- The LLM generates hypotheses and tools; deterministic analysis confirms or refutes.
- 37.6% MORE critical vulnerabilities after 5 iterations of self-refinement without tool feedback (IEEE-ISTAS 2025). Tool grounding is mandatory, not optional.
- A confirmed pattern should always generate a codebase-wide sweep rule (Mode 2 / KNighter pattern).
- Coverage records accumulate across runs. The gap list shrinks each time.
- Annotations persist in the project directory across runs. They're the audit trail.
- After each batch, run
critique to find gaps before moving on.
Source: gadievron/raptor → .claude/skills/audit/SKILL.md
1---2name: audit3description: Hypothesis-driven, tool-grounded security review of coverage gaps4---5
6
7# /audit Skill — Systematic Code Review
8
9## [CONFIG]
10
11- Model: Opus for all code review. Sonnet for orchestration plumbing only.
12- Unit of review: directory (subsystem), not individual function.
13- Context slice: function source + 1-hop callers + 1-hop callees + checklist metadata.
14- Checklist item fields: `name`, `kind` (`"function"`/`"global"`/`"macro"`/`"class"`), `line_start`, `line_end`, `signature`, `checked_by`, `metadata` (`visibility`, `params`, `return_type`, `attributes`). The field is `kind`, not `type`. Source: `core/inventory/extractors.CodeItem`.
15- Findings format: standard `findings.json` (same as `/scan`, fed to `/validate` unchanged).
16- Annotations: markdown per source file, structured metadata in HTML comments.
17- Prerequisite: `/understand --map` must have run first. If `context-map.json` is missing from the output directory (or project siblings), auto-run it before starting the review loop.
18- Scoping: `--scope <dir>` restricts gap selection to a subdirectory (e.g. `ipc/`, `net/ipv4/`). All annotations and coverage records still write to the project-level output dir, so successive scoped runs accumulate into one audit trail.
19
20## [EXEC] Execution Rules
21
220. **No env-var prefixes on commands.** NEVER write `OUTPUT_DIR=... libexec/raptor-audit ...` or `VAR=val command`. It breaks permission patterns. Capture `OUTPUT_DIR` from `raptor-run-lifecycle start` output and pass it via `--out` flags on every subsequent command.
231. **Read code before reasoning.** Never describe or hypothesize about code you haven't read with the Read tool.
242. **LLM generates hypotheses; tools validate.** Never directly classify code as vulnerable — that gets 37% accuracy. Form a hypothesis, generate a mechanical test, run it, evaluate the result.
253. **Pure self-critique is prohibited.** All iteration MUST include tool feedback. "Review again" without running a tool is forbidden. Generate a new Semgrep rule, CodeQL query, or SMT check instead. Run all tool invocations through `libexec/raptor-audit sweep` (or the Python API `packages.semgrep.runner` / `packages.coccinelle.runner`) — never call `semgrep` or `spatch` directly via Bash. This ensures results are logged to the audit trail automatically.
264. **Tool evidence is the verdict.** If a tool confirms a hypothesis, the finding includes the tool's output as proof. If a tool refutes it, the hypothesis is discarded — no "but I still think..."
275. **Annotations record what was tested, not opinions.** Each annotation lists the hypotheses tested, the tools run, and their results. A reviewer can re-run the generated rules to verify.
286. **One status per function.** Call `libexec/raptor-audit record` exactly once per reviewed function. Status is `clean` (no issues), `dormant` (real bug but currently unreachable/dead code), `suspicious` (concern but not confirmed), `finding` (tool-confirmed reachable vulnerability), or `error` (review blocked). Use `dormant` — not `clean` — when a function has a genuine bug that is unreachable today (dead code, no callers, commented out). A dormant bug becomes a finding when reachability changes.
297. **Findings need concrete evidence.** A finding must cite: the vulnerable code (file:line), what assumption is violated, and the tool output that confirms it. No "this could be dangerous if..."
308. **Checker synthesis for patterns.** When a confirmed finding suggests a repeatable pattern (e.g., "unchecked return used as index"), generate a codebase-wide Semgrep or Coccinelle rule and run it. One hypothesis → sweep the whole codebase.
319. **Run lifecycle.** Start with `raptor-run-lifecycle start`, end with `raptor-run-lifecycle complete`. On failure: `raptor-run-lifecycle fail`.
3210. **Do NOT narrate gate compliance.** Only show substantive work — the hypothesis, the tool, the result. No "I am now following EXEC rule 2..."
33
34## [GATES] Must-Pass Gates
35
36- **G1 [HYPOTHESIS-FIRST]**: Every suspicion MUST be framed as a testable hypothesis before any finding is emitted. "X looks dangerous" is not a hypothesis. "If input Y reaches sink Z without check W, CWE-N applies" is.
37- **G2 [TOOL-GROUNDED]**: Every finding MUST have at least one mechanical validation (Semgrep match, CodeQL path, Coccinelle hit, SMT sat result, or compilation test). Ungrounded findings are annotation-only (`suspicious`), never `finding`.
38- **G3 [NO-SELF-CRITIQUE-LOOP]**: Iteration without tool feedback is prohibited. If re-reviewing, generate a NEW tool invocation.
39- **G4 [EVIDENCE-IN-ANNOTATION]**: The annotation body MUST include tool names and results, not just prose.
40- **G5 [READ-FIRST]**: Code must be read with the Read tool before any hypothesis is formed about it.
41- **G6 [ASSUMPTION-TRUST]**: For every function, identify what it trusts (inputs, return values, global state, caller guarantees) and ask what happens when each trust is violated.
42- **G7 [REACHABILITY]**: Findings must be reachable. The `orchestrator` chokepoint mechanically enforces this:
43 - **Two-signal hard gate**: zero static callers AND binary oracle `absent` (full DWARF) → `orchestrator` refuses `finding`, forces `dormant`. The compiler deleted the function — the bug is real but unexploitable.
44 - **Soft gate**: zero static callers, no binary oracle data → `orchestrator` requires `--reach-via` explaining how the function is reachable (callback, HTTP route, exported API, cross-language binding, dynamic dispatch). This prevents honeyslop (planted dead code with obvious bugs) from inflating finding counts.
45 - Entry points (from context-map) and functions with static callers bypass this gate automatically.
46
47## [STYLE]
48
49- Status values in JSON: snake_case (`clean`, `dormant`, `suspicious`, `finding`, `error`)
50- Status in human output: Title Case (`Clean`, `Suspicious`, `Finding`, `Error`)
51- No red/green indicators (perspective-dependent)
52- Annotations are markdown prose — no JSON in annotation bodies
53- Findings in `findings.json` use standard RAPTOR schema
54
55## [STRATEGIES]
56
57Review strategies are selected per-function based on file paths, parameter types, and return types. Multiple strategies can apply to the same function.
58
59| Strategy | When | Key questions |
60|----------|------|---------------|
61| **General** | Default for all code | What does it trust? What happens when assumptions are violated? What's surprising? |
62| **Input handling** | Parsers, protocol handlers, decoders | Input format/size assumptions? Length fields trusted before use? |
63| **Concurrency** | Lock APIs, mutexes, atomics | Lock windows? Concurrent interleavings? Memory barriers? |
64| **Memory** | Allocators, refcounts, pools | Ownership model? Symmetric refcounting? Cleanup on failure? |
65| **Auth/privilege** | Permission checks, ACLs, credentials | Check bypass? Error path security? Unvalidated transitions? |
66| **Crypto** | Crypto APIs, key material, RNG | Correct algorithm usage? Timing side channels? Key lifecycle? |
67| **Aliasing** | splice, zero-copy, scatterlist, sk_buff | Alias assumptions? Who owns backing pages? Can another subsystem write through the alias? |
68
69Strategy details and CVE exemplars are in `.claude/skills/audit/review.md`.
70
71## [CRITIQUE]
72
73After reviewing a batch of functions, run the tool-grounded critique pass:
74
75```bash
76libexec/raptor-audit critique --out "$OUTPUT_DIR"
77```
78
79This mechanically identifies:
80- **Low sweep coverage**: functions reviewed with <2 tool checks. Generate additional Semgrep/SMT/CodeQL tests for these.
81- **Mode 2 gaps**: confirmed findings without codebase-wide rules. Write a generalized checker and run it via `rules save` + `rules run`.
82- **Suspicious with untried tools**: suspicious functions where not all tool types were attempted. Try the suggested tools.
83
84The critique pass produces action items, not prose. Each item should result in a new `sweep` call, never just "review again."
85
86## [CONTEXT]
87
88The context slice (assembled by `raptor-audit context`) includes:
89- Function source lines with line numbers
90- 1-hop callers and callees from the call graph
91- Checklist metadata (signature, visibility, parameters, return type, attributes)
92- Reachable sinks from `context-map.json`
93- Pre-computed trust surface questions (per-parameter, per-callee)
94- **Strategy exemplars**: per-strategy CVE worked examples showing the reasoning chain that found the bug
95- **Flow traces**: cross-function data flow paths from `/understand --trace` that pass through this function
96- Prior labeled attempts (if available) from the shared corpus
97- Existing annotations (for re-review context)
98
99The context is strategy-aware: a function taking `(char *buf, size_t len)` gets the input handling strategy exemplar (CVE-2023-0179) alongside the general exemplar. Functions in aliasing-relevant code get CVE-2026-31431 (CopyFail).
100
101## [REMIND]
102
103- The LLM generates hypotheses and tools; deterministic analysis confirms or refutes.
104- 37.6% MORE critical vulnerabilities after 5 iterations of self-refinement without tool feedback (IEEE-ISTAS 2025). Tool grounding is mandatory, not optional.
105- A confirmed pattern should always generate a codebase-wide sweep rule (Mode 2 / KNighter pattern).
106- Coverage records accumulate across runs. The gap list shrinks each time.
107- Annotations persist in the project directory across runs. They're the audit trail.
108- After each batch, run `critique` to find gaps before moving on.
109
110---
111
112**Source:** [`gadievron/raptor`](https://github.com/gadievron/raptor) → `.claude/skills/audit/SKILL.md`