Debug
Purpose
Systematic debugging skill. Four phases, always in order. NEVER fix symptoms -- always find and fix the root cause. After 2 failed fix attempts, escalate to the user.
When to Use
- Test failures (expected vs actual mismatch)
- Runtime errors (exceptions, crashes, hangs)
- Regressions (worked before, broken now)
- Unexpected behavior (no error, but wrong result)
Process
Phase 1: Symptom Analysis (WHAT, WHEN, WHERE)
Gather facts before forming hypotheses:
- WHAT: exact error message, stack trace, log output
- WHEN: always? intermittent? after a specific change? under load?
- WHERE: which file, function, line? which test? which environment?
- SINCE WHEN:
git log --oneline -20 -- what changed recently?
Output: symptom report with all facts classified as KNOWN or SUSPECTED.
Phase 2: Reproduction (MINIMAL REPRO)
Make the bug reproducible with the smallest possible case:
- Run the failing test or reproduce the error
- If not reproducible: document exact conditions and STOP (cannot debug what cannot be reproduced)
- Strip to minimal repro: remove unrelated code, simplify inputs, isolate the component
- Confirm: the minimal repro fails consistently
Output: exact command to reproduce the failure.
Phase 3: Root Cause (WHY)
Apply the 5 Whys to move from symptom to cause:
- Why does it fail? -> [immediate cause]
- Why does that happen? -> [deeper cause]
- Why does that happen? -> [root cause]
(Continue until you reach a cause you can fix directly)
Techniques (use as appropriate):
- Binary search: comment out code, add assertions to narrow the location
- Git bisect:
git bisect start HEAD <known-good> to find the breaking commit
- Print tracing: add targeted print/log statements at decision points
- Diff analysis:
git diff <known-good>..HEAD -- <file> to see what changed
- Assumption check: list every assumption the code makes, verify each one
Classification: identify the root cause category:
- Logic error (wrong condition, off-by-one, missing case)
- State corruption (mutation, shared state, race condition)
- Contract violation (caller sends wrong type, missing field)
- Environment (missing dependency, wrong version, config)
- Data (unexpected input, encoding, edge case)
Output: root cause statement (1-2 sentences, specific and testable).
Phase 4: Solution Design (FIX + REGRESSION TEST)
- Design the fix: minimal change that addresses the root cause
- Fix the ROOT CAUSE, not the symptom
- One logical change only
- If the fix is large, the root cause analysis may be wrong -- revisit Phase 3
- Write regression test: a test that fails without the fix and passes with it
- Apply the fix
- Verify: regression test passes AND all existing tests pass
- Check for siblings: does the same bug pattern exist elsewhere? (
grep for similar code)
Escalation Protocol
| Attempt |
Action |
| 1st fix fails |
Try a different approach (not the same thing again) |
| 2nd fix fails |
STOP. Escalate to user with: symptom, repro, root cause analysis, 2 approaches tried |
Never retry the same approach. Never loop silently.
5 Whys Example
Symptom: test_parse_config_handles_empty fails with KeyError
Why 1: config["database"] raises KeyError
Why 2: parse_config returns empty dict when file is empty
Why 3: the YAML parser returns None for empty files, not empty dict
Root cause: missing None -> {} coercion after yaml.safe_load()
Fix: add `config = yaml.safe_load(f) or {}` instead of `config = yaml.safe_load(f)`
Common Mistakes
- Fixing the symptom (add a try/except) instead of the root cause
- Not writing a regression test for the fix
- Guessing without reproducing first
- Changing multiple things at once (change one thing, verify, repeat)
- Retrying the same approach that already failed
- Not checking for sibling bugs (same pattern elsewhere)
Integration
- Called by:
/ai-dispatch (debug tasks), ai-build agent (when tests fail), user directly
- Calls: test runners (to reproduce),
/ai-test (regression test)
- Transitions to:
ai-build (fix implementation), /ai-commit (after verified fix)
$ARGUMENTS
1---2name: ai-debug3description: Use when investigating unexpected behavior, test failures, runtime errors, or regressions. Systematic 4-phase diagnosis: symptom analysis, reproduction, root cause, solution.4---5
6
7# Debug
8
9## Purpose
10
11Systematic debugging skill. Four phases, always in order. NEVER fix symptoms -- always find and fix the root cause. After 2 failed fix attempts, escalate to the user.
12
13## When to Use
14
15- Test failures (expected vs actual mismatch)
16- Runtime errors (exceptions, crashes, hangs)
17- Regressions (worked before, broken now)
18- Unexpected behavior (no error, but wrong result)
19
20## Process
21
22### Phase 1: Symptom Analysis (WHAT, WHEN, WHERE)
23
24Gather facts before forming hypotheses:
25
261. **WHAT**: exact error message, stack trace, log output
272. **WHEN**: always? intermittent? after a specific change? under load?
283. **WHERE**: which file, function, line? which test? which environment?
294. **SINCE WHEN**: `git log --oneline -20` -- what changed recently?
30
31Output: symptom report with all facts classified as KNOWN or SUSPECTED.
32
33### Phase 2: Reproduction (MINIMAL REPRO)
34
35Make the bug reproducible with the smallest possible case:
36
371. Run the failing test or reproduce the error
382. If not reproducible: document exact conditions and STOP (cannot debug what cannot be reproduced)
393. Strip to minimal repro: remove unrelated code, simplify inputs, isolate the component
404. Confirm: the minimal repro fails consistently
41
42Output: exact command to reproduce the failure.
43
44### Phase 3: Root Cause (WHY)
45
46Apply the 5 Whys to move from symptom to cause:
47
481. **Why** does it fail? -> [immediate cause]
492. **Why** does that happen? -> [deeper cause]
503. **Why** does that happen? -> [root cause]
51 (Continue until you reach a cause you can fix directly)
52
53**Techniques** (use as appropriate):
54- **Binary search**: comment out code, add assertions to narrow the location
55- **Git bisect**: `git bisect start HEAD <known-good>` to find the breaking commit
56- **Print tracing**: add targeted print/log statements at decision points
57- **Diff analysis**: `git diff <known-good>..HEAD -- <file>` to see what changed
58- **Assumption check**: list every assumption the code makes, verify each one
59
60**Classification**: identify the root cause category:
61- Logic error (wrong condition, off-by-one, missing case)
62- State corruption (mutation, shared state, race condition)
63- Contract violation (caller sends wrong type, missing field)
64- Environment (missing dependency, wrong version, config)
65- Data (unexpected input, encoding, edge case)
66
67Output: root cause statement (1-2 sentences, specific and testable).
68
69### Phase 4: Solution Design (FIX + REGRESSION TEST)
70
711. **Design the fix**: minimal change that addresses the root cause
72 - Fix the ROOT CAUSE, not the symptom
73 - One logical change only
74 - If the fix is large, the root cause analysis may be wrong -- revisit Phase 3
752. **Write regression test**: a test that fails without the fix and passes with it
763. **Apply the fix**
774. **Verify**: regression test passes AND all existing tests pass
785. **Check for siblings**: does the same bug pattern exist elsewhere? (`grep` for similar code)
79
80## Escalation Protocol
81
82| Attempt | Action |
83|---------|--------|
84| 1st fix fails | Try a different approach (not the same thing again) |
85| 2nd fix fails | STOP. Escalate to user with: symptom, repro, root cause analysis, 2 approaches tried |
86
87Never retry the same approach. Never loop silently.
88
89## 5 Whys Example
90
91```
92Symptom: test_parse_config_handles_empty fails with KeyError
93Why 1: config["database"] raises KeyError
94Why 2: parse_config returns empty dict when file is empty
95Why 3: the YAML parser returns None for empty files, not empty dict
96Root cause: missing None -> {} coercion after yaml.safe_load()
97Fix: add `config = yaml.safe_load(f) or {}` instead of `config = yaml.safe_load(f)`
98```
99
100## Common Mistakes
101
102- Fixing the symptom (add a try/except) instead of the root cause
103- Not writing a regression test for the fix
104- Guessing without reproducing first
105- Changing multiple things at once (change one thing, verify, repeat)
106- Retrying the same approach that already failed
107- Not checking for sibling bugs (same pattern elsewhere)
108
109## Integration
110
111- **Called by**: `/ai-dispatch` (debug tasks), `ai-build agent` (when tests fail), user directly
112- **Calls**: test runners (to reproduce), `/ai-test` (regression test)
113- **Transitions to**: `ai-build` (fix implementation), `/ai-commit` (after verified fix)
114
115$ARGUMENTS