Systematic Debugging
Overview
Random fixes waste time and create new bugs. Quick patches mask underlying issues.
Core principle: ALWAYS find root cause before attempting fixes.
When to Use
Use for ANY technical issue: test failures, unexpected behavior, pipeline errors, build failures, Galaxy workflow errors, notebook exceptions, environment issues.
Use ESPECIALLY when:
- "Just one quick fix" seems obvious
- You've already tried multiple fixes
- Previous fix didn't work
- You don't fully understand the issue
Supporting Files
- root-cause-tracing.md - Trace bugs backward through call chain to find the original trigger. Instrumentation techniques, stack trace analysis.
- defense-in-depth.md - Add validation at multiple layers after finding root cause. Entry point, business logic, environment guards, debug logging.
The Four Phases
Complete each phase before proceeding to the next.
Phase 1: Root Cause Investigation
BEFORE attempting ANY fix:
Read error messages carefully
- Don't skip past errors or warnings — they often contain the solution
- Read stack traces completely
- Note line numbers, file paths, error codes
Reproduce consistently
- Can you trigger it reliably? What are the exact steps?
- If not reproducible, gather more data — don't guess
Check recent changes
- Git diff, recent commits, new dependencies
- Config changes, environmental differences
Gather evidence in multi-component systems
- For pipelines (Galaxy workflow → tool → data), log what enters and exits each component
- Run once with diagnostics to see WHERE it breaks
- Then investigate that specific component
Trace data flow
- Where does the bad value originate? (See root-cause-tracing.md)
- Keep tracing up the call chain until you find the source
- Fix at source, not at symptom
Phase 2: Pattern Analysis
- Find working examples — similar working code in same codebase
- Compare against references — read reference implementation completely, don't skim
- Identify differences — list every difference, however small
- Understand dependencies — settings, config, environment, assumptions
Phase 3: Hypothesis and Testing
- Form single hypothesis — "I think X is the root cause because Y"
- Test minimally — smallest possible change, one variable at a time
- Verify — did it work? If not, form NEW hypothesis. Don't pile fixes on top.
- When you don't know — say so. Don't pretend. Research more.
Phase 4: Implementation
- Create failing test/reproduction — simplest possible, automated if possible
- Implement single fix — address root cause, ONE change, no "while I'm here" improvements
- Verify fix — test passes? No other tests broken? Issue resolved?
- If fix doesn't work:
- Count fixes attempted
- If < 3: return to Phase 1, re-analyze with new information
- If >= 3: STOP — question the architecture (see below)
When 3+ Fixes Fail
Pattern indicating architectural problem:
- Each fix reveals new issues in different places
- Fixes require "massive refactoring"
- Each fix creates new symptoms elsewhere
STOP and discuss with the user before attempting more fixes. This is not a failed hypothesis — it's a wrong approach.
Red Flags — STOP and Return to Phase 1
If you catch yourself thinking:
- "Quick fix for now, investigate later"
- "Just try changing X and see if it works"
- "It's probably X, let me fix that"
- "I don't fully understand but this might work"
- Proposing solutions before tracing data flow
- "One more fix attempt" when already tried 2+
Common Rationalizations
| Excuse |
Reality |
| "Issue is simple, don't need process" |
Simple issues have root causes too |
| "Emergency, no time for process" |
Systematic is FASTER than guess-and-check |
| "Just try this first, then investigate" |
First fix sets the pattern. Do it right. |
| "Multiple fixes at once saves time" |
Can't isolate what worked. Causes new bugs. |
| "I see the problem, let me fix it" |
Seeing symptoms != understanding root cause |
Quick Reference
| Phase |
Key Activities |
Done when |
| 1. Root Cause |
Read errors, reproduce, check changes, trace data |
Understand WHAT and WHY |
| 2. Pattern |
Find working examples, compare |
Differences identified |
| 3. Hypothesis |
Form theory, test minimally |
Confirmed or new hypothesis |
| 4. Implementation |
Create test, fix, verify |
Bug resolved, tests pass |
Attribution
Adapted from obra/superpowers systematic-debugging skill.
1---2name: systematic-debugging3description: Structured 4-phase debugging methodology. Use when encountering any bug, test failure, unexpected behavior, or pipeline error — before proposing fixes. Enforces root cause investigation first.4---5
6# Systematic Debugging
7
8## Overview
9
10Random fixes waste time and create new bugs. Quick patches mask underlying issues.
11
12**Core principle:** ALWAYS find root cause before attempting fixes.
13
14## When to Use
15
16Use for ANY technical issue: test failures, unexpected behavior, pipeline errors, build failures, Galaxy workflow errors, notebook exceptions, environment issues.
17
18**Use ESPECIALLY when:**
19- "Just one quick fix" seems obvious
20- You've already tried multiple fixes
21- Previous fix didn't work
22- You don't fully understand the issue
23
24## Supporting Files
25
26- **[root-cause-tracing.md](root-cause-tracing.md)** - Trace bugs backward through call chain to find the original trigger. Instrumentation techniques, stack trace analysis.
27- **[defense-in-depth.md](defense-in-depth.md)** - Add validation at multiple layers after finding root cause. Entry point, business logic, environment guards, debug logging.
28
29## The Four Phases
30
31Complete each phase before proceeding to the next.
32
33### Phase 1: Root Cause Investigation
34
35**BEFORE attempting ANY fix:**
36
371. **Read error messages carefully**
38 - Don't skip past errors or warnings — they often contain the solution
39 - Read stack traces completely
40 - Note line numbers, file paths, error codes
41
422. **Reproduce consistently**
43 - Can you trigger it reliably? What are the exact steps?
44 - If not reproducible, gather more data — don't guess
45
463. **Check recent changes**
47 - Git diff, recent commits, new dependencies
48 - Config changes, environmental differences
49
504. **Gather evidence in multi-component systems**
51 - For pipelines (Galaxy workflow → tool → data), log what enters and exits each component
52 - Run once with diagnostics to see WHERE it breaks
53 - Then investigate that specific component
54
555. **Trace data flow**
56 - Where does the bad value originate? (See [root-cause-tracing.md](root-cause-tracing.md))
57 - Keep tracing up the call chain until you find the source
58 - Fix at source, not at symptom
59
60### Phase 2: Pattern Analysis
61
621. **Find working examples** — similar working code in same codebase
632. **Compare against references** — read reference implementation completely, don't skim
643. **Identify differences** — list every difference, however small
654. **Understand dependencies** — settings, config, environment, assumptions
66
67### Phase 3: Hypothesis and Testing
68
691. **Form single hypothesis** — "I think X is the root cause because Y"
702. **Test minimally** — smallest possible change, one variable at a time
713. **Verify** — did it work? If not, form NEW hypothesis. Don't pile fixes on top.
724. **When you don't know** — say so. Don't pretend. Research more.
73
74### Phase 4: Implementation
75
761. **Create failing test/reproduction** — simplest possible, automated if possible
772. **Implement single fix** — address root cause, ONE change, no "while I'm here" improvements
783. **Verify fix** — test passes? No other tests broken? Issue resolved?
794. **If fix doesn't work:**
80 - Count fixes attempted
81 - If < 3: return to Phase 1, re-analyze with new information
82 - **If >= 3: STOP — question the architecture** (see below)
83
84### When 3+ Fixes Fail
85
86Pattern indicating architectural problem:
87- Each fix reveals new issues in different places
88- Fixes require "massive refactoring"
89- Each fix creates new symptoms elsewhere
90
91**STOP and discuss with the user before attempting more fixes.** This is not a failed hypothesis — it's a wrong approach.
92
93## Red Flags — STOP and Return to Phase 1
94
95If you catch yourself thinking:
96- "Quick fix for now, investigate later"
97- "Just try changing X and see if it works"
98- "It's probably X, let me fix that"
99- "I don't fully understand but this might work"
100- Proposing solutions before tracing data flow
101- "One more fix attempt" when already tried 2+
102
103## Common Rationalizations
104
105| Excuse | Reality |
106|--------|---------|
107| "Issue is simple, don't need process" | Simple issues have root causes too |
108| "Emergency, no time for process" | Systematic is FASTER than guess-and-check |
109| "Just try this first, then investigate" | First fix sets the pattern. Do it right. |
110| "Multiple fixes at once saves time" | Can't isolate what worked. Causes new bugs. |
111| "I see the problem, let me fix it" | Seeing symptoms != understanding root cause |
112
113## Quick Reference
114
115| Phase | Key Activities | Done when |
116|-------|---------------|-----------|
117| 1. Root Cause | Read errors, reproduce, check changes, trace data | Understand WHAT and WHY |
118| 2. Pattern | Find working examples, compare | Differences identified |
119| 3. Hypothesis | Form theory, test minimally | Confirmed or new hypothesis |
120| 4. Implementation | Create test, fix, verify | Bug resolved, tests pass |
121
122## Attribution
123
124Adapted from [obra/superpowers](https://github.com/obra/superpowers/) systematic-debugging skill.