Skill: agent-ops-debugging
Systematic debugging approaches for isolating and fixing software defects
Purpose
Systematic problem isolation, root cause analysis, and defect resolution. Use when something isn't working and the cause is unclear.
Core Principles
1. Understand Before Acting
- Reproduce the issue: Can you consistently trigger the problem?
- Define expected vs actual: What should happen vs what is happening?
- Gather context: When does this occur? Under what conditions?
- Recent changes: What changed before this appeared?
2. Isolate the Problem
- Binary search: Comment out half the code, test, repeat
- Minimize reproduction: Create minimal test case
- Control variables: Change one thing at a time
- Eliminate noise: Remove unrelated factors
3. Form Hypotheses
- State your assumption: "I believe X is causing Y because..."
- Make predictions: "If my hypothesis is true, then Z should happen"
- Test predictions: Verify or refute each hypothesis
- Iterate: Refine hypothesis based on evidence
4. Fix and Verify
- Address root cause: Not just symptoms
- Minimize changes: Smallest fix that resolves the issue
- Add tests: Prevent regression
- Verify fix: Test the specific scenario and related scenarios
Systematic Debugging Process
Phase 1: Problem Definition
- Describe the bug in one sentence
- List reproduction steps (minimal set)
- Specify expected behavior
- Capture actual behavior (screenshots, logs, error messages)
- Identify scope: How widespread is this?
Phase 2: Information Gathering
- Check logs: Application logs, system logs, crash reports
- Inspect state: Database records, cache contents, file system
- Review code: Recent changes, related code paths
- Compare environments: Dev vs staging vs production differences
- Monitor resources: CPU, memory, disk, network during issue
Phase 3: Hypothesis Formation
Common failure patterns:
| Pattern |
Symptoms |
Where to Look |
| Timing issues |
Intermittent, "works sometimes" |
Race conditions, deadlocks, timeouts |
| State corruption |
Wrong data, unexpected mutations |
Shared state, caches, global variables |
| Resource exhaustion |
Slows down, eventually fails |
Memory leaks, connection pools |
| Configuration |
Works elsewhere, fails here |
Environment variables, settings files |
| Dependencies |
Broke after update |
Library versions, API changes |
| Assumption violations |
Edge case failures |
Code assumes something that isn't true |
Phase 4: Hypothesis Testing
- Add logging: Instrument code to verify assumptions
- Use debugger: Set breakpoints, inspect variables, step through
- Write tests: Create failing test that reproduces bug
- Simplify: Remove complexity while preserving failure
- Verify: Confirm hypothesis explains all symptoms
Phase 5: Resolution
- Implement fix: Address root cause, not symptoms
- Add regression test: Ensure bug doesn't return
- Review similar code: Check for same issue elsewhere
- Document: Add comments, update docs if behavior changed
- Verify: Test fix works and doesn't break other things
Debugging by Symptom
"It Works on My Machine"
| Check |
Action |
| Environment differences |
Python versions, OS, dependencies |
| Uncommitted config |
Local settings, .env files |
| Race conditions |
Timing-dependent issues |
| Data differences |
Test with production data subset |
| Resource constraints |
Production may have different limits |
Intermittent Failures
| Check |
Action |
| Shared state |
Global variables, singletons, caches |
| Timing |
Race conditions, timeouts, async issues |
| Randomness |
Random seeds, shuffling, sampling |
| Resource cleanup |
Are resources properly released? |
| External dependencies |
Network calls, third-party services |
Performance Degradation
| Check |
Action |
| Profile first |
Measure before optimizing |
| O(n²) |
Nested loops, repeated work |
| I/O |
Database queries, file reads, network |
| Memory |
Leaks, large objects, excessive allocations |
| Caching |
Repeated expensive operations |
Memory Leaks
| Check |
Action |
| Profile memory |
Track allocations over time |
| Circular references |
GC can't collect cycles |
| Event listeners |
Detached handlers keeping objects alive |
| Caches |
Growing without bounds |
| Static collections |
Accumulating entries |
Deadlocks
| Check |
Action |
| Lock order |
Identify held locks, acquisition order |
| Cycles |
A waits for B, B waits for A |
| Timeouts |
Are operations waiting indefinitely? |
| Hold-and-wait |
Holding one lock while waiting for another |
Tool-Specific Guidance
Print/Log Statements
# Strategic placement with unique markers
print(f"[DEBUG-001] user_id={user_id}, state={state}")
# Include enough context
logger.debug(f"Processing item {i}/{total}: {item.id}")
# Remove after debugging!
Debugger
- Set breakpoints at suspicious locations, not everywhere
- Watch expressions for specific variables
- Check call stack to understand how you got here
- Step carefully through suspicious code
Tests for Debugging
- Write failing test that captures bug reproduction
- Use
git bisect to find when bug was introduced
- Mock external dependencies to isolate
- Property-based testing finds edge cases
Anti-Patterns to Avoid
| Anti-Pattern |
Problem |
Better Approach |
| Shotgun debugging |
Random changes hoping something works |
Form hypothesis, test, refine |
| Symptom treatment |
Adding error handling to hide failures |
Fix underlying cause |
| Assuming |
"This variable can't be null" |
Add assertion to verify |
| Overcomplicating |
Complex debugging infrastructure |
Start simple, add tools as needed |
| Ignoring evidence |
Dismissing data that doesn't fit |
Revise hypothesis to explain all |
Debugging Checklist
Before declaring "debugged":
When to Escalate
Consider asking for help if:
- After 2 hours without progress
- Issue is in unfamiliar technology stack
- Problem involves complex distributed systems
- Security implications
- Production outage
- Going in circles (revisiting same hypotheses)
Recording Debug Sessions
Track in .agent/focus.md:
## Debugging: [Issue Description]
**Symptom**: [What's happening]
**Expected**: [What should happen]
**Reproduction**: [Steps to trigger]
### Hypotheses
1. [Hypothesis] → [TESTED: result]
2. [Hypothesis] → [PENDING]
### Evidence Gathered
- Log at X showed Y
- Variable Z had value W
### Resolution
[Root cause and fix applied]
1---2name: agent-ops-debugging3description: Systematic debugging approaches for isolating and fixing software defects. Use when something isn't working and the cause is unclear.4---56# Skill: agent-ops-debugging78> Systematic debugging approaches for isolating and fixing software defects910---1112## Purpose1314Systematic problem isolation, root cause analysis, and defect resolution. Use when something isn't working and the cause is unclear.1516---1718## Core Principles1920### 1. Understand Before Acting2122- **Reproduce the issue**: Can you consistently trigger the problem?23- **Define expected vs actual**: What should happen vs what is happening?24- **Gather context**: When does this occur? Under what conditions?25- **Recent changes**: What changed before this appeared?2627### 2. Isolate the Problem2829- **Binary search**: Comment out half the code, test, repeat30- **Minimize reproduction**: Create minimal test case31- **Control variables**: Change one thing at a time32- **Eliminate noise**: Remove unrelated factors3334### 3. Form Hypotheses3536- **State your assumption**: "I believe X is causing Y because..."37- **Make predictions**: "If my hypothesis is true, then Z should happen"38- **Test predictions**: Verify or refute each hypothesis39- **Iterate**: Refine hypothesis based on evidence4041### 4. Fix and Verify4243- **Address root cause**: Not just symptoms44- **Minimize changes**: Smallest fix that resolves the issue45- **Add tests**: Prevent regression46- **Verify fix**: Test the specific scenario and related scenarios4748---4950## Systematic Debugging Process5152### Phase 1: Problem Definition53541. **Describe the bug** in one sentence552. **List reproduction steps** (minimal set)563. **Specify expected behavior**574. **Capture actual behavior** (screenshots, logs, error messages)585. **Identify scope**: How widespread is this?5960### Phase 2: Information Gathering61621. **Check logs**: Application logs, system logs, crash reports632. **Inspect state**: Database records, cache contents, file system643. **Review code**: Recent changes, related code paths654. **Compare environments**: Dev vs staging vs production differences665. **Monitor resources**: CPU, memory, disk, network during issue6768### Phase 3: Hypothesis Formation6970Common failure patterns:7172| Pattern | Symptoms | Where to Look |73|---------|----------|---------------|74| **Timing issues** | Intermittent, "works sometimes" | Race conditions, deadlocks, timeouts |75| **State corruption** | Wrong data, unexpected mutations | Shared state, caches, global variables |76| **Resource exhaustion** | Slows down, eventually fails | Memory leaks, connection pools |77| **Configuration** | Works elsewhere, fails here | Environment variables, settings files |78| **Dependencies** | Broke after update | Library versions, API changes |79| **Assumption violations** | Edge case failures | Code assumes something that isn't true |8081### Phase 4: Hypothesis Testing82831. **Add logging**: Instrument code to verify assumptions842. **Use debugger**: Set breakpoints, inspect variables, step through853. **Write tests**: Create failing test that reproduces bug864. **Simplify**: Remove complexity while preserving failure875. **Verify**: Confirm hypothesis explains all symptoms8889### Phase 5: Resolution90911. **Implement fix**: Address root cause, not symptoms922. **Add regression test**: Ensure bug doesn't return933. **Review similar code**: Check for same issue elsewhere944. **Document**: Add comments, update docs if behavior changed955. **Verify**: Test fix works and doesn't break other things9697---9899## Debugging by Symptom100101### "It Works on My Machine"102103| Check | Action |104|-------|--------|105| Environment differences | Python versions, OS, dependencies |106| Uncommitted config | Local settings, .env files |107| Race conditions | Timing-dependent issues |108| Data differences | Test with production data subset |109| Resource constraints | Production may have different limits |110111### Intermittent Failures112113| Check | Action |114|-------|--------|115| Shared state | Global variables, singletons, caches |116| Timing | Race conditions, timeouts, async issues |117| Randomness | Random seeds, shuffling, sampling |118| Resource cleanup | Are resources properly released? |119| External dependencies | Network calls, third-party services |120121### Performance Degradation122123| Check | Action |124|-------|--------|125| Profile first | Measure before optimizing |126| O(n²) | Nested loops, repeated work |127| I/O | Database queries, file reads, network |128| Memory | Leaks, large objects, excessive allocations |129| Caching | Repeated expensive operations |130131### Memory Leaks132133| Check | Action |134|-------|--------|135| Profile memory | Track allocations over time |136| Circular references | GC can't collect cycles |137| Event listeners | Detached handlers keeping objects alive |138| Caches | Growing without bounds |139| Static collections | Accumulating entries |140141### Deadlocks142143| Check | Action |144|-------|--------|145| Lock order | Identify held locks, acquisition order |146| Cycles | A waits for B, B waits for A |147| Timeouts | Are operations waiting indefinitely? |148| Hold-and-wait | Holding one lock while waiting for another |149150---151152## Tool-Specific Guidance153154### Print/Log Statements155156```python157# Strategic placement with unique markers158print(f"[DEBUG-001] user_id={user_id}, state={state}")159160# Include enough context161logger.debug(f"Processing item {i}/{total}: {item.id}")162163# Remove after debugging!164```165166### Debugger167168- Set breakpoints at suspicious locations, not everywhere169- Watch expressions for specific variables170- Check call stack to understand how you got here171- Step carefully through suspicious code172173### Tests for Debugging174175- Write failing test that captures bug reproduction176- Use `git bisect` to find when bug was introduced177- Mock external dependencies to isolate178- Property-based testing finds edge cases179180---181182## Anti-Patterns to Avoid183184| Anti-Pattern | Problem | Better Approach |185|--------------|---------|-----------------|186| **Shotgun debugging** | Random changes hoping something works | Form hypothesis, test, refine |187| **Symptom treatment** | Adding error handling to hide failures | Fix underlying cause |188| **Assuming** | "This variable can't be null" | Add assertion to verify |189| **Overcomplicating** | Complex debugging infrastructure | Start simple, add tools as needed |190| **Ignoring evidence** | Dismissing data that doesn't fit | Revise hypothesis to explain all |191192---193194## Debugging Checklist195196Before declaring "debugged":197198- [ ] Root cause identified, not just symptom treated199- [ ] Fix is minimal and targeted200- [ ] Regression test added201- [ ] Related code checked for same issue202- [ ] Documentation updated if needed203- [ ] Fix verified in realistic scenario204- [ ] No new issues introduced205206---207208## When to Escalate209210Consider asking for help if:211212- After 2 hours without progress213- Issue is in unfamiliar technology stack214- Problem involves complex distributed systems215- Security implications216- Production outage217- Going in circles (revisiting same hypotheses)218219---220221## Recording Debug Sessions222223Track in `.agent/focus.md`:224225```markdown226## Debugging: [Issue Description]227228**Symptom**: [What's happening]229**Expected**: [What should happen]230**Reproduction**: [Steps to trigger]231232### Hypotheses2331. [Hypothesis] → [TESTED: result]2342. [Hypothesis] → [PENDING]235236### Evidence Gathered237- Log at X showed Y238- Variable Z had value W239240### Resolution241[Root cause and fix applied]242```