When to use
Whenever you are given an error, stack trace, failing test, or "it's broken" report and need to find and fix the real cause. Resist changing code before you understand the failure.
Not for: adding new features, speculative refactors, or performance tuning where nothing is actually broken. If there's no defect to reproduce, this isn't the skill.
Method
- Reproduce. Pin exact steps, inputs, environment, and version that trigger it, plus frequency (always / intermittent). A bug you can't reproduce, you can't confirm fixed. If it's intermittent, capture logs across several runs before theorizing.
- Read the evidence. Parse the full stack trace top to bottom; find the first frame in your own code. Capture the error text verbatim — don't paraphrase it away.
- Isolate. Shrink the surface: bisect commits, disable branches, add targeted logging, or minimize the input to the smallest failing case.
- Form one hypothesis at a time. State what you believe is wrong and what output would confirm or refute it. If the observed output contradicts the hypothesis, the output wins — discard it and form the next one; never edit code to fit a theory you haven't confirmed.
- Diagnose root cause, not symptom. Ask why the bad state arose, not just where it surfaced. Separate trigger from underlying defect.
- Fix minimally, then verify by re-running the original repro. Add a regression test that fails before the fix and passes after. If the repro still fails, you fixed the wrong thing — return to step 4.
Example
Symptom: 500 on POST /cart; "TypeError: cannot read 'price' of undefined"
Repro: add SKU that was deleted mid-session → always fails.
Isolate: first own-code frame = cart.ts:42, item lookup returns undefined.
Hypothesis: deleted SKUs aren't filtered before price read. Confirmed: log
shows lookup miss returns undefined, not a guard.
Root cause: cart.ts:42 assumes every SKU still exists.
Fix: skip + warn on missing SKU. Repro now returns 200.
Regression test: test_cart_with_deleted_sku (red before, green after).
Pitfalls
- Fixing the symptom (swallowing the null) instead of the cause (why it's null).
- Changing several things at once so you can't tell which one worked.
- Declaring it fixed without re-running the exact repro — "should be fixed" is not verification.
- Paraphrasing the error and chasing the wrong function; use the verbatim message and first own-code frame.
- Debugging the wrong environment. The error is in prod/the live process but you're reading local code or a stale build; confirm the running version (commit, restart time, loaded config) matches what you're reading before theorizing.
- Trusting the report over the log. "It's broken when I click X" is a starting point, not evidence — pull the actual log/trace first; user reports routinely misattribute the trigger.
Environment-mismatch checks (do these before step 3)
When symptoms make no sense against the code you're reading, suspect a version/config skew:
- Confirm the running process's version: commit hash, process start time vs last deploy/restart, and whether a source change actually reached the runtime (compiled bundle vs live source).
- Confirm the effective config: env vars in the LIVE process (not just the .env file — files only reach the process after a restart), feature flags, and which database/endpoint it's actually pointed at.
- Reproduce in the same environment where the failure was reported; a repro that only works locally is a different bug.
Output format
Symptom: <observed behavior + verbatim error text>
Reproduction: <exact steps, inputs, env, frequency>
Root cause: <the actual defect, file:line>
Fix: <what changed and why it addresses the cause>
Verification: <repro re-run result + regression test name>
Prevention: <guard, test, or invariant to stop recurrence>
1---2name: eng-debug-23description: Diagnose and fix a bug through a disciplined reproduce, isolate, diagnose, fix, verify loop driven by evidence and one hypothesis at a time — not guesswork.4---56## When to use78Whenever you are given an error, stack trace, failing test, or "it's broken" report and need to find and fix the real cause. Resist changing code before you understand the failure.910**Not for:** adding new features, speculative refactors, or performance tuning where nothing is actually broken. If there's no defect to reproduce, this isn't the skill.1112## Method13141. **Reproduce.** Pin exact steps, inputs, environment, and version that trigger it, plus frequency (always / intermittent). A bug you can't reproduce, you can't confirm fixed. If it's intermittent, capture logs across several runs before theorizing.152. **Read the evidence.** Parse the full stack trace top to bottom; find the first frame in your own code. Capture the error text verbatim — don't paraphrase it away.163. **Isolate.** Shrink the surface: bisect commits, disable branches, add targeted logging, or minimize the input to the smallest failing case.174. **Form one hypothesis at a time.** State what you believe is wrong and what output would confirm or refute it. If the observed output contradicts the hypothesis, the output wins — discard it and form the next one; never edit code to fit a theory you haven't confirmed.185. **Diagnose root cause, not symptom.** Ask why the bad state arose, not just where it surfaced. Separate trigger from underlying defect.196. **Fix minimally, then verify** by re-running the original repro. Add a regression test that fails before the fix and passes after. If the repro still fails, you fixed the wrong thing — return to step 4.2021## Example2223```24Symptom: 500 on POST /cart; "TypeError: cannot read 'price' of undefined"25Repro: add SKU that was deleted mid-session → always fails.26Isolate: first own-code frame = cart.ts:42, item lookup returns undefined.27Hypothesis: deleted SKUs aren't filtered before price read. Confirmed: log28 shows lookup miss returns undefined, not a guard.29Root cause: cart.ts:42 assumes every SKU still exists.30Fix: skip + warn on missing SKU. Repro now returns 200.31Regression test: test_cart_with_deleted_sku (red before, green after).32```3334## Pitfalls3536- **Fixing the symptom** (swallowing the null) instead of the cause (why it's null).37- **Changing several things at once** so you can't tell which one worked.38- **Declaring it fixed without re-running the exact repro** — "should be fixed" is not verification.39- **Paraphrasing the error** and chasing the wrong function; use the verbatim message and first own-code frame.40- **Debugging the wrong environment.** The error is in prod/the live process but you're reading local code or a stale build; confirm the running version (commit, restart time, loaded config) matches what you're reading before theorizing.41- **Trusting the report over the log.** "It's broken when I click X" is a starting point, not evidence — pull the actual log/trace first; user reports routinely misattribute the trigger.4243## Environment-mismatch checks (do these before step 3)4445When symptoms make no sense against the code you're reading, suspect a version/config skew:46471. Confirm the running process's version: commit hash, process start time vs last deploy/restart, and whether a source change actually reached the runtime (compiled bundle vs live source).482. Confirm the effective config: env vars in the LIVE process (not just the .env file — files only reach the process after a restart), feature flags, and which database/endpoint it's actually pointed at.493. Reproduce in the same environment where the failure was reported; a repro that only works locally is a different bug.5051## Output format5253```54Symptom: <observed behavior + verbatim error text>55Reproduction: <exact steps, inputs, env, frequency>56Root cause: <the actual defect, file:line>57Fix: <what changed and why it addresses the cause>58Verification: <repro re-run result + regression test name>59Prevention: <guard, test, or invariant to stop recurrence>60```