Numerics Debugging (DebugMode-based)
Per-op activation capture + comparison toolkit. Captures activations on a
designated step via torch.utils._debug_mode.DebugMode, then diffs two
captures into an HTML report to surface numerics divergence between runs
that should agree (bitwise, or within float32 reduction-order noise).
Two pieces live in scripts/:
activation_tracer.py— runtime capture, driven fromProfilerbyActivationCaptureProfiler. Output:{dump_folder}/numerics/rank_{N}_activations.log.compare_numerics.py— diffs two logs, produces an HTML report. Standard library only, no torch import.
They sit outside the torchtitan package on purpose. Nothing in core
torchtitan or graph_trainer references them; an agent must edit torchtitan
to wire the tracer in before a capture run, and revert the edits when done
(they don't belong on main). Because .claude is not a valid Python
package name, activation_tracer is imported by putting scripts/ on
sys.path rather than by dotted module path — see
references/patching.md.
The two runs being compared must use the same dtype and seed. The matcher keys on shape + float64 L1 norm; a precision change (bf16 vs fp32) makes every row diverge and the matcher degrades to the structural-only
statspass.
Workflow
- Patch torchtitan to wire the capture into
Profiler(andgraph_trainerif you're capturing the traced path). Full patch set: references/patching.md. - Capture twice, once per run you want to compare:
The capture step is./run_train.sh \ --dump_folder ./outputs/run_A \ --training.steps 2 \ --profiler.dump_numerics \ --profiler.profile_freq 2 \ --debug.seed 42 \ --debug.deterministic \ --training.mixed_precision_param float32profile_freq. Withprofile_freq=2andtraining.steps=2, step 1 warms up and step 2 is the snapshot. Capture adds ~10–40% memory only on the capture step (stats are computed inline in float64; tensors aren't held). - Diff the two logs:
Openpython .claude/skills/numerics_debugging/scripts/compare_numerics.py \ outputs/run_A/numerics/rank_0_activations.log \ outputs/run_B/numerics/rank_0_activations.log \ --name1 run_A --name2 run_B \ -o diff.htmldiff.html. Each row pairs one op from each run; cells turn red when a stat diverges; the "Match method" chip shows which of the four matching passes (override / exact key / fuzzy key / stats) paired the row.
Customizing
Excluded ops, numel / dtype filter, hash function, the manual-override file format, and HTML appearance are all tunable. See references/customization.md, which also catalogs the common eager-vs-traced mismatch patterns (AC-recompute FQN drift, per-layer counter shifts, collective renaming) you'll see in the diff and how to express them as overrides.