Fret perf + Tracy (end-to-end attribution workflow)
When to use
Use this skill when you want to go beyond “numbers + worst bundle” and answer:
- What exactly ran on the UI thread in the worst frame?
- Was the hitch real CPU work or scheduling noise?
- Which spans/roots/widgets explain the bundle’s top metrics?
This is especially useful for:
- Tail spikes (max/p95) that are hard to reason about from aggregates alone.
- “Fast alone, slow in suite” behavior (state contamination).
- Renderer regressions where
paintdominates and you need a timeline.
Quick start
- Run a perf suite with a baseline gate:
cargo run -p fretboard-dev --release -- diag perf ui-gallery-steady --repeat 3 --warmup-frames 5 --dir target/fret-diag --perf-baseline docs/workstreams/perf-baselines/ui-gallery-steady.windows-rtx4090.v1.json --env FRET_DIAG_SCRIPT_AUTO_DUMP=0 --env FRET_DIAG_SEMANTICS=0 --env FRET_UI_GALLERY_VIEW_CACHE=1 --env FRET_UI_GALLERY_VIEW_CACHE_SHELL=1 --launch -- cargo run -p fret-ui-gallery --release
- For each failure, jump straight to the evidence bundle:
- Open
target/fret-diag/check.perf_thresholds.json - For each
failures[].evidence_bundle, run:cargo run -p fretboard-dev --release -- diag stats <bundle.json> --sort cpu_cycles --top 30
- Reproduce the same script under Tracy:
cargo run -p fretboard-dev --release -- diag repro <script.json> --with tracy --dir target/fret-diag --env FRET_DIAG_SCRIPT_AUTO_DUMP=0 --env FRET_DIAG_SEMANTICS=0 --env FRET_UI_GALLERY_VIEW_CACHE=1 --env FRET_UI_GALLERY_VIEW_CACHE_SHELL=1 --launch -- cargo run -p fret-ui-gallery --release
Then open Tracy and look for the same span taxonomy you see in bundles (fret.frame, fret.ui.layout.*, fret.ui.paint.*, cache-root spans, renderer spans).
Workflow
- Decide what you’re optimizing
- Tail smoothness: use
diag perfwithmaxbaselines and correlate the single worst bundle. - Typical perf: seed and gate p95/p90 baselines; don’t overfit to one spike.
- Attribute from bundle first (cheap)
- Use
diag stats --sort cpu_cyclesto separate “real work” vs “wall-clock noise”. - Check phase split (
layoutvspaint), then the hot breakdown:- Layout:
layout.engine_solve, invalidation walks, request/build roots, view-cache. - Paint:
paint.widget, cache replay/misses, scene encoding, text prepare.
- Layout:
- Move to Tracy only when you have a hypothesis
- Use Tracy to validate:
- “Which span is actually the long pole?”
- “Is it a single big region or many small regions?”
- “Are we blocked on GPU submission or doing CPU work?”
- Add instrumentation safely (low overhead by default)
- Prefer
tracingspans that are:- Cheap when disabled (
tracing::enabled!fast-path). - Stable names (so comparison across runs is meaningful).
- Field-light (avoid formatting strings on hot paths).
- Cheap when disabled (
- Prefer timing helpers that use
fret_core::time::Instant:fret_perf::measure(...)fret_perf::measure_span(...)
- Keep “profiling mode” explicit:
- Environment-gated instrumentation (e.g.
FRET_LAYOUT_PROFILE=1) should be treated as profiling-only and is expected to change perf numbers.
- Environment-gated instrumentation (e.g.
- Close the loop
- Re-run
diag perfwithout profiling flags to confirm you improved the gate surface. - Keep changes reversible and leave an evidence note in the active workstream doc.
Evidence anchors
- Perf gate evidence:
target/fret-diag/check.perf_thresholds.json - Bundle attribution:
cargo run -p fretboard-dev --release -- diag stats <bundle.json>(ortarget/release/fretboard-dev[.exe] ...if already built) - Tracy usage + span taxonomy:
docs/tracy.mddocs/adr/0036-observability-tracing-and-ui-inspector-hooks.mddocs/adr/0181-ui-automation-and-debug-recipes-v1.md
- Layout and cache-root spans:
crates/fret-ui/src/tree/layout/mod.rscrates/fret-ui/src/tree/paint/mod.rs
Examples
- Example: correlate a worst frame with a Tracy timeline
- User says: "We have the worst bundle—what actually ran on the UI thread?"
- Actions: reproduce the same script, capture a Tracy trace, and map spans to the bundle's top phases.
- Result: root-cause attribution that survives "it depends" discussions.
Common pitfalls
- Running perf gates with profiling flags on (e.g.
FRET_LAYOUT_PROFILE=1): you’ll “fail” due to instrumentation overhead. - Capturing Tracy without a script repro: you’ll get a noisy timeline that’s hard to compare.
- Adding spans that allocate/format on hot paths: prefer IDs and small integers; avoid building debug paths unless diagnostics are on.
Troubleshooting
- Symptom: traces are noisy and not comparable.
- Fix: keep the same script and fixed setup; avoid mixing profiling-only flags with perf gate runs.
- Symptom: the act of profiling changes perf.
- Fix: use profiling as attribution only; confirm improvements with
diag perfwithout instrumentation flags.
- Fix: use profiling as attribution only; confirm improvements with
Related skills
fret-diag-workflow: scripts, bundles, triage, and perf baselines/gates.fret-perf-optimization: turning hitches into durable perf contracts and landing reversible fixes.fret-framework-maintainer-guide: contract-first changes and evidence discipline.