SGLang Diffusion Benchmark and Profile
Use this skill when measuring denoise performance, finding the slow op, checking whether an existing fast path can solve it, or verifying that a hotspot is real before any kernel work in sglang.multimodal_gen.
This skill is diagnosis-first. It owns:
- checked-in denoise benchmark presets
- same-GPU quality/BCG applicability checks with repeated lossless, extra-high, and high rows
- perf dump collection and before/after comparison
torch.profiler trace capture and quick hotspot ranking
- mapping hot kernels back to known fast paths and fusion families
- packaging confirmed kernel work with enough evidence for the appropriate kernel, Nsight, or framework-specific optimization workflow
This skill does not own low-level kernel authoring or standalone Nsight workflows.
Preflight
Before running any benchmark, profiler, or kernel-validation command:
- use
scripts/diffusion_skill_env.py to derive the repo root from sglang.__file__
- verify the repo is writable
- export
HF_TOKEN before using gated Hugging Face models such as black-forest-labs/FLUX.*
- export
FLASHINFER_DISABLE_VERSION_CHECK=1
- set
SGLANG_DIFFUSION_SYNC_STAGE_PROFILING=1 when comparing stage-level
denoise/decode timings; the preset helper sets it by default unless the
caller explicitly overrides it
- for downloaded checkpoints, use the preset helper's task-owned
--model-cache-root together with --cleanup-model-cache; verify the JSONL
ledger reports zero residual weight files before moving to the next model
- choose idle GPU(s) before starting perf work; for a comparison matrix, hold
the same GPU set and verify it has no foreign process at every run boundary
Native Backend Gate
All diffusion benchmark and profiling results owned by this skill must come from the native SGLang diffusion backend.
Treat any of the following as a hard stop condition:
Falling back to diffusers backend
Using diffusers backend
Loaded diffusers pipeline
If any benchmark, perf-dump, or torch.profiler command prints one of those signals:
- stop the workflow immediately
- do not keep the generated numbers or traces as SGLang benchmark evidence
- do not continue to hotspot classification or kernel work
- first fix model resolution, pipeline selection, overlay/materialization, or other backend-selection issues so the model runs on the native SGLang diffusion path
Main Reference
- benchmark-and-profile.md — canonical denoise benchmark, perf dump, and
torch.profiler workflow; uses checked-in nightly-aligned presets plus current-source extras such as LongCat image/edit, Qwen base edit/layered, SD3.5, SANA-Video/SANA-WM, LingBot Video/World, Cosmos3 Edge/Super I2V/distilled and the explicit Super TP2 x CFG2 comparator, LTX-2.5 and its diffusion decoder, MiniMax-H3, FLUX.2 Klein, Ideogram4, ERNIE/GLM/SANA image models, FastWan2.1/2.2, the Blackwell-only Wan2.2 NVFP4 comparator, LTX-2.3, HunyuanVideo, MOVA, Helios, image edit, Hunyuan3D shape, and a separate Pi0.5 action-policy lane
- existing-fast-paths.md — map bottlenecks to existing fused kernels, MoE routing, packed QKV paths, fused
QK norm + RoPE, distributed overlap patterns, and open optimization PRs before proposing new code
- scripts/diffusion_skill_env.py — preflight helper: repo root discovery from the skill's owning checkout before falling back to
sglang.__file__, write-access probe, benchmark/profile output directories, idle GPU selection
- scripts/bench_diffusion_denoise.py — end-to-end denoise benchmark preset runner via
sglang generate; defaults to eager/lossless, supports explicit quality and BCG comparators plus a same-GPU applicability matrix, rejects invalid BCG capture/fallback logs and late high-quality DiT fusion mounts, forces H3 to its eager consistency mode, enables synchronized stage attribution, validates nightly preset drift, and can clean one isolated model cache after the full matrix in a finally block with a JSONL ledger
Opportunity Discovery Rule
Before calling a diffusion hotspot "new", first classify it with existing-fast-paths.md.
Always rule out these existing families first:
- HunyuanVideo VAE GroupNorm+SiLU
- LTX upsampler GroupNorm+SiLU
- Z-Image bf16-native Triton RMSNorm scale/tanh-residual modulation
- SANA packed self-attention Q/K/V and cross-attention K/V GEMMs
- SANA-Video's packed projections and request-scoped BF16-input linear
attention at
quality=extra-high or quality=high; keep the second attention GEMM in FP32 and
compare against quality=lossless before changing its precision further
- SANA-Video reuse of SANA's bit-exact bias/activation, residual-gate, and
LayerNorm-modulation fast paths before adding video-only kernels
- MiniMax-H3 indexed modulation, fused QK norm + RoPE, packed Ulysses QKV,
USP relayout, and batched TP AdaLN collectives
- bit-exact diffusion adaLN modulation and fused LayerNorm + modulation for
FLUX.1, GLM-Image, and SANA
- request-scoped DiT and VAE fast paths at
quality=extra-high or quality=high
- LingBot Video's default-on fused group-limited top-k expert selection before
treating its router's
topk/mask/gather chain as a new hotspot
- Wan causal-VAE cache/padding and DupUp3D data-movement fusions
- fused diffusion
QK norm + RoPE
- LTX2 split RoPE
- LTX2 residual-gate add
- LTX-2.5 diffusion-decoder NATTEN selection before interpreting a
FlexAttention fallback trace
- varlen USP attention pack/scatter
- NVFP4 / Nunchaku packed QKV
- Nunchaku fused GELU MLP
- Ulysses / USP attention overlap
- turbo-layer async all-to-all overlap
torch.compile compute / communication reorder
- breakable CUDA graph capture for supported fixed-resolution pipelines
- dual-stream diffusion execution
The checked-in helper defaults to eager. Use --torch-compile only for a
controlled comparator, never for the eager ground truth. The legacy
--no-torch-compile spelling remains accepted but is redundant.
For kernel/BCG discovery, run --quality-bcg-matrix. It executes Eager/BCG as
A-B-B-A at lossless, then repeats the pair at extra-high and high, on
one locked GPU set and one isolated checkpoint cache. The extra-high/high+BCG
rows are applicability checks, not presumed-valid performance cells. A BCG row is invalid unless the log
contains [Diffusion BCG] captured and contains no support-disable,
capture-failure, serving-signature-miss, or late quality-fusion marker. In
particular, a request-scoped DiT fusion mounted after lossless warmup capture
would be bypassed by replay; reject that row even when capture and signature
checks pass. For video presets, the helper declares both the request resolution
and --warmup-num-frames so the synthetic BCG warmup captures the requested
temporal shape. Treat any remaining temporal or conditioning signature miss as
Eager fallback, not as a valid BCG measurement.
A zero process exit is not sufficient evidence: every accepted row must also
contain its requested perf dump and a generated image, video, audio, or 3D mesh
file.
The helper gives every cell a unique output name and rejects missing artifacts.
On machines with a read-only Hugging Face cache, combine
--model-cache-root <task-owned-dir> with one or more
--seed-model-cache-root <read-only-HF-home-or-hub> options. The helper exposes
cached repos through a task-owned copy-on-write directory overlay, downloads
misses only into the isolated cache, and removes links plus new downloads in
its normal cleanup finally block without modifying the seed cache.
Keep prompt, negative prompt, seed, shape, steps, guidance, dtype, topology,
and residency fixed. Lossless comparisons require byte-identical artifacts.
For quality=extra-high and quality=high, report aggregate and worst-frame SSIM/PSNR; the repository
defaults are 0.95/28 dB for images and 0.92/24 dB for video unless the model's
checked-in consistency metadata defines a different threshold. A performance
PR needs repeated saved-request e2e improvement of at least 1.5%, a
representative profile, and before/after image or video evidence.
MiniMax-H3 is always an eager consistency case on current main. Use
--model minimax-h3-t2va; its preset writes the H3 request fields through a
generated config and suppresses the helper's global compile default. Do not
turn the model's nominal BCG support gate into a performance claim: prompt-
dependent packed-sequence host boundaries can differ between warmup and the
serving request. A valid H3 BCG experiment must prove that every captured
segment replays, keeps the MP4 byte-identical, and does not trade latency for
the extra graph memory.
For FLUX-family manual profiling runs with a quantized transformer override:
- use
sglang generate directly
- pass the override as
--transformer-path <dir>
- prefer
--prompt-path <file> when also fixing --output-file-name
- if the base model is already cached locally and the machine has unreliable HF access, use the local cached
--model-path plus HF_HUB_OFFLINE=1
- remember that
--profile changes latency substantially; use the non-profile perf dump for the real before/after benchmark claim
1---2name: sglang-diffusion-benchmark-profile3description: Use when benchmarking denoise latency or profiling a diffusion bottleneck in SGLang.4---5
6# SGLang Diffusion Benchmark and Profile
7
8Use this skill when measuring denoise performance, finding the slow op, checking whether an existing fast path can solve it, or verifying that a hotspot is real before any kernel work in `sglang.multimodal_gen`.
9
10This skill is diagnosis-first. It owns:
11- checked-in denoise benchmark presets
12- same-GPU quality/BCG applicability checks with repeated lossless, extra-high, and high rows
13- perf dump collection and before/after comparison
14- `torch.profiler` trace capture and quick hotspot ranking
15- mapping hot kernels back to known fast paths and fusion families
16- packaging confirmed kernel work with enough evidence for the appropriate kernel, Nsight, or framework-specific optimization workflow
17
18This skill does not own low-level kernel authoring or standalone Nsight workflows.
19
20## Preflight
21
22Before running any benchmark, profiler, or kernel-validation command:
23- use `scripts/diffusion_skill_env.py` to derive the repo root from `sglang.__file__`
24- verify the repo is writable
25- export `HF_TOKEN` before using gated Hugging Face models such as `black-forest-labs/FLUX.*`
26- export `FLASHINFER_DISABLE_VERSION_CHECK=1`
27- set `SGLANG_DIFFUSION_SYNC_STAGE_PROFILING=1` when comparing stage-level
28 denoise/decode timings; the preset helper sets it by default unless the
29 caller explicitly overrides it
30- for downloaded checkpoints, use the preset helper's task-owned
31 `--model-cache-root` together with `--cleanup-model-cache`; verify the JSONL
32 ledger reports zero residual weight files before moving to the next model
33- choose idle GPU(s) before starting perf work; for a comparison matrix, hold
34 the same GPU set and verify it has no foreign process at every run boundary
35
36## Native Backend Gate
37
38All diffusion benchmark and profiling results owned by this skill must come from the native SGLang diffusion backend.
39
40Treat any of the following as a hard stop condition:
41- `Falling back to diffusers backend`
42- `Using diffusers backend`
43- `Loaded diffusers pipeline`
44
45If any benchmark, perf-dump, or `torch.profiler` command prints one of those signals:
46- stop the workflow immediately
47- do not keep the generated numbers or traces as SGLang benchmark evidence
48- do not continue to hotspot classification or kernel work
49- first fix model resolution, pipeline selection, overlay/materialization, or other backend-selection issues so the model runs on the native SGLang diffusion path
50
51## Main Reference
52
53- [benchmark-and-profile.md](benchmark-and-profile.md) — canonical denoise benchmark, perf dump, and `torch.profiler` workflow; uses checked-in nightly-aligned presets plus current-source extras such as LongCat image/edit, Qwen base edit/layered, SD3.5, SANA-Video/SANA-WM, LingBot Video/World, Cosmos3 Edge/Super I2V/distilled and the explicit Super TP2 x CFG2 comparator, LTX-2.5 and its diffusion decoder, MiniMax-H3, FLUX.2 Klein, Ideogram4, ERNIE/GLM/SANA image models, FastWan2.1/2.2, the Blackwell-only Wan2.2 NVFP4 comparator, `LTX-2.3`, HunyuanVideo, MOVA, Helios, image edit, Hunyuan3D shape, and a separate Pi0.5 action-policy lane
54- [existing-fast-paths.md](existing-fast-paths.md) — map bottlenecks to existing fused kernels, MoE routing, packed QKV paths, fused `QK norm + RoPE`, distributed overlap patterns, and open optimization PRs before proposing new code
55- [scripts/diffusion_skill_env.py](scripts/diffusion_skill_env.py) — preflight helper: repo root discovery from the skill's owning checkout before falling back to `sglang.__file__`, write-access probe, benchmark/profile output directories, idle GPU selection
56- [scripts/bench_diffusion_denoise.py](scripts/bench_diffusion_denoise.py) — end-to-end denoise benchmark preset runner via `sglang generate`; defaults to eager/lossless, supports explicit quality and BCG comparators plus a same-GPU applicability matrix, rejects invalid BCG capture/fallback logs and late high-quality DiT fusion mounts, forces H3 to its eager consistency mode, enables synchronized stage attribution, validates nightly preset drift, and can clean one isolated model cache after the full matrix in a `finally` block with a JSONL ledger
57
58## Opportunity Discovery Rule
59
60Before calling a diffusion hotspot "new", first classify it with `existing-fast-paths.md`.
61
62Always rule out these existing families first:
63- HunyuanVideo VAE GroupNorm+SiLU
64- LTX upsampler GroupNorm+SiLU
65- Z-Image bf16-native Triton RMSNorm scale/tanh-residual modulation
66- SANA packed self-attention Q/K/V and cross-attention K/V GEMMs
67- SANA-Video's packed projections and request-scoped BF16-input linear
68 attention at `quality=extra-high` or `quality=high`; keep the second attention GEMM in FP32 and
69 compare against `quality=lossless` before changing its precision further
70- SANA-Video reuse of SANA's bit-exact bias/activation, residual-gate, and
71 LayerNorm-modulation fast paths before adding video-only kernels
72- MiniMax-H3 indexed modulation, fused QK norm + RoPE, packed Ulysses QKV,
73 USP relayout, and batched TP AdaLN collectives
74- bit-exact diffusion adaLN modulation and fused LayerNorm + modulation for
75 FLUX.1, GLM-Image, and SANA
76- request-scoped DiT and VAE fast paths at `quality=extra-high` or `quality=high`
77- LingBot Video's default-on fused group-limited top-k expert selection before
78 treating its router's `topk`/mask/gather chain as a new hotspot
79- Wan causal-VAE cache/padding and DupUp3D data-movement fusions
80- fused diffusion `QK norm + RoPE`
81- LTX2 split RoPE
82- LTX2 residual-gate add
83- LTX-2.5 diffusion-decoder NATTEN selection before interpreting a
84 FlexAttention fallback trace
85- varlen USP attention pack/scatter
86- NVFP4 / Nunchaku packed QKV
87- Nunchaku fused GELU MLP
88- Ulysses / USP attention overlap
89- turbo-layer async all-to-all overlap
90- `torch.compile` compute / communication reorder
91- breakable CUDA graph capture for supported fixed-resolution pipelines
92- dual-stream diffusion execution
93
94The checked-in helper defaults to eager. Use `--torch-compile` only for a
95controlled comparator, never for the eager ground truth. The legacy
96`--no-torch-compile` spelling remains accepted but is redundant.
97
98For kernel/BCG discovery, run `--quality-bcg-matrix`. It executes Eager/BCG as
99A-B-B-A at `lossless`, then repeats the pair at `extra-high` and `high`, on
100one locked GPU set and one isolated checkpoint cache. The extra-high/high+BCG
101rows are applicability checks, not presumed-valid performance cells. A BCG row is invalid unless the log
102contains `[Diffusion BCG] captured` and contains no support-disable,
103capture-failure, serving-signature-miss, or late quality-fusion marker. In
104particular, a request-scoped DiT fusion mounted after lossless warmup capture
105would be bypassed by replay; reject that row even when capture and signature
106checks pass. For video presets, the helper declares both the request resolution
107and `--warmup-num-frames` so the synthetic BCG warmup captures the requested
108temporal shape. Treat any remaining temporal or conditioning signature miss as
109Eager fallback, not as a valid BCG measurement.
110
111A zero process exit is not sufficient evidence: every accepted row must also
112contain its requested perf dump and a generated image, video, audio, or 3D mesh
113file.
114The helper gives every cell a unique output name and rejects missing artifacts.
115
116On machines with a read-only Hugging Face cache, combine
117`--model-cache-root <task-owned-dir>` with one or more
118`--seed-model-cache-root <read-only-HF-home-or-hub>` options. The helper exposes
119cached repos through a task-owned copy-on-write directory overlay, downloads
120misses only into the isolated cache, and removes links plus new downloads in
121its normal cleanup finally block without modifying the seed cache.
122
123Keep prompt, negative prompt, seed, shape, steps, guidance, dtype, topology,
124and residency fixed. Lossless comparisons require byte-identical artifacts.
125For `quality=extra-high` and `quality=high`, report aggregate and worst-frame SSIM/PSNR; the repository
126defaults are 0.95/28 dB for images and 0.92/24 dB for video unless the model's
127checked-in consistency metadata defines a different threshold. A performance
128PR needs repeated saved-request e2e improvement of at least 1.5%, a
129representative profile, and before/after image or video evidence.
130
131MiniMax-H3 is always an eager consistency case on current main. Use
132`--model minimax-h3-t2va`; its preset writes the H3 request fields through a
133generated config and suppresses the helper's global compile default. Do not
134turn the model's nominal BCG support gate into a performance claim: prompt-
135dependent packed-sequence host boundaries can differ between warmup and the
136serving request. A valid H3 BCG experiment must prove that every captured
137segment replays, keeps the MP4 byte-identical, and does not trade latency for
138the extra graph memory.
139
140For FLUX-family manual profiling runs with a quantized transformer override:
141- use `sglang generate` directly
142- pass the override as `--transformer-path <dir>`
143- prefer `--prompt-path <file>` when also fixing `--output-file-name`
144- if the base model is already cached locally and the machine has unreliable HF access, use the local cached `--model-path` plus `HF_HUB_OFFLINE=1`
145- remember that `--profile` changes latency substantially; use the non-profile perf dump for the real before/after benchmark claim