Vllm Omni Diffusion Benchmark Profile

Benchmark, profile, and tune vLLM-Omni diffusion models, especially Wan/Qwen/Helios style pipelines on GPU or NPU. Use this when measuring denoise latency, collecting torch or torch_npu profiler traces, reading ASCEND_PROFILER_OUTPUT artifacts, comparing before/after performance, or diagnosing bottlenecks such as SP communication, VAE convs, data transforms, offload overlap, and RoPE overhead. Use when this capability is needed.

tomevault-io Updated

File contents

tomevault-io/skills-registry/tree/main/hsliuustc0106--vllm-omni-skills--vllm-omni-skills commit 4f6a1263c1

Frequently asked questions

npx skillmds@latest add tomevault-io/vllm-omni-diffusion-benchmark-profile