Profile Model Performance

Inspect and baseline performance for FlashDreams-style model integrations and interactive demos: map the generation path, add trustworthy timing splits, build focused probes, and identify whether decode, model/denoise, cache, data transfer, or presentation dominates. Use when starting performance work on an existing model runner, demo, serving path, or downstream integration before implementing speedups. Pair with `apply-inference-optimizations` after the bottleneck is known and `validate-performance-quality` for benchmark and quality gates.

NVIDIA b48b11a 5.5 KB Updated 2.2k repo stars

File contents

nvidia/flashdreams/tree/main/skills/profile-model-performance commit b48b11aea2

Frequently asked questions

npx skillmds@latest add nvidia/profile-model-performance