Benchmark Model Kernels

Inspect Hugging Face decoder layers on meta tensors and plan or run per-rank BF16, FP8, and NVFP4 GEMM or fused-MoE microbenchmarks with the bundled scripts and a local FlashInfer checkout. Use when choosing a model, GPU, TP, EP, or M/token-concurrency sweep; deriving common fused QKV and gate/up shapes without loading checkpoint weights; or using FlashInfer benchmark utilities. Do not use for end-to-end server throughput or request latency.

NVIDIA a054bbb 5 files · 108.0 KB Updated 2.2k repo stars

File contents

nvidia/model-optimizer/tree/main/plugins/modelopt/skills/benchmark-model-kernels commit a054bbb308

Frequently asked questions

npx skillmds@latest add nvidia/benchmark-model-kernels