Quantization

Work on vLLM-Omni quantization for diffusion, autoregressive, omni, or multi-stage models. Use when choosing or adding methods such as fp8, int8, gguf, mxfp8, mxfp4, mxfp4_dualscale, ModelOpt, AutoRound, INC, msModelSlim, awq, or gptq; debugging quantized loading; or validating memory, speed, and output quality.

vllm-project b58daaa 6 files · 23.5 KB Updated

File contents

vllm-project/vllm-omni/tree/main/.claude/skills/quantization commit b58daaa281

Frequently asked questions

npx skillmds add vllm-project-vllm-omni/quantization