Vllm Benchmarking

Run production vLLM benchmarks — `vllm bench` (serve, throughput, latency, sweep, startup, mm-processor), request-rate vs max-concurrency semantics, TTFT/TPOT/ITL/E2EL percentiles, goodput SLO measurement, prefix-cache workloads, air-gapped operation (HF_ENDPOINT, ModelScope, hf-mirror, offline cache). Methodology split — SLO health checks vs A/B change sweeps — plus pitfalls that produce misleading numbers (no warmup, wrong tokenizer, random-as-prod, `--request-rate inf` alone).

air-gapped f2c2fa7 11 files · 73.2 KB Updated

File contents

air-gapped/skills/tree/main/.claude/skills/vllm-benchmarking commit f2c2fa746a

Frequently asked questions

npx skillmds@latest add air-gapped/vllm-benchmarking