Vllm Bench Serve

Benchmark vLLM or OpenAI-compatible serving endpoints using vllm bench serve. Supports multiple datasets (random, sharegpt, sonnet, HF), backends (openai, openai-chat, vllm-pooling, embeddings), throughput/latency testing with request-rate control, and result saving. Use when benchmarking LLM serving performance, measuring TTFT/TPOT, or load testing inference APIs.

vllm-project Updated

File contents

vllm-project/vllm-skills/tree/main/plugins/vllm-skills/skills/vllm-bench-serve commit 74556b2272

Frequently asked questions

npx skillmds@latest add vllm-project/vllm-bench-serve