Vllm Bench

Use vllm-bench to benchmark OpenAI-compatible or vLLM serving endpoints, especially multi-turn chat load tests. Use when the user asks for vllm bench, vllm-bench, multi-turn benchmark, chat serving benchmark, TTFT/TPOT/throughput measurement, concurrency sweep, load test, 压测, 多轮压测, or wants help installing, running --help, choosing flags, validating a local PegaInfer/vLLM server, saving JSON results, or interpreting vllm-bench metrics.

openinfer-project be4060b 2 files · 9.4 KB Updated

File contents

openinfer-project/openinfer/tree/main/.claude/skills/vllm-bench commit be4060b473

Frequently asked questions

npx skillmds@latest add openinfer-project/vllm-bench-2