Vllm Bench Serve

Interactive online benchmark orchestrator for vLLM inference services using `vllm bench serve`. Supports single benchmarks, multi-case batch execution with result aggregation, and auto-optimization to find optimal concurrency/throughput under latency SLO constraints (TTFT, TPOT, P99, success rate). Use this skill whenever the user wants to benchmark, stress test, or measure performance of a running LLM/multimodal/embedding inference service — even if they don't say "vllm bench serve" explicitly. Do NOT use for offline inference throughput, service deployment/startup, profiling/tracing, health checks only, or analyzing existing benchmark results without running new tests.

ascend-ai-coding 89002e9 18 files · 137.6 KB Updated

File contents

ascend-ai-coding/awesome-ascend-skills/tree/main/skills/inference/vllm-bench-serve commit 89002e93d0

Frequently asked questions

npx skillmds@latest add ascend-ai-coding/vllm-bench-serve