LLM Serving Auto Benchmark

Framework-independent LLM serving benchmark skill for comparing SGLang, vLLM, TensorRT-LLM, TokenSpeed, or another serving framework. Use when a user wants to find the best deployment command for one model across multiple serving frameworks under the same workload, GPU budget, and latency SLA.

bbuf d13ae94 50 files · 230.2 KB Updated

File contents

bbuf/ai-infra-auto-driven-skills/tree/main/skills/llm-serving-auto-benchmark commit d13ae9405f

Frequently asked questions

npx skillmds@latest add bbuf/llm-serving-auto-benchmark