LLM Serving Auto Benchmark

Framework-independent LLM serving benchmark skill for comparing SGLang, vLLM, TensorRT-LLM, or another serving framework. Use when a user wants to find the best deployment command for one model across multiple serving frameworks under the same workload, GPU budget, and latency SLA.

kunpengcompute 8c720e4 46 files · 179.2 KB Updated

File contents

kunpengcompute/sglang/tree/main/.claude/skills/llm-serving-auto-benchmark commit 8c720e4952

Frequently asked questions

npx skillmds@latest add kunpengcompute/llm-serving-auto-benchmark