LLM Serving Auto Benchmark

Framework-independent LLM serving benchmark for SGLang, vLLM, TensorRT-LLM. Config-driven search_space over launch flags under shared workload, GPU budget, and latency SLA. Use for cross-framework deployment comparison, cookbook sweeps, or finding best serve command before QPS binary search. Pairs with model-perf-binary-search for SLO-max-QPS tuning after a winner is chosen.

Saddss Updated

File contents

Saddss/cursor-skills/tree/main/skills/llm-serving-auto-benchmark commit 462c5cae2d

Frequently asked questions

npx skillmds@latest add saddss/llm-serving-auto-benchmark