LLM Serving Capacity Planner

Parse SGLang/vLLM startup logs to explain GPU memory use and request capacity. Use for KV cache budget, mem-fraction-static comparisons, OOM triage, and max-concurrency estimates.

bbuf a6c088b 4 files · 65.2 KB Updated

File contents

bbuf/ai-infra-auto-driven-skills/tree/main/skills/llm-serving-capacity-planner commit a6c088ba1b

Frequently asked questions

npx skillmds@latest add bbuf/llm-serving-capacity-planner