Vllm Sota Humanize Loop

Run an autonomous Humanize-governed vLLM SOTA performance loop for one LLM model: first perform the fixed fair vLLM/SGLang/TensorRT-LLM deployment search and benchmark, then start one RLCR loop that repeatedly decides the gap, profiles the current bottleneck, runs layer/kernel pipeline analysis, patches vLLM code, optionally uses ncu-report-skill for kernel evidence, and revalidates until vLLM matches or beats the best observed framework under the same workload and SLA.

bbuf da492c1 2 files · 31.5 KB Updated

File contents

bbuf/ai-infra-auto-driven-skills/tree/main/skills/vllm-sota-humanize-loop commit da492c1b58

Frequently asked questions

npx skillmds@latest add bbuf/vllm-sota-humanize-loop