Agentsop Vllm

Decision SOP for serving LLMs with vLLM. Covers PagedAttention mental model, quantization/parallelism/batching tradeoffs, OOM triage, and when NOT to use vLLM. Activates when a coder-agent is choosing or tuning an inference engine, debugging vLLM throughput/latency/OOM, or comparing vLLM against TGI/SGLang/TensorRT-LLM/llama.cpp.

agentsope cb2852e 8 files · 63.2 KB Updated

File contents

agentsope/SkillAlchemy/tree/main/skills/agentsop-vllm commit cb2852e68b

Frequently asked questions

npx skillmds@latest add agentsope/agentsop-vllm