Vllm Docs

Use when working with vLLM inference engine: OpenAI-compatible serving, model deployment, quantization (AWQ, GPTQ, FP8, GGUF, INT4/INT8), speculative decoding, LoRA adapters, structured outputs, tool calling, multimodal inputs, distributed serving (tensor/pipeline/expert/context parallel), Docker/Kubernetes deployment, engine configuration, memory optimization, PagedAttention, offline inference, CLI usage, or troubleshooting vLLM issues.

wenerme 9c68180 168 files · 1.5 MB Updated

File contents

wenerme/ai/tree/main/skills/vllm-docs commit 9c68180bb0

Frequently asked questions

npx skillmds@latest add wenerme/vllm-docs