Vllm

Serves LLMs with vLLM PagedAttention and continuous batching: OpenAI-compatible /v1/chat/completions, GPTQ/AWQ/FP8 quantization, tensor parallelism, and offline LLM.generate. Use when deploying high-throughput GPU inference endpoints or fitting 30B-70B models. Not for TRL training, llama.cpp CPU/edge, or using vLLM only as an inspect-ai/lighteval backend.

Kayforkind 69af1c8 5 files · 36.6 KB Updated

File contents

Kayforkind/skill-slice commit 69af1c858d

Frequently asked questions

npx skillmds@latest add kayforkind/vllm