Vllm Setup

Deploy a vLLM inference server on an NVIDIA DGX Station GB300 with validated container, GPU targeting, and tuning parameters. Use when the user asks to serve a model with vLLM, start a vLLM endpoint, or set up OpenAI-compatible inference on DGX Station.

NVIDIA 15d2609 2.9 KB Updated 2.2k repo stars

File contents

nvidia/dgx-spark-playbooks/tree/main/nvidia/station-ai-skills/assets/skills/vllm-setup commit 15d260960f

Frequently asked questions

npx skillmds@latest add nvidia/vllm-setup