Sglang Setup

Deploy an SGLang inference server on an NVIDIA DGX Station GB300 with the cu130 container, RadixAttention prefix caching, and structured JSON output support. Use when the user asks to serve a model with SGLang, start an SGLang endpoint, or needs structured-output inference on DGX Station.

NVIDIA 41913a3 4.1 KB Updated 2.2k repo stars

File contents

nvidia/dgx-spark-playbooks/tree/main/nvidia/station-ai-skills/assets/skills/sglang-setup commit 41913a358a

Frequently asked questions

npx skillmds@latest add nvidia/sglang-setup