Dgx Station Inference

Resolve, tune, preflight, launch, verify, inspect, and stop exact-model inference on NVIDIA DGX Station through dgx-assist. Use for vLLM or SGLang container selection, NGC versus upstream, GPU memory utilization, CPU or KV offload, HBM fit, KV-cache sizing, ISL or context length, prefix caching, chunked prefill, batching, concurrency, performance tuning, serving or deploying a named model, an OpenAI-compatible endpoint, Station recipe models, or an owned inference service. Require an exact model ID for recipe resolution or model-specific tuning, and never recommend or substitute a different model.

NVIDIA e9ac961 6 files · 13.9 KB Updated 2.2k repo stars

File contents

nvidia/dgx-spark-playbooks/tree/main/nvidia/station-ai-skills/assets/skills/dgx-station-inference commit e9ac961974

Frequently asked questions

npx skillmds@latest add nvidia/dgx-station-inference