SGLang Setup on DGX Station
Deploy an SGLang inference server on DGX Station with validated configuration.
Steps
Find the GB300 GPU index. Run:
nvidia-smi --query-gpu=index,name --format=csv,noheader
Identify the device index for the GB300 (typically device 1). Use this index for --gpus below. Do NOT use --gpus all — mixed coherency will cause CUDA failures.
Ask the user which model to serve. If they don't have a preference, suggest:
Qwen/Qwen3-8B — small, fast, good for testing
Qwen/Qwen3-32B — medium, good balance
meta-llama/Llama-3.1-70B-Instruct — large general-purpose
Check if the user has an HF_TOKEN. Pass inline with -e HF_TOKEN="...".
Deploy the container. Use this validated configuration:
docker pull lmsysorg/sglang:latest-cu130
docker run -d \
--name sglang-server \
--gpus '"device=<GB300_INDEX>"' \
--ipc host \
--ulimit memlock=-1 \
--ulimit stack=67108864 \
-p 30000:30000 \
-e HF_TOKEN="<TOKEN>" \
-v "$HOME/.cache/huggingface/hub:/root/.cache/huggingface/hub" \
lmsysorg/sglang:latest-cu130 \
sglang serve --model-path "<MODEL>" \
--host 0.0.0.0 \
--port 30000 \
--context-length 32768 \
--mem-fraction-static 0.85
Container version: Use lmsysorg/sglang:latest-cu130. The cu130 tag is required for Blackwell SM103 support.
First launch downloads the model and compiles kernels. This takes extra time — subsequent starts are faster.
Wait for the server to be ready. Monitor logs:
docker logs -f sglang-server
Test the server:
curl http://localhost:30000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "<MODEL>",
"messages": [{"role": "user", "content": "Hello"}],
"max_tokens": 64
}'
Report the result to the user, including:
- Model loaded and serving on port 30000
- How to stop:
docker stop sglang-server && docker rm sglang-server
Key features
- RadixAttention — automatic KV cache reuse across requests sharing prefixes. On by default, no flag needed. Verify with:
docker logs sglang-server 2>&1 | grep "cached-token" | tail -5
- Structured JSON output — use
response_format.json_schema in API requests for guaranteed valid JSON.
- Chunked prefill — add
--chunked-prefill-size 8192 to break long prefills into chunks, reducing time-to-first-token.
Tuning parameters
| Parameter |
Default |
Agent workloads |
Throughput workloads |
--context-length |
32768 |
32768-65536 |
8192-16384 |
--mem-fraction-static |
0.85 |
0.80-0.85 |
0.85-0.88 |
--chunked-prefill-size |
off |
4096-8192 |
8192 |
--enable-metrics |
off |
Optional |
Recommended |
Structured output example
curl http://localhost:30000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "<MODEL>",
"messages": [{"role": "user", "content": "List three programming languages."}],
"max_tokens": 512,
"response_format": {
"type": "json_schema",
"json_schema": {
"name": "languages",
"schema": {
"type": "object",
"properties": {
"languages": {
"type": "array",
"items": {
"type": "object",
"properties": {
"name": {"type": "string"},
"primary_use": {"type": "string"}
},
"required": ["name", "primary_use"]
}
}
},
"required": ["languages"]
}
}
}
}'
1---2name: sglang-setup3description: Deploy an SGLang inference server on an NVIDIA DGX Station GB300 with the cu130 container, RadixAttention prefix caching, and structured JSON output support. Use when the user asks to serve a model with SGLang, start an SGLang endpoint, or needs structured-output inference on DGX Station.4---56# SGLang Setup on DGX Station78Deploy an SGLang inference server on DGX Station with validated configuration.910## Steps11121. **Find the GB300 GPU index.** Run:13 ```bash14 nvidia-smi --query-gpu=index,name --format=csv,noheader15 ```16 Identify the device index for the GB300 (typically device 1). Use this index for `--gpus` below. Do NOT use `--gpus all` — mixed coherency will cause CUDA failures.17182. **Ask the user which model to serve.** If they don't have a preference, suggest:19 - `Qwen/Qwen3-8B` — small, fast, good for testing20 - `Qwen/Qwen3-32B` — medium, good balance21 - `meta-llama/Llama-3.1-70B-Instruct` — large general-purpose22233. **Check if the user has an HF_TOKEN.** Pass inline with `-e HF_TOKEN="..."`.24254. **Deploy the container.** Use this validated configuration:2627 ```bash28 docker pull lmsysorg/sglang:latest-cu1302930 docker run -d \31 --name sglang-server \32 --gpus '"device=<GB300_INDEX>"' \33 --ipc host \34 --ulimit memlock=-1 \35 --ulimit stack=67108864 \36 -p 30000:30000 \37 -e HF_TOKEN="<TOKEN>" \38 -v "$HOME/.cache/huggingface/hub:/root/.cache/huggingface/hub" \39 lmsysorg/sglang:latest-cu130 \40 sglang serve --model-path "<MODEL>" \41 --host 0.0.0.0 \42 --port 30000 \43 --context-length 32768 \44 --mem-fraction-static 0.8545 ```4647 **Container version:** Use `lmsysorg/sglang:latest-cu130`. The `cu130` tag is required for Blackwell SM103 support.4849 **First launch** downloads the model and compiles kernels. This takes extra time — subsequent starts are faster.50515. **Wait for the server to be ready.** Monitor logs:52 ```bash53 docker logs -f sglang-server54 ```55566. **Test the server:**57 ```bash58 curl http://localhost:30000/v1/chat/completions \59 -H "Content-Type: application/json" \60 -d '{61 "model": "<MODEL>",62 "messages": [{"role": "user", "content": "Hello"}],63 "max_tokens": 6464 }'65 ```66677. **Report the result** to the user, including:68 - Model loaded and serving on port 3000069 - How to stop: `docker stop sglang-server && docker rm sglang-server`7071## Key features7273- **RadixAttention** — automatic KV cache reuse across requests sharing prefixes. On by default, no flag needed. Verify with: `docker logs sglang-server 2>&1 | grep "cached-token" | tail -5`74- **Structured JSON output** — use `response_format.json_schema` in API requests for guaranteed valid JSON.75- **Chunked prefill** — add `--chunked-prefill-size 8192` to break long prefills into chunks, reducing time-to-first-token.7677## Tuning parameters7879| Parameter | Default | Agent workloads | Throughput workloads |80|-----------|---------|-----------------|---------------------|81| `--context-length` | 32768 | 32768-65536 | 8192-16384 |82| `--mem-fraction-static` | 0.85 | 0.80-0.85 | 0.85-0.88 |83| `--chunked-prefill-size` | off | 4096-8192 | 8192 |84| `--enable-metrics` | off | Optional | Recommended |8586## Structured output example8788```bash89curl http://localhost:30000/v1/chat/completions \90 -H "Content-Type: application/json" \91 -d '{92 "model": "<MODEL>",93 "messages": [{"role": "user", "content": "List three programming languages."}],94 "max_tokens": 512,95 "response_format": {96 "type": "json_schema",97 "json_schema": {98 "name": "languages",99 "schema": {100 "type": "object",101 "properties": {102 "languages": {103 "type": "array",104 "items": {105 "type": "object",106 "properties": {107 "name": {"type": "string"},108 "primary_use": {"type": "string"}109 },110 "required": ["name", "primary_use"]111 }112 }113 },114 "required": ["languages"]115 }116 }117 }118 }'119```