Sglang

Fast structured generation and serving for LLMs with RadixAttention prefix caching. Use for JSON/regex outputs, constrained decoding, agentic workflows with tool calls, or when you need 5× faster inference than vLLM with prefix sharing. Powers 300,000+ GPUs at xAI, AMD, NVIDIA, and LinkedIn.

nota-america Updated

File contents

nota-america/forgecat-agent-profiles/tree/main/profiles/orchestra-research/ai-research-skills/for-codex/.agents/skills/ai-research-skills/12-inference-serving/sglang commit 017f611d9f

Frequently asked questions

npx skillmds@latest add nota-america/sglang