Minicpm5 Deploy Sglang

Serve MiniCPM5-1B or MiniCPM5-2B via SGLang as an OpenAI-compatible HTTP server with RadixAttention prefix cache and built-in MiniCPM5 tool-call parsing. Use when the user asks for "SGLang", "RadixAttention", "prefix cache", batch evaluation, tool calling, or wants a high-concurrency NVIDIA-GPU server alternative to vLLM.

OpenBMB 7c09a01 5.0 KB Updated

File contents

openbmb/minicpm/tree/main/skills/minicpm5-deploy-sglang commit 7c09a01e07

Frequently asked questions

npx skillmds@latest add openbmb/minicpm5-deploy-sglang