Minicpm5 Deploy Vllm

Serve MiniCPM5-1B or MiniCPM5-2B via vLLM as an OpenAI-compatible HTTP server. Use when the user wants high-throughput production serving on NVIDIA GPU, asks for "vLLM", "OpenAI server", "REST API for MiniCPM5", or "production deployment".

OpenBMB faa1562 3.2 KB Updated

File contents

openbmb/minicpm/tree/main/skills/minicpm5-deploy-vllm commit faa1562ca8

Frequently asked questions

npx skillmds@latest add openbmb/minicpm5-deploy-vllm