Vllm Ascend Server

Deploy vLLM inference services on Ascend NPU servers with automatic model detection and optimized configuration. Supports local and remote deployment across bare metal, containers, and Docker images. Handles model discovery, quantization auto-detection, tensor parallelism configuration, graph/eager mode selection, and service health verification. Use when users need to: (1) Start or deploy vLLM server on NPU, (2) Launch LLM inference service, (3) Configure multi-card tensor parallel deployment, (4) Enable speculative decoding (Eagle) or quantization, (5) Run vllm offline batch inference, (6) Check or test vLLM service status.

ascend-ai-coding 76fc3c5 24 files · 85.6 KB Updated

File contents

ascend-ai-coding/awesome-ascend-skills/tree/main/skills/inference/vllm-ascend-server commit 76fc3c5f1b

Frequently asked questions

npx skillmds@latest add ascend-ai-coding/vllm-ascend-server