Vllm Ascend

vLLM Ascend plugin for LLM inference serving on Huawei Ascend NPU. Use for offline batch inference, API server deployment, quantization inference (with msmodelslim quantized models), tensor/pipeline parallelism for distributed serving, and OpenAI-compatible API endpoints. Supports Qwen, DeepSeek, GLM, LLaMA models with Ascend-optimized kernels.

ascend-ai-coding d061afd 10 files · 70.4 KB Updated

File contents

ascend-ai-coding/awesome-ascend-skills/tree/main/skills/inference/vllm-ascend commit d061afd5b4

Frequently asked questions

npx skillmds@latest add ascend-ai-coding/vllm-ascend