External Cannbot Model Model Infer Kvcache

基于 PyTorch 框架的昇腾 NPU 模型推理 KVCache 优化技能。分析和优化 LLM 文本推理模型中的 KVCache 实现,包括连续缓存、分页注意力(Paged Attention)配合 FA 融合算子、MLA 压缩缓存。触发场景包括:KVCache 管理实现、paged attention、KV 压缩、FA 融合算子、OOM/性能问题、block_table/slot_mapping 构造。基于本仓库已有模型的 KVCache 实现经验,按模型类型和场景推荐最佳方案。

ascend-ai-coding d4e2386 2 files · 23.0 KB Updated

File contents

ascend-ai-coding/awesome-ascend-skills/tree/main/external/cannbot/model/model-infer-kvcache commit d4e23865d4

Frequently asked questions

npx skillmds@latest add ascend-ai-coding/external-cannbot-model-model-infer-kvcache