Triton Ascend Kernels

Provide guidance for writing and benchmarking optimized triton kernels for ascend npus(910B, 910C). Support models like Qwen, Deepseek and so on. The triton kernels could be integrated into vllm inference engine or used separately with torch. This skill should be used when the user asks to "write/optimize triton kernels" for Ascend NPUs.

small-cat Updated

File contents

small-cat/my-awesome-skills/tree/main/skills/triton-ascend-kernels commit aa585e9d34

Frequently asked questions

npx skillmds@latest add small-cat/triton-ascend-kernels