High Performance Gpu Operator Development

Use this skill when the user wants to generate or optimize high-performance GPU kernels using languages like Triton or native C++/CUDA. This is applicable for requests like 'write a fast Triton kernel for matrix operations', 'optimize my GPU code for throughput', 'create a custom operator that outperforms the standard library', or 'debug a memory error in a parallel kernel'. It is specifically triggered when performance bottlenecks (latency/bandwidth) or hardware-specific optimizations (register pressure/memory coalescing) are the primary concerns rather than general Python logic.

dingxingdi 8c41cbb 5 files · 19.6 KB Updated

File contents

dingxingdi/paper_fast_search_backup/tree/main/skill_bank_evolved/code/skills/high-performance-gpu-operator-development commit 8c41cbb9b4

Frequently asked questions

npx skillmds@latest add dingxingdi/high-performance-gpu-operator-development