Kernel Loop

Iterative GPU kernel optimization orchestrator for CUDA/CUTLASS/CuTe DSL/Triton kernels. Use for measured, one-change-at-a-time optimization loops with correctness, NCU profiling, KBS evidence, hypothesis discipline, hard iteration gates, final benchmarking, and a traceable report.

fmh66 acae759 5 files · 26.1 KB Updated

File contents

fmh66/kernel-opt-agent/tree/main/skills/kernel-loop commit acae75908f

Frequently asked questions

npx skillmds@latest add fmh66/kernel-loop