Tritonify

Agent-driven Triton/CUDA kernel optimization: a roofline-targeted trial-loop that treats cuBLAS/cuDNN/Liger as baselines to BEAT — never claiming an unmeasured speedup, never calling an op impossible without checking GPU access. Use to write, optimize, fuse, profile, port, or speed up any Triton/CUDA kernel or LLM op — GEMM, MLP, MoE, attention, activation, fused/custom loss, quantized.

IsNoobgrammer a146e90 9 files · 50.7 KB Updated

File contents

IsNoobgrammer/skills-for-agents/tree/main/skills/tritonify commit a146e90785

Frequently asked questions

npx skillmds@latest add isnoobgrammer/tritonify