Kernel Benchmark

Standalone kernel benchmarking skill for cuda-cpp, cutlass, cute-dsl, and triton implementations. Use when the user wants to compare a custom CUDA/CUTLASS .cu kernel or CuTe DSL/Triton .py kernel against selectable PyTorch eager, torch.compile, or FlashInfer baselines, validate correctness, measure execution time with KernelBench-style CUDA event timing, or generate benchmark.md for kernel optimization results.

fmh66 0157839 4 files · 40.5 KB Updated

File contents

fmh66/kernel-opt-agent/tree/main/skills/kernel-benchmark commit 015783957f

Frequently asked questions

npx skillmds@latest add fmh66/kernel-benchmark