Amd Kernel Optimization

Optimize inference latency and throughput of PyTorch models on AMD GPUs (MI250/MI300/MI350) with ROCm. Use when profiling and optimizing GEMM, attention, elementwise ops, torch.compile, CUDAGraphs, or Triton kernels on AMD hardware. Covers the full optimize cycle: benchmark → profile → analyze → implement → verify. Also covers benchmarking methodology and common pitfalls that waste time.

wenyi-li Updated

File contents

wenyi-li/awesome-agent-kernel-skills/tree/main/amd-kernel-optimization commit f481b36d02

Frequently asked questions

npx skillmds@latest add wenyi-li/amd-kernel-optimization