Mla Analysis

MLA (Multi-Latent Attention) cost models, regime analysis, and kernel selection guide. Use when: (1) reasoning about which kernel approach to use for a given regime, (2) understanding cost model tradeoffs between FlashMLA, FlashAttention, and MLAvar6+, (3) analyzing roofline behavior across decode/speculative/prefill regimes, (4) setting optimization targets, (5) understanding MLA math and absorption trick.

pepperu96 1931ced 3.5 KB Updated

File contents

pepperu96/hyper-mla/tree/main/.claude/skills/mla-analysis commit 1931ced7c3

Frequently asked questions

npx skillmds@latest add pepperu96/mla-analysis