Xe2 Sdp Hd256

Use this skill when writing, optimizing, benchmarking, or debugging Flash Attention SDP kernels with head dimension 256 (HD=256) targeting Intel Xe2 (Lunar Lake/LNL, Battlemage/BMG) GPU using SYCL ESIMD. Xe2 is the GPU architecture; LNL and BMG are product names. Covers the S^T (transposed scores) architecture, oneDNN-inspired v2 kernel design, GQA support, softmax optimization, lsc_slm_scatter S transpose elimination, ISA-level analysis, and the complete optimization journey from 64 to 88 TFLOPS. Use whenever the user mentions HD=256 SDP, head_dim=256 attention, rev256, onednn_v2 kernel, S transpose, s_scatter, s_gather, lsc_slm_scatter, lsc_slm_gather, or large head dimension flash attention on Intel GPU.

ModelTC d30871a 15 files · 300.6 KB Updated

File contents

modeltc/lightx2v/tree/main/.claude/skills/lightx2v_kernel_skills/Intel_XPU/kernel_specific_skills/xe2-sdp-hd256 commit d30871a39a

Frequently asked questions

npx skillmds@latest add modeltc/xe2-sdp-hd256