Xe2 Esimd Gemv

Use this skill when writing, optimizing, benchmarking, or debugging W4A16 or W8A16 GEMV kernels targeting Intel Xe2 (Lunar Lake/LNL, Battlemage/BMG) GPU using SYCL ESIMD. Xe2 is the GPU architecture; LNL and BMG are product names. Also covers general FP16 GEMV patterns. Covers quantized weight dequantization, SIMD vs scalar interleaving, K-split SLM reduction, VL/ROWS tuning, workgroup decomposition, uint4 unpacking, FP32 accumulation, SLM barriers, performance methodology, and all hardware constraints.

ModelTC dc5984c 9 files · 70.8 KB Updated

File contents

modeltc/lightx2v/tree/main/.claude/skills/lightx2v_kernel_skills/Intel_XPU/kernel_specific_skills/xe2-esimd-gemv commit dc5984c4b6

Frequently asked questions

npx skillmds@latest add modeltc/xe2-esimd-gemv