Onednn Fp8 Gemm

Use this skill when implementing, optimizing, or debugging quantized GEMM kernels using oneDNN on Intel Xe2 (Lunar Lake/LNL, Battlemage/BMG) or newer Intel XPU. Xe2 is the GPU architecture; LNL and BMG are product names. Covers FP16/BF16 x FP8_E4M3 with per-N scale, FP16 x FP8 with block-wise scale along K, FP16 x INT4 (U4) with block-wise scale + zero-point, 2D block quantization emulation via repeat_interleave, bias fusion, and the critical API differences between set_scales_mask (JIT) vs set_scales (ref fallback). Use whenever the user mentions oneDNN FP8 GEMM, quantized matmul, W8A16, W4A16, per-N scale, block-wise FP8, block-wise INT4, 2D block quantization, or dnnl matmul primitive on Intel GPU.

ModelTC 8da2633 6 files · 66.3 KB Updated

File contents

modeltc/lightx2v/tree/main/.claude/skills/lightx2v_kernel_skills/Intel_XPU/kernel_specific_skills/onednn-fp8-gemm commit 8da26331a0

Frequently asked questions

npx skillmds@latest add modeltc/onednn-fp8-gemm