Sparsed Sparse Attention Diffusion Lm

Achieve up to 1.50x speedup in diffusion language models by computing head-specific sparse attention patterns once during early denoising steps and reusing them across all subsequent iterations, while preserving full attention in critical early phases to maintain generation quality and accuracy.

adu2021 a28b246 16.9 KB Updated

File contents

adu2021/skillxiv/tree/main/skills/skillxiv-v0.0.2-claude-opus-4.6/sparsed-sparse-attention-diffusion-lm commit a28b246e54

Frequently asked questions

npx skillmds@latest add adu2021/sparsed-sparse-attention-diffusion-lm