Sparse Sparse Attention

Overcome the attention gap in sparse transformers by training with both full and sparse attention simultaneously, aligned through bidirectional losses that encourage naturally sparser distributions while maintaining learning capability, enabling efficient inference without capability degradation.

adu2021 080d473 16.5 KB Updated

File contents

adu2021/skillxiv/tree/main/skills/skillxiv-v0.0.2-claude-opus-4.6/sparse-sparse-attention commit 080d4734e9

Frequently asked questions

npx skillmds@latest add adu2021/sparse-sparse-attention