Token Sparse Attention Long Context

Dynamically select important tokens at the attention head level, performing dense attention only on selected tokens and scattering results back. Achieves 3.23x attention speedup at 128K context with 1% accuracy loss through layer-wise representation stability analysis.

adu2021 7fb6aa5 13.7 KB Updated

File contents

adu2021/skillxiv/tree/main/skills/skillxiv-v0.0.2-claude-opus-4.6/token-sparse-attention-long-context commit 7fb6aa5b3e

Frequently asked questions

npx skillmds@latest add adu2021/token-sparse-attention-long-context