Lychee Decode Sparse Kv Sharing

Classify attention heads into retrieval (full attention) and sparse (token-selected) roles using HardKuma distribution for differentiable discrete optimization. Sparse heads reuse KV pairs from retrieval heads, reducing cache by 90% while maintaining quality through joint training.

adu2021 d622013 15.1 KB Updated

File contents

adu2021/skillxiv/tree/main/skills/skillxiv-v0.0.2-claude-opus-4.6/lychee-decode-sparse-kv-sharing commit d622013f8b

Frequently asked questions

npx skillmds@latest add adu2021/lychee-decode-sparse-kv-sharing