More Than Quick Glance

Implement LASER-KV-style KV-cache compression for LLM inference pipelines using block-wise accumulative budgeting and hybrid exact-attention/LSH token selection. Use when: 'optimize KV cache for long context', 'compress KV cache without losing accuracy', 'implement LASER-KV', 'reduce LLM memory for 128k context', 'block-wise cache eviction strategy', 'fix attention-score greedy bias in cache pruning'.

ndpvt-web 1c8686d 16.7 KB Updated

File contents

ndpvt-web/arxiv-claude-skills/tree/main/skills/more-than-quick-glance commit 1c8686d019

Frequently asked questions

npx skillmds@latest add ndpvt-web/more-than-quick-glance