Visor Sparse Vision Language Interaction

Optimize vision-language model inference by sparsifying the interactions between vision and language tokens instead of compressing images. Uses a dynamic policy to allocate visual computation per sample based on complexity, enabling a universal network across different compute budgets. Maintains high-resolution reasoning when needed. Use when deploying VLMs under varying compute constraints, need per-sample efficiency adaptation, or want to preserve fine visual details while reducing compute.

adu2021 2ace466 10.3 KB Updated

File contents

adu2021/skillxiv/tree/main/skills/skillxiv-v0.0.3-claude-opus-4.6/visor-sparse-vision-language-interaction commit 2ace4664e6

Frequently asked questions

npx skillmds@latest add adu2021/visor-sparse-vision-language-interaction