Mm Inference Eval

This evaluation probes the effectiveness and efficiency of dynamic sparse attention methods for long-context vision-language models (VLMs). It tests the model's ability to perform long-video understanding, retrieve specific visual or mixed-modality information from extremely long contexts (Needle in a Haystack), and maintain accuracy while reducing computational cost and latency. Use when the user wants to benchmark on Video Understanding Benchmarks, V-NIAH, MM-NIAH, or asks about evaluating this task. Reports task accuracy / official benchmark score.

qhjqhj00 4f30cc3 4.0 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/mm-inference-eval commit 4f30cc3663

Frequently asked questions

npx skillmds add qhjqhj00/mm-inference-eval