Flashvlm Eval

Evaluates the robustness and efficiency of text-guided visual token pruning in large multimodal models. It probes whether aggressive token compression (retaining 32–128 tokens for images, 114–455 for video) degrades performance on image and video question-answering tasks, and measures cross-modal grounding quality via spatial alignment and semantic overlap metrics. Use when the user wants to benchmark on VQAv2, GQA, VizWiz, ScienceQA-IMG, TextVQA, POPE, MME, MMBench, MMBench-CN, MM Vet, TGIF-QA, MSVDQA, MSRVTT-QA, ActivityNet-QA, or asks about evaluating this task. Reports accuracy, average_accuracy.

qhjqhj00 481b5fc 4.5 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/flashvlm-eval commit 481b5fcae0

Frequently asked questions

npx skillmds add qhjqhj00/flashvlm-eval