Vbenchcomp Eval

Evaluates video language models by disentangling question types into LLM-Answerable, Semantic, Temporal, and Others. It isolates true temporal and spatial understanding from language priors and static visual cues by computing accuracy exclusively on the Semantic and Temporal subsets. Use when the user wants to benchmark on LongVideoBench, Egoschema, NextQA, VideoMME, MLVU, LVBench, PerceptionTest, or asks about evaluating this task. Reports VBenchComp score.

qhjqhj00 2be0800 3.7 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/vbenchcomp-eval commit 2be0800add

Frequently asked questions

npx skillmds add qhjqhj00/vbenchcomp-eval