Video Oasis Eval

This protocol audits video understanding benchmarks to measure genuine spatio-temporal reasoning versus shortcut reliance. It filters out samples solvable without video context and evaluates models under diagnostic conditions (e.g., blind, audio-only, center-frame) to quantify performance degradation and dependency on actual video content. Use when the user wants to benchmark on EgoSchema, ImplicitQA, VSI-Bench, TVBench, VCR-Bench, RTV-Bench, Video-Holmes, MINERVA, MMR-V, VideoMME, MVBench, LVBench, LongVideoBench, MLVU, or asks about evaluating this task. Reports accuracy.

qhjqhj00 579ef71 4.3 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/video-oasis-eval commit 579ef7108e

Frequently asked questions

npx skillmds add qhjqhj00/video-oasis-eval