Omnivideobench Eval

Evaluates multimodal large language models' ability to jointly reason across visual and audio modalities in long-duration videos. It probes capabilities like cross-modal alignment, temporal dependency modeling, and understanding of low-semantic acoustic cues such as music and ambient sounds. Use when the user wants to benchmark on OmniVideoBench, or asks about evaluating this task. Reports accuracy.

qhjqhj00 a261d73 3.1 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/omnivideobench-eval commit a261d739ec

Frequently asked questions

npx skillmds add qhjqhj00/omnivideobench-eval