Longvideo Bench Eval

This benchmark evaluates long-context video-language understanding by testing a model's ability to retrieve specific moments from lengthy videos and reason over multimodal details. It distinguishes between single-moment visual perception and multi-moment relational reasoning across 17 fine-grained categories. Use when the user wants to benchmark on LongVideoBench, or asks about evaluating this task. Reports accuracy.

qhjqhj00 76a33ae 2.5 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/longvideo-bench-eval commit 76a33aec83

Frequently asked questions

npx skillmds add qhjqhj00/longvideo-bench-eval