H2vu Benchmark Eval

Evaluates multimodal large language models on hierarchical and holistic video understanding, specifically probing temporal reasoning, countercommonsense comprehension, trajectory state tracking, and first-person streaming video analysis. Use when the user wants to benchmark on H²VU, or asks about evaluating this task. Reports accuracy.

qhjqhj00 8294905 2.5 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/h2vu-benchmark-eval commit 8294905576

Frequently asked questions

npx skillmds add qhjqhj00/h2vu-benchmark-eval