Egoavu Bench Eval

Evaluates multimodal large language models' ability to perform joint audio-visual reasoning on egocentric videos, including action/object/sound recognition, temporal reasoning, hallucination detection, and dense audio-visual narration. It specifically probes whether models can correctly associate environmental sounds with their visual sources and maintain temporal alignment without relying heavily on visual cues. Use when the user wants to benchmark on EgoAVU-Bench, or asks about evaluating this task. Reports Accuracy.

qhjqhj00 9e3f1a4 4.3 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/egoavu-bench-eval commit 9e3f1a4474

Frequently asked questions

npx skillmds add qhjqhj00/egoavu-bench-eval