Av Speakerbench Eval

This benchmark probes fine-grained audiovisual reasoning in multimodal large language models, specifically requiring them to jointly determine who is speaking, what is being said, and when events occur within real-world video clips. It evaluates cross-modal fusion, temporal grounding, and speaker-centric perception through multiple-choice questions validated by human experts. Use when the user wants to benchmark on AV-SpeakerBench, or asks about evaluating this task. Reports MCQ accuracy.

qhjqhj00 c921a3a 2.9 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/av-speakerbench-eval commit c921a3a085

Frequently asked questions

npx skillmds add qhjqhj00/av-speakerbench-eval