Cg Bench Eval

Evaluates multimodal large language models on clue-grounded audio-visual counting tasks over long videos. It probes the model's ability to integrate audio and visual cues to locate temporal segments and accurately count events, objects, or attributes within those segments. Use when the user wants to benchmark on CG-Bench, or asks about evaluating this task. Reports counting_accuracy.

qhjqhj00 94a7f20 2.9 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/cg-bench-eval commit 94a7f209f6

Frequently asked questions

npx skillmds add qhjqhj00/cg-bench-eval