Omnimodal QA Eval

This benchmark probes an agent's ability to perform long-horizon question answering over audio-video streams. It specifically tests fine-grained multimodal retrieval, multi-turn tool calling, and budget-aware reasoning when key evidence is scattered across time. Use when the user wants to benchmark on OmniVideoBench, WorldSense, Daily-Omni, or asks about evaluating this task. Reports accuracy.

qhjqhj00 d7fff70 2.5 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/omnimodal-qa-eval commit d7fff705a0

Frequently asked questions

npx skillmds add qhjqhj00/omnimodal-qa-eval