Omnigaiatoolreasoning Eval

Evaluates native omni-modal AI agents' ability to perform multi-hop cross-modal reasoning across video, audio, and image inputs. It probes their capacity to integrate external tools (web search, browser, code execution) for evidence gathering and to produce verifiable open-form answers under varying task difficulties. Use when the user wants to benchmark on OmniGAIA, or asks about evaluating this task. Reports Pass@1.

qhjqhj00 cf24bae 2.8 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/omnigaiatoolreasoning-eval commit cf24baeb06

Frequently asked questions

npx skillmds add qhjqhj00/omnigaiatoolreasoning-eval