Art Audio Reasoning Eval

This benchmark evaluates multimodal large language models on their ability to perform cross-modal audio reasoning. It requires models to integrate multiple audio cues (e.g., speech, environmental sounds, speaker identity) and apply logical inference to answer questions, rather than just performing isolated audio tasks like transcription or classification. Use when the user wants to benchmark on ART (Audio Reasoning Tasks), or asks about evaluating this task. Reports Absolute accuracy.

qhjqhj00 c37f781 3.5 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/art-audio-reasoning-eval commit c37f7814b6

Frequently asked questions

npx skillmds add qhjqhj00/art-audio-reasoning-eval