Jamendo Mt QA Eval

Evaluates audio-language models on multi-track comparative reasoning by asking them to compare two music tracks and answer questions. It probes the model's ability to perform grounded, sentence-level comparative explanations versus simple binary or short-answer discrimination. Use when the user wants to benchmark on Jamendo-MT-QA, or asks about evaluating this task. Reports accuracy, LLM-as-a-Judge score.

qhjqhj00 08ad0f7 3.5 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/jamendo-mt-qa-eval commit 08ad0f7183

Frequently asked questions

npx skillmds add qhjqhj00/jamendo-mt-qa-eval