Halluaudio Eval

This benchmark probes the hallucination detection capabilities of Large Audio-Language Models (LALMs) across speech, environmental sound, and music domains. It systematically induces hallucinations using adversarial prompts and mixed-audio inputs to evaluate response correctness, affirmative bias, and refusal behavior beyond standard accuracy. Use when the user wants to benchmark on HalluAudio, or asks about evaluating this task. Reports Accuracy.

qhjqhj00 ff498aa 4.9 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/halluaudio-eval commit ff498aa27c

Frequently asked questions

npx skillmds add qhjqhj00/halluaudio-eval