Brace Hallucination Eval

Probes models' robustness in detecting subtle hallucinations in audio captions, specifically those introduced via LLM-driven noun substitution. It measures the ability to identify semantically flawed or factually incorrect descriptions against audio ground truth. Use when the user wants to benchmark on BRACE-Hallucination, or asks about evaluating this task. Reports F1-score.

qhjqhj00 2cd7643 2.6 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/brace-hallucination-eval commit 2cd7643ba4

Frequently asked questions

npx skillmds add qhjqhj00/brace-hallucination-eval