Omnidpo Hallucination Eval

This evaluation protocol assesses the ability of omni-modal large language models to avoid hallucinating non-existent audio or visual content when presented with contradictory or incomplete multimodal inputs. It specifically probes whether models can correctly identify real elements while resisting false affirmations of missing ones across text, vision, and audio modalities. Use when the user wants to benchmark on AVHBench, CMM, or asks about evaluating this task. Reports F1 Score, Hallucination Resistance (HR).

qhjqhj00 4d064fc 4.5 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/omnidpo-hallucination-eval commit 4d064fc458

Frequently asked questions

npx skillmds add qhjqhj00/omnidpo-hallucination-eval