Avere Emotion Reasoning Eval

This benchmark probes multimodal large language models' ability to reason about emotions from audio and video inputs while avoiding spurious cue associations and hallucinations. It specifically tests whether models can correctly align relevant audiovisual cues with emotional labels and resist over-reliance on textual priors or irrelevant modalities. Use when the user wants to benchmark on EmoReAlM, DFEW, RAVDESS, MER2023, EMER, or asks about evaluating this task. Reports average accuracy.

qhjqhj00 482d6d8 5.0 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/avere-emotion-reasoning-eval commit 482d6d8a85

Frequently asked questions

npx skillmds add qhjqhj00/avere-emotion-reasoning-eval