Biasinear Eval

Evaluates the robustness and sensitivity of multimodal large language models (MLLMs) to perturbations in spoken multiple-choice questions. It probes how models handle variations in language, accent, speaker gender, and answer option ordering, measuring both absolute correctness and prediction stability across conditions. Use when the user wants to benchmark on BiasInEar, or asks about evaluating this task. Reports Question Entropy.

qhjqhj00 ba35a04 3.9 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/biasinear-eval commit ba35a04106

Frequently asked questions

npx skillmds add qhjqhj00/biasinear-eval