Vilbias Eval

This benchmark probes a model's ability to detect framing bias in multimodal news content (text-image pairs) and generate grounded, correct rationales for its decisions. It evaluates both closed-ended classification accuracy and open-ended reasoning quality using an LLM-as-judge protocol. Use when the user wants to benchmark on ViLBias, or asks about evaluating this task. Reports Accuracy.

qhjqhj00 573e085 4.3 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/vilbias-eval commit 573e085844

Frequently asked questions

npx skillmds add qhjqhj00/vilbias-eval