Clash Eval

This benchmark probes multimodal large language models' ability to detect cross-modal contradictions between images and text. It evaluates whether models can identify inconsistencies when either modality contains errors or hallucinations, rather than assuming one modality is ground truth. The task reveals systematic modality biases and category-specific reasoning weaknesses. Use when the user wants to benchmark on CLASH, or asks about evaluating this task. Reports accuracy.

qhjqhj00 0548b54 3.0 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/clash-eval commit 0548b54f86

Frequently asked questions

npx skillmds add qhjqhj00/clash-eval