Prismm Bench Eval

Evaluates large multimodal models' ability to detect, correct, and reason over real-world multimodal inconsistencies in scientific papers. It probes inter-modal mismatch detection, structured reasoning, and robustness to linguistic shortcuts versus genuine visual grounding. Use when the user wants to benchmark on PRISMM-Bench, or asks about evaluating this task. Reports Accuracy (%).

qhjqhj00 aa38e9b 3.5 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/prismm-bench-eval commit aa38e9bbb0

Frequently asked questions

npx skillmds add qhjqhj00/prismm-bench-eval