Mm Judgebench Eval

Evaluates the cross-lingual generalization and robustness of Large Vision-Language Models (LVLMs) acting as automated judges. It probes their ability to correctly rank paired multimodal responses across 25 languages while measuring susceptibility to positional and length biases. Use when the user wants to benchmark on MM-JudgeBench, or asks about evaluating this task. Reports average accuracy.

qhjqhj00 cfb2351 3.4 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/mm-judgebench-eval commit cfb23519d2

Frequently asked questions

npx skillmds add qhjqhj00/mm-judgebench-eval