M2cqa Eval

Evaluates vision-language models' ability to correctly identify true statements about images while rejecting culturally plausible but visually incorrect counterfactual statements. It specifically probes grounding failures and cultural reasoning biases across multiple languages and dialects. Use when the user wants to benchmark on M²CQA, or asks about evaluating this task. Reports CFHR.

qhjqhj00 49bf2bd 3.4 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/m2cqa-eval commit 49bf2bde54

Frequently asked questions

npx skillmds add qhjqhj00/m2cqa-eval