Alm Bench Eval

This benchmark evaluates the cultural and linguistic reasoning capabilities of large multimodal models across 100 languages. It probes visual understanding and cultural knowledge through generic and culturally specific domains, testing both closed-form and open-ended question answering. Use when the user wants to benchmark on ALM-bench, or asks about evaluating this task. Reports accuracy.

qhjqhj00 5e3c17e 2.6 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/alm-bench-eval commit 5e3c17e7da

Frequently asked questions

npx skillmds add qhjqhj00/alm-bench-eval