Kaleidoscope Eval

Evaluates multilingual vision-language reasoning by testing models on multiple-choice questions about images entirely in their native language. It probes cultural and linguistic authenticity, assessing how well models handle complex multimodal reasoning without relying on English translations. Use when the user wants to benchmark on Kaleidoscope, or asks about evaluating this task. Reports accuracy.

qhjqhj00 52dcd1f 4.0 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/kaleidoscope-eval commit 52dcd1f970

Frequently asked questions

npx skillmds add qhjqhj00/kaleidoscope-eval