Omibench Eval

Evaluates large vision-language models on Olympiad-level multi-image reasoning tasks across biology, chemistry, mathematics, and physics. It probes the model's ability to integrate complementary visual and textual evidence across multiple images to generate stepwise rationales and select or produce correct final answers. Use when the user wants to benchmark on OMIBench, or asks about evaluating this task. Reports accuracy.

qhjqhj00 ec9502f 2.7 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/omibench-eval commit ec9502f0cc

Frequently asked questions

npx skillmds add qhjqhj00/omibench-eval