Olympiadbench Eval

This benchmark evaluates the scientific reasoning capabilities of large multimodal models (LMMs) and large language models (LLMs) on Olympiad-level mathematics and physics problems. It specifically probes bilingual (English and Chinese) text-and-image problem solving, computational correctness, and logical consistency in complex, expert-annotated scenarios. Use when the user wants to benchmark on OlympiadBench, or asks about evaluating this task. Reports micro-average accuracy.

qhjqhj00 37ee23a 3.5 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/olympiadbench-eval commit 37ee23a1b2

Frequently asked questions

npx skillmds add qhjqhj00/olympiadbench-eval