Csmbench Eval

Evaluates large multimodal models' ability to perceive, interpret, and reason about scientific figures across four hierarchical physical scales (atomic, micro, meso, macro) in materials science. It probes both discriminative visual matching and open-ended scientific narrative generation. Use when the user wants to benchmark on CSMBench, or asks about evaluating this task. Reports accuracy.

qhjqhj00 b97b42b 3.8 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/csmbench-eval commit b97b42be1f

Frequently asked questions

npx skillmds add qhjqhj00/csmbench-eval