Matscibench Eval

This benchmark evaluates the reasoning capabilities of large language models in materials science, covering six primary fields and 31 sub-fields. It probes domain knowledge, mathematical/formula reasoning, and multimodal visual comprehension through expert-curated problems with three-tier difficulty classifications. Use when the user wants to benchmark on MatSciBench, or asks about evaluating this task. Reports Accuracy Score (%).

qhjqhj00 44b6871 3.7 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/matscibench-eval commit 44b6871da0

Frequently asked questions

npx skillmds add qhjqhj00/matscibench-eval