Mme Sci Eval

Evaluates multimodal large language models on scientific reasoning across four disciplines (math, physics, chemistry, biology) and five languages. It probes cross-lingual consistency, modality robustness (text-only vs. image-only vs. image-text), and fine-grained domain knowledge under varying visual complexity. Use when the user wants to benchmark on MME-SCI, or asks about evaluating this task. Reports accuracy.

qhjqhj00 4ceb5ca 3.2 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/mme-sci-eval commit 4ceb5ca44f

Frequently asked questions

npx skillmds add qhjqhj00/mme-sci-eval