Mascqa Eval

Evaluates large language models' domain-specific reasoning and numerical problem-solving capabilities in materials science and metallurgical engineering. It probes their ability to accurately answer multiple-choice, matching, and numerical questions, highlighting gaps in scientific reasoning and computational precision. Use when the user wants to benchmark on MaScQA, or asks about evaluating this task. Reports accuracy.

qhjqhj00 1797234 3.1 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/mascqa-eval commit 1797234137

Frequently asked questions

npx skillmds add qhjqhj00/mascqa-eval