Atlas Eval

This benchmark probes frontier scientific reasoning across multiple disciplines (e.g., physics, chemistry, biology, computer science, mathematics) using original, multi-step problems. It evaluates a model's ability to generate complex, open-ended, LaTeX-formatted answers and assesses both solution accuracy and inference stability across multiple sampling runs. Use when the user wants to benchmark on ATLAS, or asks about evaluating this task. Reports Accuracy.

qhjqhj00 dac6b69 3.2 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/atlas-eval commit dac6b69200

Frequently asked questions

npx skillmds add qhjqhj00/atlas-eval