Llama Berry Eval

Evaluates LLMs on complex mathematical reasoning using search-based inference (SR-MCTS) rather than direct generation. It measures success rates across varying difficulty levels, from grade-school math to Olympiad-level problems, by testing both majority-vote and best-of-k strategies. Use when the user wants to benchmark on AIME24, AMC23, Math Odyssey, GPQA Diamond, OlympiadBench, College Math, MMLU STEM, GSM8K, GSMHard, MATH500, or asks about evaluating this task. Reports major@k.

qhjqhj00 17cfa2b 3.5 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/llama-berry-eval commit 17cfa2b2ed

Frequently asked questions

npx skillmds add qhjqhj00/llama-berry-eval