Latent Reasoning Benchmarks Eval

Evaluates a language model's reasoning, coding, and general knowledge capabilities using a suite of standard academic benchmarks. It specifically probes how test-time compute scaling (via recurrent depth) impacts performance across mathematical, coding, and commonsense reasoning tasks. Use when the user wants to benchmark on GSM8K, MATH (Minerva), MathQA, MBPP, HumanEval, ARC-E, ARC-C, HellaSwag, MMLU, OBQA, PiQA, SciQ, WinoGrande, or asks about evaluating this task. Reports flexible extract accuracy.

qhjqhj00 7bdcacc 4.6 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/latent-reasoning-benchmarks-eval commit 7bdcacce31

Frequently asked questions

npx skillmds add qhjqhj00/latent-reasoning-benchmarks-eval