Extra Cot Eval

Evaluates the ability of large language models to generate mathematically reasoned chain-of-thought outputs that are compressed to a target token budget while preserving logical fidelity and answer accuracy. Use when the user wants to benchmark on GSM8K, MATH-500, AMC2023, or asks about evaluating this task. Reports Acc@all.

qhjqhj00 89b8378 3.5 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/extra-cot-eval commit 89b8378a49

Frequently asked questions

npx skillmds add qhjqhj00/extra-cot-eval