Long Cot Reasoning Eval

This evaluation probes a model's ability to perform complex, multi-step reasoning across mathematics, coding, and scientific domains. It specifically measures the capacity to generate long chain-of-thought traces and produce correct final answers or executable code under strict generation constraints. Use when the user wants to benchmark on AIME24, AIME25, GPQA Diamond, LiveCodeBench v5, LiveCodeBench v6, or asks about evaluating this task. Reports average accuracy.

qhjqhj00 9c02922 3.2 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/long-cot-reasoning-eval commit 9c02922109

Frequently asked questions

npx skillmds add qhjqhj00/long-cot-reasoning-eval