Reasoning Accuracy Eval

Evaluates large language models' logical reasoning and problem-solving capabilities across mathematical, algorithmic, and creative tasks. It measures both the correctness of final answers and the computational efficiency of the reasoning process. Use when the user wants to benchmark on Game of 24, BIG-Bench (subset), Python Puzzles, MGSM, Shakespearean Sonnet Writing, or asks about evaluating this task. Reports Acc_logic.

qhjqhj00 3693ebf 3.6 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/reasoning-accuracy-eval commit 3693ebffe9

Frequently asked questions

npx skillmds add qhjqhj00/reasoning-accuracy-eval