Cot Eval

Evaluates the zero-shot and few-shot reasoning capabilities of language models, specifically probing their ability to generate step-by-step chain-of-thought rationales and produce correct answers across classification and generation tasks. Use when the user wants to benchmark on BigBench Hard (BBH), P3, MGSM, or asks about evaluating this task. Reports accuracy.

qhjqhj00 2418d0d 4.6 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/cot-eval commit 2418d0dd92

Frequently asked questions

npx skillmds add qhjqhj00/cot-eval