Mistake Finding Eval

Evaluates an LLM's ability to detect and locate the first logical error in a multi-step chain-of-thought reasoning trace. It probes whether models can accurately identify specific reasoning steps that contain mistakes or correctly assert that a trace is entirely correct. Use when the user wants to benchmark on BIG-Bench Mistake, or asks about evaluating this task. Reports accuracy.

qhjqhj00 5d59a73 3.2 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/mistake-finding-eval commit 5d59a7360c

Frequently asked questions

npx skillmds add qhjqhj00/mistake-finding-eval