Bbh Prompting Eval

Tests the impact of prompt structure and logical validity on language model reasoning performance. Specifically, it compares answer-only, standard chain-of-thought, and logically invalid chain-of-thought prompting strategies on complex reasoning tasks. Use when the user wants to benchmark on BIG-Bench Hard, or asks about evaluating this task. Reports accuracy.

qhjqhj00 0da5d93 2.6 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/bbh-prompting-eval commit 0da5d93036

Frequently asked questions

npx skillmds add qhjqhj00/bbh-prompting-eval