Neural Cot Search Eval

Evaluates large language models' ability to perform multi-step reasoning across diverse domains including mathematics, commonsense, and expert knowledge. It specifically probes the model's capacity to generate accurate solutions while minimizing computational cost (token usage) through dynamic reasoning path search. Use when the user wants to benchmark on AMC23, ARC-C, GPQA, GSM8K, or asks about evaluating this task. Reports Efficiency Metric ($\eta$).

qhjqhj00 e47919c 4.3 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/neural-cot-search-eval commit e47919c766

Frequently asked questions

npx skillmds add qhjqhj00/neural-cot-search-eval