Sketch Of Thought Eval

Evaluates the reasoning efficiency and accuracy of LLMs under cognitive-inspired prompting constraints. It probes the model's ability to produce structured, concise reasoning chains while maintaining correctness across mathematical, commonsense, logical, multi-hop, scientific, medical, multilingual, and multimodal tasks. Use when the user wants to benchmark on GSM8K, SVAMP, AQUA-RAT, DROP, CommonsenseQA, OpenbookQA, StrategyQA, LogiQA, ReClor, HotPotQA, MuSiQue-Ans, QASC, Worldtree, PubMedQA, MedQA, MMLU, MMMLU, GQA, ScienceQA, or asks about evaluating this task. Reports accuracy.

qhjqhj00 a0e45f4 5.0 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/sketch-of-thought-eval commit a0e45f4d8a

Frequently asked questions

npx skillmds add qhjqhj00/sketch-of-thought-eval