Reasoning Benchmarks Eval

Evaluates language model reasoning and general knowledge across multiple standard benchmarks. It measures the impact of data mixture optimization on model performance using accuracy scores on commonsense, scientific, and factual QA tasks. Use when the user wants to benchmark on PIQA, ARC_C, ARC_E, HellaSwag, WinoGrande, SIQA, MMLU, or asks about evaluating this task. Reports test accuracy.

qhjqhj00 66fcf93 2.7 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/reasoning-benchmarks-eval commit 66fcf937d6

Frequently asked questions

npx skillmds add qhjqhj00/reasoning-benchmarks-eval