Standard LLM Benchmarks Eval

Evaluates language model capabilities across knowledge, reasoning, instruction following, and safety using a standard suite of downstream benchmarks. It measures both pretraining quality and post-adaptation performance on established NLP and coding tasks. Use when the user wants to benchmark on MMLU, HellaSwag, ARC-Challenge, ARC-Easy, PIQA, WinoGrande, GSM8k, BBH, HumanEval, AlpacaEval 1.0, XSTest, IFEval, or asks about evaluating this task. Reports exact-match accuracy.

qhjqhj00 f9ce8ef 3.9 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/standard-llm-benchmarks-eval commit f9ce8ef736

Frequently asked questions

npx skillmds add qhjqhj00/standard-llm-benchmarks-eval