Lm Eval Harness Benchmarks Eval

Evaluates generative language models on a suite of multiple-choice and open-ended benchmarks covering reasoning, commonsense, multitask proficiency, and truthfulness. It measures accuracy across diverse domains to assess generalization and the impact of data combination strategies. Use when the user wants to benchmark on AI2 Reasoning Challenge (ARC), HellaSwag, MMLU, TruthfulQA, BigBench, HumanEval, or asks about evaluating this task. Reports accuracy.

qhjqhj00 7de76c3 3.2 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/lm-eval-harness-benchmarks-eval commit 7de76c3ed9

Frequently asked questions

npx skillmds add qhjqhj00/lm-eval-harness-benchmarks-eval