Mosaicml Eval

Evaluates downstream language model capabilities across 33 question-answering tasks. It measures how effectively data pruning strategies improve general performance compared to unpruned baselines, using a normalized accuracy metric that accounts for random guessing baselines. Use when the user wants to benchmark on MosaicML evaluation gauntlet, or asks about evaluating this task. Reports average normalized accuracy.

qhjqhj00 228f635 3.0 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/mosaicml-eval commit 228f635e4f

Frequently asked questions

npx skillmds add qhjqhj00/mosaicml-eval