Dclm Benchmark Eval

Evaluates the effectiveness of data curation strategies for language models by training base models on curated corpora and measuring performance on 53 downstream tasks. It isolates data quality effects from architectural and computational variables using fixed training recipes across multiple compute scales. Use when the user wants to benchmark on DCLM downstream tasks, or asks about evaluating this task. Reports MMLU 5-shot accuracy.

qhjqhj00 a92d8a0 3.9 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/dclm-benchmark-eval commit a92d8a049c

Frequently asked questions

npx skillmds add qhjqhj00/dclm-benchmark-eval