Skywork Benchmark Eval

Evaluates bilingual foundation models on general knowledge, Chinese domain-specific reasoning, mathematical problem-solving, and language modeling capabilities using standardized benchmarks and custom held-out text corpora. Use when the user wants to benchmark on MMLU, CEVAL, CMMLU, GSM8K, Custom Chinese LM Testset, or asks about evaluating this task. Reports 5-shot accuracy.

qhjqhj00 7c3b3a6 3.4 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/skywork-benchmark-eval commit 7c3b3a6a7c

Frequently asked questions

npx skillmds add qhjqhj00/skywork-benchmark-eval