Chinese LLM Benchmarks Eval

Evaluates the knowledge, reasoning, instruction-following, and safety alignment capabilities of Chinese instruction-tuned LLMs across academic, professional, open-ended, and safety-critical domains. Use when the user wants to benchmark on C-Eval, CMMLU, BELLE-EVAL, SafetyBench, or asks about evaluating this task. Reports log-likelihood.

qhjqhj00 f325011 2.8 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/chinese-llm-benchmarks-eval commit f3250113a1

Frequently asked questions

npx skillmds add qhjqhj00/chinese-llm-benchmarks-eval