Opencompass Downstream Eval

Evaluates language model performance across five diverse downstream benchmarks spanning commonsense reasoning, science QA, and complex reasoning. It specifically probes how dynamic expert routing mechanisms adapt to input difficulty compared to fixed Top-K routing. Use when the user wants to benchmark on PIQA, Hellaswag, ARC-e, CommonsenseQA, BBH, or asks about evaluating this task. Reports accuracy (score).

qhjqhj00 148e344 2.8 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/opencompass-downstream-eval commit 148e344f40

Frequently asked questions

npx skillmds add qhjqhj00/opencompass-downstream-eval