Auto Dataset Update Eval

Evaluates LLMs on automatically updated benchmark datasets (BIG-bench, MMLU) to measure evaluation stability, data leakage mitigation, and cognitive-level difficulty control via mimicking and extending generation strategies. Use when the user wants to benchmark on BIG-bench, MMLU, or asks about evaluating this task. Reports full-mark rate (%).

qhjqhj00 5791811 2.5 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/auto-dataset-update-eval commit 579181150f

Frequently asked questions

npx skillmds add qhjqhj00/auto-dataset-update-eval