Helm Lite Eval

This evaluation probes the robustness of open benchmarks against test-set memorization and data leakage. It measures whether small language models can artificially inflate leaderboard scores by overfitting directly to public test sets, revealing flaws in current benchmarking practices. Use when the user wants to benchmark on HELM-lite, or asks about evaluating this task. Reports Exact Match.

qhjqhj00 b2e796d 3.0 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/helm-lite-eval commit b2e796de4d

Frequently asked questions

npx skillmds add qhjqhj00/helm-lite-eval