Eval Integrity

Audit an LLM evaluation or benchmark repo for integrity and credibility practices. Use when asked to "audit my benchmark," "is my eval trustworthy," "check my leaderboard for contamination," "review this benchmark's methodology," or "what would a reviewer attack in my eval." Greps the target repo for evidence across seven dimensions (pre-registration, contamination, holdout hygiene, judge validity, statistical honesty, reproducibility, leaderboard exclusions) and emits a scored report with file:line evidence, severity, and concrete fixes.

conorbronsdon 5f7a9aa 4 files · 268.4 KB Updated

File contents

conorbronsdon/agent-skills/tree/main/eval-integrity commit 5f7a9aa852

Frequently asked questions

npx skillmds@latest add conorbronsdon/eval-integrity