Benchmark Judge Run

Use when a benchmark execution record or variant run needs a canonical judged record with route-first classification, five-dimension rubric scoring, post-score failure tags, separate invariance or robustness notes, and optional scaffold-to-completed finalization.

LoogacyStudio 8bd2080 4 files · 13.5 KB Updated

File contents

LoogacyStudio/skills/tree/main/.github/skills/benchmark-judge-run commit 8bd2080fb1

Frequently asked questions

npx skillmds@latest add loogacystudio/benchmark-judge-run