Nasde Benchmark Calibration

Calibrate assessment rubrics by reviewing agent work in GitHub/GitLab PRs and feeding human comments back into the rubric. Use this skill when the user wants to: - Calibrate, tune, or sanity-check assessment criteria / dimensions of a benchmark - Review trial diffs alongside the LLM-as-a-Judge scores in a PR/MR - Investigate why judge scores feel off, too harsh, too lenient, or misaligned with how a human would grade the code - Pull review comments back from PRs/MRs and turn them into concrete rubric edits Even if the user doesn't say "calibrate" — if they're worried the LLM judge's scores diverge from human judgment, or want to align scores with a real developer's opinion before freezing a benchmark, this skill applies.

NoesisVision 77cadcb 7.6 KB Updated

File contents

NoesisVision/nasde-toolkit/tree/main/.claude/skills/nasde-benchmark-calibration commit 77cadcbbc0

Frequently asked questions

npx skillmds@latest add noesisvision/nasde-benchmark-calibration