Noisy But Valid Robust

Statistically certify LLM safety/quality using imperfect LLM judges with guaranteed Type-I error control. Implements the "Noisy but Valid" hypothesis testing framework: calibrate a judge's TPR/FPR on a small human-labeled set, then run a variance-corrected test on a large judge-labeled dataset. Use when: "certify my model's failure rate", "validate LLM safety with an LLM judge", "statistical test with noisy labels", "is my model below the safety threshold", "evaluate LLM with imperfect judge", "calibrate judge accuracy and run hypothesis test".

ndpvt-web Updated

File contents

ndpvt-web/arxiv-claude-skills/tree/main/skills/noisy-but-valid-robust commit 2375681a21

Frequently asked questions

npx skillmds@latest add ndpvt-web/noisy-but-valid-robust