Precise Reducing Bias Evaluations

Implement the PRECISE framework to debias LLM-as-judge evaluations of search, ranking, and RAG systems by combining a small human-annotated gold set with large-scale LLM judgments using Prediction-Powered Inference. Triggers: 'evaluate search quality with LLM judges', 'debias LLM relevance judgments', 'estimate Precision@K with fewer annotations', 'combine human and LLM labels for ranking evaluation', 'reduce annotation cost for retrieval evaluation', 'correct LLM judge bias in ranking metrics'

ndpvt-web f3d6f99 14.2 KB Updated

File contents

ndpvt-web/arxiv-claude-skills/tree/main/skills/precise-reducing-bias-evaluations commit f3d6f99303

Frequently asked questions

npx skillmds@latest add ndpvt-web/precise-reducing-bias-evaluations