Evaluating They Not Know

Build statistically efficient LLM evaluation pipelines that combine direct accuracy with pairwise comparison signals as control variates. Use when the user asks to 'evaluate LLM accuracy on a benchmark', 'rank models with small sample sizes', 'reduce variance in LLM evaluation', 'build a model comparison pipeline', 'get tighter confidence intervals for model performance', or 'statistically compare reasoning models'.

ndpvt-web fdc7144 12.7 KB Updated

File contents

ndpvt-web/arxiv-claude-skills/tree/main/skills/evaluating-they-not-know commit fdc71441a1

Frequently asked questions

npx skillmds@latest add ndpvt-web/evaluating-they-not-know