Multi Prompt Eval

This protocol evaluates how accurately a statistical estimation method can reconstruct the full performance distribution and specific quantiles of large language models across hundreds of prompt templates, using a fraction of the standard evaluation budget. It probes the robustness of LLM performance metrics against arbitrary prompt selection and measures the efficiency of borrowing strength across prompts and examples. Use when the user wants to benchmark on MMLU, BIG-bench Hard, LMentry, or asks about evaluating this task. Reports Wasserstein-1 distance ($W_1$).

qhjqhj00 753bd30 3.8 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/multi-prompt-eval commit 753bd30b34

Frequently asked questions

npx skillmds add qhjqhj00/multi-prompt-eval