Prove whether a prompt or model variant really won before shipping with promptstats

Run statistically sound comparisons on eval results so prompt and model changes are judged by confidence bounds, not bar-chart vibes.

agentskillexchange Updated 28 repo stars

File contents

agentskillexchange/skills/tree/main/skills/prove-whether-a-prompt-or-model-variant-really-won-before-shipping-with-promptstats commit 3e99f251cf

Frequently asked questions

npx skillmds@latest add agentskillexchange/prove-whether-a-prompt-or-model-variant-really-won-before-sh