Padosoft AI Evaluation

Use this skill when measuring whether a model-backed feature works — building or changing an evaluation harness, a golden dataset, a metric, a scoring report, a regression gate on prompt or model changes. Also when the user asks how to know a prompt change made things better, how to stop a model upgrade silently regressing, what to put in a test set, how to score free-form output, or wants red-team coverage. It covers datasets and reports as versioned artifacts, the isolation an evaluation run needs, cohorts, and what must never leak into a report. Do not use it to write application tests (padosoft-test-integrity), to choose a model, or to design prompts.

padosoft Updated

File contents

padosoft/skills/tree/main/skills/padosoft-ai-evaluation commit c0af0da715

Frequently asked questions

npx skillmds@latest add padosoft/padosoft-ai-evaluation