Agent Evaluation Designer

Use this skill whenever the user wants to evaluate, test, or validate an AI agent, decide whether an agent is ready to ship or go live, choose how to grade an agent's answers (exact match, similarity, meaning, keywords, quality, or custom), design a test set of questions and expected answers, or interpret evaluation results into a go/no-go decision. Invoke it before the user hand-builds tests or declares an agent "done."

kody-w abc90e6 3 files · 60.0 KB Updated

File contents

kody-w/rapp-skills-legacy/tree/main/cat-agent-skills/agent-evaluation-designer commit abc90e6d43

Frequently asked questions

npx skillmds@latest add kody-w/agent-evaluation-designer