Agent Eval Designer

Use when asked to design an evaluation for an AI agent or skill's real-world reliability — across repeated runs, under noisy tools and environments, against realistic tasks — rather than to verify a single change worked once. Distinct from a one-off QA pass; produces a repeatable eval suite with metrics, noise controls, and human-review gates for high-risk cases.

PresidenteOG Updated

File contents

PresidenteOG/claude-skills/tree/main/agent-eval-designer commit e367cabbd0

Frequently asked questions

npx skillmds@latest add presidenteog/agent-eval-designer