Agent Eval Cases

Decide which AI agent behaviors are worth an eval case, then write those cases — harness-, framework-, and language-agnostic. Use when writing a first eval suite for an agent, adding cases to an existing one, reviewing eval cases or scorers someone else wrote, choosing between a deterministic check and an LLM judge, deciding how many times to repeat a case, or reading a red run and working out whether the agent or the grader is wrong. Also use when a tool or prompt change "needs an eval" and it is not clear what to actually test, or when an agent misbehaves in a way no unit test can catch. Covers cases, tasks, scorers, graders, evaluators, judges, and pass@k. Does not build eval harnesses — it detects one and asks before anything gets built.

agentailor Updated

File contents

agentailor/skills/tree/main/agent-eval-cases commit 508941ae68

Frequently asked questions

npx skillmds@latest add agentailor/agent-eval-cases