Agent Evaluation

Design evaluation systems for AI agents by choosing the right grader mix, benchmark shape, harness boundaries, and production feedback loop. Use when the user needs eval planning for coding agents, research agents, conversational agents, or computer-use agents, even if they ask in terms like benchmark, grader, harness, regression suite, eval roadmap, red-team tasks, online evals, or agent quality monitoring. Not for fixing the underlying product feature or writing the feature tests themselves.

akillness a0cb563 5 files · 16.1 KB Updated 42 repo stars

File contents

akillness/oh-my-gods/tree/main/.god-skills/agent-evaluation commit a0cb5632ff

Frequently asked questions

npx skillmds@latest add akillness/agent-evaluation