Agent Evaluation

Design and implement comprehensive evaluation systems for AI agents. Use when building evals for coding agents, conversational agents, research agents, or computer-use agents. Covers grader types, benchmarks, 8-step roadmap, and production integration.

autohandai Updated

File contents

autohandai/community-skills/tree/main/agent-evaluation commit 02234c6e3c

Frequently asked questions

npx skillmds@latest add autohandai-community-skills/agent-evaluation