Agent Evaluation

Designs and runs reproducible evaluations for AI agents, prompts, tools, skills, and model-backed workflows using realistic datasets, isolated baselines, objective assertions, rubric grading, trajectory analysis, cost/latency tracking, and regression comparison. Use when measuring agent quality, optimizing skill triggering, comparing prompts or models, or gating an AI feature release. Not for ordinary deterministic unit tests.

thiientv Updated

File contents

thiientv/godmode/tree/main/skills/agent-evaluation commit 788810d5c7

Frequently asked questions

npx skillmds@latest add thiientv/agent-evaluation