Agent Eval

Reproducible evaluation harness for coding agents, prompts, and skills — head-to-head comparison with pass rate, cost, time, and consistency metrics captured in git worktrees. Use when comparing coding agents (Claude Code, Aider, Codex), regression-testing your own skills after changes, measuring variance across repeated runs of the same prompt, validating a model upgrade before adopting it, or A/B testing two versions of a prompt. Trigger on "compare Claude Code vs aider", "is my skill still passing", "run regression evals on this skill", "how much variance does this prompt have", "did the model upgrade regress anything". Use when this capability is needed.

tomevault-io Updated

File contents

tomevault-io/skills-registry/tree/main/dakaneye--claude-skills--agent-eval commit bea6fe16eb

Frequently asked questions

npx skillmds@latest add tomevault-io/agent-eval-5