Agent Eval

Run head-to-head comparisons of coding agents (Claude Code, Aider, Codex, etc.) on custom tasks, reporting pass rate, cost, time, and consistency metrics. USE WHEN choosing between coding agents or benchmarking agent performance on representative tasks. Use when this capability is needed.

tomevault-io Updated

File contents

tomevault-io/skills-registry/tree/main/sheshiyer--skill-clusters--agent-eval commit fcbd0ca594

Frequently asked questions

npx skillmds@latest add tomevault-io/agent-eval-8