Agent Eval

Run head-to-head comparisons of coding agents (Claude Code, Aider, Codex, etc.) on custom tasks, reporting pass rate, cost, time, and consistency metrics. USE WHEN choosing between coding agents or benchmarking agent performance on representative tasks.

Sheshiyer 8c3e6f8 4.5 KB Updated

File contents

Sheshiyer/skill-clusters/tree/main/skills/agent-eval commit 8c3e6f8faf

Frequently asked questions

npx skillmds@latest add sheshiyer/agent-eval