Agent Eval

Head-to-head comparison of coding agents (Claude Code, Aider, Codex, etc.) on custom tasks with pass rate, cost, time, and consistency metrics. Use when choosing between coding agents, or when a change to an agent setup needs measured pass rate, cost, and time rather than an impression.

JunMystery c082704 4.5 KB Updated

File contents

JunMystery/Agent-Guidance-Rust/tree/main/skills/agent-eval commit c082704f19

Frequently asked questions

npx skillmds@latest add junmystery/agent-eval