# Agent Evals

> Build repeatable evaluations for an agent from production failures, deterministic verifiers, cost, latency, and trace evidence.

- Skill: `getedgehq/agent-evals` (Agent Skill)
- Install (CLI): `npx skillmds@latest add getedgehq/agent-evals`
- Raw SKILL.md: https://api.skillmd.com/api/skills/getedgehq/agent-evals/raw
- Safety review: pending
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: AI & ML
- Author: getedgehq (https://skillmd.com/u/getedgehq)
- Updated: 2026-09-22
- Page: https://skillmd.com/skills/getedgehq/agent-evals

---


Turn representative failures into isolated tasks with objective checks and comparable runs.

