Turn representative failures into isolated tasks with objective checks and comparable runs.
Agent Evals
Build repeatable evaluations for an agent from production failures, deterministic verifiers, cost, latency, and trace evidence.
Agent Evals by getedgehq · a8f0143
npx skillmds@latest add getedgehq/agent-evals File contents
---name: agent-evalsdescription: Build repeatable evaluations for an agent from production failures, deterministic verifiers, cost, latency, and trace evidence.---Turn representative failures into isolated tasks with objective checks and comparable runs.
getedgehq/skills/tree/main/evals/skillneed/fixtures/skillneed-catalog-routing/catalog/agent-evals commit a8f0143435
Frequently asked questions
Run npx skillmds@latest add getedgehq/agent-evals in your terminal (requires Node.js), paste this page's agent-chat prompt into Claude, Cursor, or any MCP-connected agent, or download the SKILL.md file and copy it into your agent's skills directory.
Build repeatable evaluations for an agent from production failures, deterministic verifiers, cost, latency, and trace evidence. It is listed under AI & ML on SkillMD.
This skill has not completed SkillMD's automated safety review yet. Capability flags: docs only. SkillMD never runs a skill's scripts for you; review the SKILL.md before installing.
This skill is tagged as working with Claude Code, Claude.ai, OpenAI Codex. SKILL.md is an open format, so most agents that read a skills directory can load it too.
Yes. Installing skills from SkillMD is free, and the skill stays under its author's original license.
getedgehq (@getedgehq) published this skill. Their other Agent Skills are listed on their SkillMD profile.