Agent Eval

Designs and runs evaluations for LLM or agent outputs — builds rubrics, sets up LLM-as-judge scoring, creates regression test sets, and reports pass rates with concrete failure examples. Use this skill whenever the user wants to evaluate, test, grade, or score an agent's or LLM's outputs; asks "how do I know if this is working," "is this any good," "set up an eval," or "did this prompt change make things worse"; needs a rubric for judging quality; wants to compare two prompts, models, or outputs; or wants to catch regressions before shipping a change. Also trigger when reviewing agent trajectories specifically — did the agent pick the right tool, the right arguments, the right sequence — not just the final output.

Hefrock a4d4f99 25 files · 167.5 KB Updated

File contents

Hefrock/agent-skills/tree/main/skills/agent-eval commit a4d4f9982f

Frequently asked questions

npx skillmds@latest add hefrock/agent-eval