Agent Evaluation Harness

Design an evaluation harness for an LLM agent before shipping it. Use when the user is building or rewriting an agent, deciding ship/no-ship, debugging regressions, or mentions golden sets, eval suites, regression tests, trace-level evals, LLM-as-judge, scoring rubrics, or asks "how do I test this agent?" / "how do I know if my agent got better?".

cobusgreyling d3a39fd 3 files · 9.0 KB Updated

File contents

cobusgreyling/agent-skills/tree/main/skills/agent-evaluation-harness commit d3a39fd783

Frequently asked questions

npx skillmds@latest add cobusgreyling/agent-evaluation-harness