Agent Eval

Use when measuring whether an LLM or agent system actually got better and gating merges on it: golden sets, fixing an inflated LLM-as-judge, scoring RAG (faithfulness, contextual recall) or agent trajectories (tool correctness, completion), or picking an eval framework. NOT building the agent loop, tools or RAG plumbing (that is `building-agents`).

ericrisco 0e9308a 6 files · 33.3 KB Updated

File contents

ericrisco/rsc-harness/tree/main/skills/agent-eval commit 0e9308a126

Frequently asked questions

npx skillmds@latest add ericrisco/agent-eval