Agent Eval

Build an evaluation harness for an LLM agent, prompt template, or tool-using workflow — a task set, deterministic and judge-based scoring, per-tag metrics, and a regression gate — then run it. Use it when a prompt or model change needs to be measured instead of eyeballed, or when a project ships LLM behavior with no eval set at all.

timurgaleev e4d22f9 26.7 KB Updated

File contents

timurgaleev/vibestack/tree/main/skills/agent-eval commit e4d22f90fa

Frequently asked questions

npx skillmds@latest add timurgaleev/agent-eval