Subagent Eval

Empirically evaluate and optimise a Claude Code sub-agent via a blinded head-to-head against the production baseline, then pick its model on data. Use when tuning a sub-agent's prompt, choosing or re-checking its model (e.g. a new model release), proving an agent change before committing, or testing whether a specialist beats native Explore. To author a new agent first, read AUTHORING.md. Use when this capability is needed.

tomevault-io 19d88f8 2 files · 7.8 KB Updated

File contents

tomevault-io/skills-registry/tree/main/samjmarshall--rekurve--subagent-eval commit 19d88f8305

Frequently asked questions

npx skillmds@latest add tomevault-io/subagent-eval