Kelly Agent Eval

Review board (Busabase App-in-Skill) that runs a fixed suite of mock test cases against a baseline vs candidate agent version and surfaces rubric-scored regressions before a release. Use when the user invokes $kelly-agent-eval or /kelly-agent-eval, wants to review agent-version regressions, compare baseline vs candidate quality, triage a release, or record a release approve/block decision. Deterministic mock rubric scores only — not a real LLM-judge call, and it never deploys anything.

mr-kelly Updated

File contents

mr-kelly/skills/tree/main/skills/kelly-agent-eval commit 47f3ae6251

Frequently asked questions

npx skillmds@latest add mr-kelly/kelly-agent-eval