Skill Eval Loop

Prove whether a candidate skill actually improves agent performance before adopting it - blind A/B evals on reconstructed real tasks, a three-tier adoption gate, and production rechecks. Use when asked to "eval this skill", "does this skill help", "test this skill before installing", "run the auto-improve loop", or "recheck probationary skills". Takes candidates from skill-miner; runs locally with the user's own agent CLI and models.

getedgehq Updated

File contents

getedgehq/skills/tree/main/skill-eval-loop commit 7e7fafeb10

Frequently asked questions

npx skillmds@latest add getedgehq/skill-eval-loop