Eval Skills

Eval and improve a skill against golden cases — run the target skill blind in a fresh, context-free subagent on each example input, grade the artifact against the expected outcome, and let the gaps drive the edits. Use when the user wants to test/eval/improve/harden a skill, says "this skill keeps producing X / keeps missing Y", or hands a skill plus example input→expected-output pairs. Pairs with [write-skills](../write-skills/SKILL.md) (the authoring principles every fix obeys).

dzhng 505e36e 7.5 KB Updated

File contents

dzhng/duet-agent/tree/main/.agents/skills/eval-skills commit 505e36edfb

Frequently asked questions

npx skillmds@latest add dzhng/eval-skills