Eval Skills

Eval and improve a skill against golden cases — run the target skill blind in a fresh, context-free subagent on each example input, grade the artifact against a per-case bar, and let the gaps drive the edits. Use when the user wants to test/eval/improve/harden a skill, says "this skill keeps producing X / keeps missing Y", or hands a skill plus cases with a bar for what good looks like. Pairs with write-skills (the authoring principles every fix obeys). Judgment-first and harness-free — prefer skill-creator for harness-based benchmarking, variance analysis, or description-optimization loops.

avivsinai Updated

File contents

avivsinai/skills-marketplace/tree/main/plugins/skill-authoring/skills/eval-skills commit 970b8347a1

Frequently asked questions

npx skillmds@latest add avivsinai/eval-skills