Skill Gauntlet

Regression-guarded blind A/B evaluation for skills, prompts, and agent configurations. Use when a candidate revision of a skill (or system prompt, rubric, or agent config) must prove it beats the current champion WITHOUT regressing guard dimensions before adoption — champion/challenger testing, pre-registered conjunctive adoption gates, blind persona-diverse judge panels, cross-check auditors, planted-flaw key verification, framing fixtures, negative fixtures, and generic controls. Triggers: "A/B test this skill", "extreme testing", "make sure we're not regressing", "test before adopting", "compare these two prompts/skills rigorously", "champion vs challenger", evolving an installed skill where a silent regression would be costly. Do NOT use for: simple factual QA, one-shot content generation, or evals where a single automated metric fully defines quality (use a test suite instead).

kimasplund dec914e 4 files · 14.9 KB Updated

File contents

kimasplund/skill-gauntlet/tree/main/ commit dec914e9c7

Frequently asked questions

npx skillmds@latest add kimasplund/skill-gauntlet