Skill Quality Assurance
Run 6 quality checks on a skill and produce an actionable improvement report.
This is the quality gate before publishing — and also useful mid-creation to catch design issues early. The "Top Improvements" section at the end maps directly to what skill-creator should fix next, making this the natural handoff between authoring and shipping.
Input
If the skill path wasn't provided, ask for it — you need the root directory containing SKILL.md.
Read all skill files before dispatching agents:
SKILL.md(required)agents/*.md(if present)references/*.md(if present)scripts/*(if present)
If a directory doesn't exist, note its absence rather than skipping the check.
Execution
Run checks 1–5 in parallel. Run check 6 after — it depends on understanding the skill's promises first.
| # | Check | Agent | What it examines |
|---|---|---|---|
| 1 | Usefulness | agents/usefulness-checker.md |
Does this skill earn its place? |
| 2 | Authoring Principles | agents/authoring-checker.md |
description ≤250 chars? Standing Mandates? Compaction-aware? effort field? |
| 3 | Agent Structure | agents/structure-reviewer.md |
Should responsibilities split into subagents? Which steps can parallelize? |
| 4 | MCP Fit | agents/mcp-advisor.md |
Which MCPs would genuinely help? (load references/mcp-catalog.md) |
| 5 | SKILL.md Weight | agents/weight-analyzer.md |
Too heavy? What belongs in references/ or scripts/? |
| 6 | Output Quality | agents/eval-agent.md |
Does the skill measurably improve output vs no-skill baseline? |
Pass the full content of all skill files to each agent as context.
Output Format
Produce a self-contained report in this format:
## Skill QA Report: [skill-name]
### 1. Usefulness — [PASS / WARN / FAIL]
[findings]
### 2. Authoring Principles — [PASS / WARN / FAIL]
[findings]
### 3. Agent Structure — [GOOD / IMPROVABLE / MISSING]
**Persona separation:** [findings]
**Parallel opportunities:** [findings]
### 4. MCP Opportunities — [NONE / OPTIONAL / RECOMMENDED]
[findings]
### 5. SKILL.md Weight — [LIGHT / OK / HEAVY / CRITICAL]
[findings]
### 6. Output Quality — [PASS / MARGINAL / FAIL]
**With-skill pass rate:** X/N assertions
**Without-skill pass rate:** X/N assertions
**Delta:** +X
**Discriminating assertions:** [what the skill enforces that baseline misses]
**Skill gaps:** [what the skill promises but doesn't deliver]
---
## Top Improvements for [skill-name]
1. 🔴 [must fix — concrete and actionable]
2. 🟡 [recommended improvement]
3. 🟢 [optional enhancement]
Priority labels: 🔴 Must fix · 🟡 Recommended · 🟢 Optional
The "Top Improvements" section is what skill-creator reads to decide what to fix next. Make it concrete and actionable — not "improve structure" but "extract the grading logic into agents/grader.md and call it from SKILL.md with a Task()".