Skill Evaluator
Evaluate skills across 25 criteria using a hybrid automated + manual approach.
Quick Start
1. Run automated checks
python3 scripts/eval-skill.py /path/to/skill
python3 scripts/eval-skill.py /path/to/skill --json # machine-readable
python3 scripts/eval-skill.py /path/to/skill --verbose # show all details
Checks: file structure, frontmatter, description quality, script syntax, dependency audit, credential scan, env var documentation.
2. Manual assessment
Use the rubric at references/rubric.md to score 25 criteria across 8 categories (0–4 each, 100 total). Each criterion has concrete descriptions per score level.
3. Write the evaluation
Copy assets/EVAL-TEMPLATE.md to the skill directory as EVAL.md. Fill in automated results + manual scores.
Evaluation Process
- Run
eval-skill.py — get the automated structural score
- Read the skill's SKILL.md — understand what it does
- Read/skim the scripts — assess code quality, error handling, testability
- Score each manual criterion using references/rubric.md — concrete criteria per level
- Prioritize findings as P0 (blocks publishing) / P1 (should fix) / P2 (nice to have)
- Write EVAL.md in the skill directory with scores + findings
Categories (8 categories, 25 criteria)
| # |
Category |
Source Framework |
Criteria |
| 1 |
Functional Suitability |
ISO 25010 |
Completeness, Correctness, Appropriateness |
| 2 |
Reliability |
ISO 25010 |
Fault Tolerance, Error Reporting, Recoverability |
| 3 |
Performance / Context |
ISO 25010 + Agent |
Token Cost, Execution Efficiency |
| 4 |
Usability — AI Agent |
Shneiderman, Gerhardt-Powals |
Learnability, Consistency, Feedback, Error Prevention |
| 5 |
Usability — Human |
Tognazzini, Norman |
Discoverability, Forgiveness |
| 6 |
Security |
ISO 25010 + OpenSSF |
Credentials, Input Validation, Data Safety |
| 7 |
Maintainability |
ISO 25010 |
Modularity, Modifiability, Testability |
| 8 |
Agent-Specific |
Novel |
Trigger Precision, Progressive Disclosure, Composability, Idempotency, Escape Hatches |
Interpreting Scores
| Range |
Verdict |
Action |
| 90–100 |
Excellent |
Publish confidently |
| 80–89 |
Good |
Publishable, note known issues |
| 70–79 |
Acceptable |
Fix P0s before publishing |
| 60–69 |
Needs Work |
Fix P0+P1 before publishing |
| <60 |
Not Ready |
Significant rework needed |
Deeper Security Scanning
This evaluator covers security basics (credentials, input validation, data safety) but for thorough security audits of skills under development, consider SkillLens (npx skilllens scan <path>). It checks for exfiltration, code execution, persistence, privilege bypass, and prompt injection — complementary to the quality focus here.
Dependencies
- Python 3.6+ (for eval-skill.py)
- PyYAML (
pip install pyyaml) — for frontmatter parsing in automated checks
1---2name: skill-evaluator3description: Evaluate Clawdbot skills for quality, reliability, and publish-readiness using a multi-framework rubric (ISO 25010, OpenSSF, Shneiderman, agent-specific heuristics). Use when asked to review, audit, evaluate, score, or assess a skill before publishing, or when checking skill quality. Runs automated structural checks and guides manual assessment across 25 criteria.4---56# Skill Evaluator78Evaluate skills across 25 criteria using a hybrid automated + manual approach.910## Quick Start1112### 1. Run automated checks1314```bash15python3 scripts/eval-skill.py /path/to/skill16python3 scripts/eval-skill.py /path/to/skill --json # machine-readable17python3 scripts/eval-skill.py /path/to/skill --verbose # show all details18```1920Checks: file structure, frontmatter, description quality, script syntax, dependency audit, credential scan, env var documentation.2122### 2. Manual assessment2324Use the rubric at [references/rubric.md](references/rubric.md) to score 25 criteria across 8 categories (0–4 each, 100 total). Each criterion has concrete descriptions per score level.2526### 3. Write the evaluation2728Copy [assets/EVAL-TEMPLATE.md](assets/EVAL-TEMPLATE.md) to the skill directory as `EVAL.md`. Fill in automated results + manual scores.2930## Evaluation Process31321. **Run `eval-skill.py`** — get the automated structural score332. **Read the skill's SKILL.md** — understand what it does343. **Read/skim the scripts** — assess code quality, error handling, testability354. **Score each manual criterion** using [references/rubric.md](references/rubric.md) — concrete criteria per level365. **Prioritize findings** as P0 (blocks publishing) / P1 (should fix) / P2 (nice to have)376. **Write EVAL.md** in the skill directory with scores + findings3839## Categories (8 categories, 25 criteria)4041| # | Category | Source Framework | Criteria |42|---|----------|-----------------|----------|43| 1 | Functional Suitability | ISO 25010 | Completeness, Correctness, Appropriateness |44| 2 | Reliability | ISO 25010 | Fault Tolerance, Error Reporting, Recoverability |45| 3 | Performance / Context | ISO 25010 + Agent | Token Cost, Execution Efficiency |46| 4 | Usability — AI Agent | Shneiderman, Gerhardt-Powals | Learnability, Consistency, Feedback, Error Prevention |47| 5 | Usability — Human | Tognazzini, Norman | Discoverability, Forgiveness |48| 6 | Security | ISO 25010 + OpenSSF | Credentials, Input Validation, Data Safety |49| 7 | Maintainability | ISO 25010 | Modularity, Modifiability, Testability |50| 8 | Agent-Specific | Novel | Trigger Precision, Progressive Disclosure, Composability, Idempotency, Escape Hatches |5152## Interpreting Scores5354| Range | Verdict | Action |55|-------|---------|--------|56| 90–100 | Excellent | Publish confidently |57| 80–89 | Good | Publishable, note known issues |58| 70–79 | Acceptable | Fix P0s before publishing |59| 60–69 | Needs Work | Fix P0+P1 before publishing |60| <60 | Not Ready | Significant rework needed |6162## Deeper Security Scanning6364This evaluator covers security basics (credentials, input validation, data safety) but for thorough security audits of skills under development, consider [SkillLens](https://www.npmjs.com/package/skilllens) (`npx skilllens scan <path>`). It checks for exfiltration, code execution, persistence, privilege bypass, and prompt injection — complementary to the quality focus here.6566## Dependencies6768- Python 3.6+ (for eval-skill.py)69- PyYAML (`pip install pyyaml`) — for frontmatter parsing in automated checks