Skill Eval

Score and grade OpenClaw skills against the AgentSkills spec, apply fixes, and explain why each change is better. Use this skill whenever someone asks to score, grade, rate, or benchmark a skill's quality — even if they just say 'is this good enough to ship' or 'how does this skill look'. Use when: (1) scoring a skill before shipping, (2) grading skill quality against the spec, (3) checking if a skill is production-ready, (4) evaluating trigger description accuracy, (5) benchmarking a CLI wrapper skill for completeness, (6) running the evaluate→fix→explain loop on any skill. Triggers on: 'skill eval', 'score this skill', 'grade this skill', 'rate this skill', 'is this skill good enough', 'is this skill production quality', 'evaluate this skill', 'how good is this skill', 'benchmark skill quality'. NOT for: creating or structurally editing skills (use skill-creator), running behavioral test cases in Docker/subagents (use Claude Code's eval framework or Skill Eval), editing agent workspace files directly.

cyperx84 ca964b8 3 files · 11.2 KB Updated

File contents

cyperx84/agent-skills/tree/main/skills/skill-eval commit ca964b8afc

Frequently asked questions

npx skillmds@latest add cyperx84/skill-eval