Skill LLM Judge

Audit Agent Skill design quality through static analysis of SKILL.md — no prompt execution, no code running. Scores 7 design dimensions (100pts): knowledge ratio, expert knowledge craft, specification compliance, progressive disclosure, pattern + freedom fit, predicted usability, and output specification. Use when reviewing a skill's DESIGN before or after functional testing. Outputs structured design_score.json for eval pipeline consumption. Do NOT use when you want to run the skill on real prompts or measure actual output quality — use skill_tester for that.

openJiuwen-ai Updated

File contents

openjiuwen-ai/agent-core/tree/main/openjiuwen/dev_tools/skill_evaluator/skills/skill_judge commit 753aa97bc0

Frequently asked questions

npx skillmds@latest add openjiuwen-ai/skill-llm-judge