Eval Agent Md

Behavioral compliance testing for any CLAUDE.md or agent definition file. Auto-generates test scenarios from your rules, runs them via LLM-as-judge scoring, and reports a compliance score with per-rule pass/fail breakdown. Optionally improves failing rules via automated mutation loop. Use when: (1) testing whether your CLAUDE.md rules are actually followed, (2) evaluating an agent definition for role-boundary compliance, (3) dogfooding a skill's own SKILL.md. Triggers on: "eval", "compliance test", "test my CLAUDE.md", "check rules", "behavioral test", "/eval-agent-md". Do not trigger for: editing or writing CLAUDE.md rules, general code review, adding linting config, or any task that is not explicitly about testing behavioral compliance.

ravnhq 4c6dc6f 238 files · 426.7 KB Updated

File contents

ravnhq/ai-toolkit/tree/main/skills/assistant/eval-agent-md commit 4c6dc6fb2c

Frequently asked questions

npx skillmds@latest add ravnhq/eval-agent-md