Judgment Eval

Evaluates agent judgment quality through scenario-based testing in-conversation. Use when the user wants to test, validate, or stress-test an agent, skill, or command definition — e.g. "test this agent", "evaluate this skill", "does this prompt handle edge cases", "check this agent's judgment", or after writing or modifying any agent/skill/command .md file. Use when this capability is needed.

tomevault-io 11e220b 2 files · 8.1 KB Updated

File contents

tomevault-io/skills-registry/tree/main/iamladi--cautious-computing-machine--sdlc-plugin--judgment-eval commit 11e220bd87

Frequently asked questions

npx skillmds@latest add tomevault-io/judgment-eval