Evaluation

This skill should be used when the user asks to "evaluate agent performance", "build test framework", "measure agent quality", "create evaluation rubrics", or mentions LLM-as-judge, multi-dimensional evaluation, agent testing, or quality gates for agent pipelines. Use when this capability is needed.

tomevault-io Updated

File contents

tomevault-io/skills-registry/tree/main/muratcankoylan--agent-skills-for-context-engineering--evaluation commit 0e9c97e4d0

Frequently asked questions

npx skillmds@latest add tomevault-io/evaluation-8