Evaluation

This skill should be used when the user asks to "evaluate agent performance", "build test framework", "measure agent quality", "create evaluation rubrics", or mentions LLM-as-judge, multi-dimensional evaluation, agent testing, or quality gates for agent pipelines. Use when this capability is needed.

tomevault-io Updated

File contents

tomevault-io/skills-registry/tree/main/aldy505--atrium--evaluation commit 44ace7b3d9

Frequently asked questions

npx skillmds@latest add tomevault-io/evaluation