Evaluation

This skill should be used when the user asks to "evaluate agent performance", "build test framework", "measure agent quality", "create evaluation rubrics", or mentions LLM-as-judge, multi-dimensional evaluation, agent testing, or quality gates for agent pipelines.

0xharryriddle e45843e 3 files · 35.9 KB Updated 3 repo stars

File contents

0xharryriddle/awesome-codex-subagents/tree/main/archive/upstream/chasebuild-agent-skills/context-engineering/skills/evaluation commit e45843ece8

Frequently asked questions

npx skillmds@latest add 0xharryriddle/evaluation