Agentic Eval

Patterns and techniques for evaluating and improving AI agent outputs. Use this skill when: - Implementing self-critique and reflection loops - Building evaluator-optimizer pipelines for quality-critical generation - Creating test-driven code refinement workflows - Designing rubric-based or LLM-as-judge evaluation systems - Adding iterative improvement to agent outputs (code, reports, analysis) - Measuring and improving agent response quality Do NOT use for unit testing, code linting, static analysis, benchmarking non-agent outputs, or simple pass/fail validation without iteration.

eggboy 5731a8b 8.8 KB Updated

File contents

eggboy/skills/tree/main/agentic-eval commit 5731a8b3db

Frequently asked questions

npx skillmds@latest add eggboy/agentic-eval