Agentic Evaluation Framework

This skill should be used when the user asks to "evaluate LLM output quality", "set up LLM-as-judge", "build an eval rubric", "compare model outputs pairwise", or "measure agent quality".

borghei 5e19948 5 files · 34.3 KB Updated

File contents

borghei/claude-skills/tree/main/engineering/agentic-evaluation-framework commit 5e199484e0

Frequently asked questions

npx skillmds@latest add borghei/agentic-evaluation-framework