Evaluation

Evaluate agent systems with quality gates and LLM-as-judge. Use when you need to measure component quality or implement quality gates. Not for simple unit testing or binary pass/fail checks without nuance.

aibot88 Updated 3 repo stars

File contents

aibot88/sec_skill_store/tree/main/skills/claudskills/evaluation-git-fg-meta-plugin-manager commit aa3eb290a7

Frequently asked questions

npx skillmds@latest add aibot88/evaluation