Advanced Evaluation

This skill should be used when the user asks to "implement LLM-as-judge", "compare model outputs", "create evaluation rubrics", "mitigate evaluation bias", or mentions direct scoring, pairwise comparison, position bias, evaluation pipelines, or automated quality assessment.

Ghosteken 6a12b4f 17.4 KB Updated

File contents

Ghosteken/agent-harness/tree/main/skills/advanced-evaluation commit 6a12b4f264

Frequently asked questions

npx skillmds@latest add ghosteken/advanced-evaluation