Eval Judge

Score LLM and agent outputs using LLM-as-judge techniques — direct scoring against rubrics or pairwise comparison between two outputs. Includes built-in bias mitigation for position bias, length bias, and self-enhancement bias. Load when the user asks to score an output, judge a response, evaluate against a rubric, compare two outputs, do direct scoring, run pairwise comparison, or says "rate this", "which response is better", "score this against the rubric", "judge this output", "LLM as judge this". Sub-skill of eval-output orchestrator.

dvy1987 Updated 3 repo stars

File contents

dvy1987/agent-loom/tree/main/.agents/skills/eval-judge commit 1161fbbcb4

Frequently asked questions

npx skillmds@latest add dvy1987/eval-judge