LLM As Judge

Built-in LLM-as-Judge evaluation for Copilot/Claude. Score LLM responses, compare models, detect hallucinations, and verify quality—all within the agent context. No external APIs. Supports single-agent and multi-agent (sequential consensus) evaluation with improved sub-agent prompts for deeper reasoning. Use when comparing responses, validating output quality, checking for hallucinations, or performing model comparisons.

nainishshafi Updated

File contents

nainishshafi/developer-productivity-skills/tree/main/.github/skills/llm-as-judge commit 7f196441c9

Frequently asked questions

npx skillmds@latest add nainishshafi/llm-as-judge