LLM Evaluator

LLM-as-a-Judge evaluator via Langfuse. Scores traces on relevance, accuracy, hallucination, and helpfulness using GPT-5-nano as judge. Supports single trace scoring, batch backfill, and test mode. Integrates with Langfuse dashboard for observability. Triggers: evaluate trace, score quality, check accuracy, backfill scores, test evaluator, LLM judge.

modbender Updated 12 repo stars

File contents

modbender/skill-library-mcp/tree/main/data/llm-evaluator-pro commit d223181a4a

Frequently asked questions

npx skillmds@latest add modbender/llm-evaluator