LLM Evaluator

LLM-as-a-Judge evaluation system using Langfuse. Score AI outputs on relevance, accuracy, hallucination, and helpfulness. Backfill scoring on historical traces. Uses GPT-5-nano for cost-efficient judging. Use when evaluating AI quality, building evals, or monitoring output accuracy.

dvcrn Updated 32 repo stars

File contents

dvcrn/openclaw-skills-marketplace/tree/main/plugins/aiwithabidi--llm-evaluator/skills/llm-evaluator commit 310b303496

Frequently asked questions

npx skillmds@latest add dvcrn/llm-evaluator