LLM Evaluator

LLM-as-a-Judge evaluator via Langfuse. Scores traces on relevance, accuracy, hallucination, and helpfulness using GPT-5-nano as judge. Supports single trace scoring, batch backfill, and test mode. Integrates with Langfuse dashboard for observability. Triggers: evaluate trace, score quality, check accuracy, backfill scores, test evaluator, LLM judge.

dvcrn Updated 32 repo stars

File contents

dvcrn/openclaw-skills-marketplace/tree/main/plugins/aiwithabidi--llm-evaluator-pro/skills/llm-evaluator commit fcda7ff5af

Frequently asked questions

npx skillmds@latest add dvcrn/llm-evaluator-2