AI Evaluation

Use when designing an evaluation framework for AI/LLM features. Covers golden dataset creation, automated scoring rubrics, hallucination detection, regression testing infrastructure, and production monitoring. Do not use for prompt design (use prompt-engineering) or RAG pipeline architecture (use rag-architecture).

dtsong Updated

File contents

dtsong/agentic-council/tree/main/skills/oracle-ai-evaluation commit ea75180b18

Frequently asked questions

npx skillmds@latest add dtsong/ai-evaluation-3