AI Evaluation

Use when designing an evaluation framework for AI/LLM features. Covers golden dataset creation, automated scoring rubrics, hallucination detection, regression testing infrastructure, and production monitoring. Do not use for prompt design (use prompt-engineering) or RAG pipeline architecture (use rag-architecture).

dtsong Updated

File contents

dtsong/my-claude-setup/tree/main/skills/council/oracle/ai-evaluation commit ab0b1449f8

Frequently asked questions

npx skillmds@latest add dtsong/ai-evaluation-2