AI Eval Playbook

Answer questions about evaluating GenAI / AI tools and products using the AI Evaluation Playbook (eval.playbook.org.ai; The Agency Fund with CGD and IDinsight, 2026) — the 4-level framework (Model / Product / User / Impact), Minimum Viable Evaluations, golden datasets, LLM-as-judge, engagement funnels, guardrail metrics, level linkages. Use when Ane asks how to evaluate an AI tool, chatbot or GenAI product, "how do I know this AI is working", "what should we measure for this AI tool", "where do I start with AI evals", or an MA/partner AI tool proposal needs an evaluation appraisal. NOT for programme or MEL evaluation design (/ann, /grill-mel, indicator-designer, evidence-synthesis own that lane) and NOT for AI used by evaluators as a working tool (wiki page ai-mel-framework governs that).

gasserane Updated

File contents

gasserane/personal-skills/tree/main/skills/ai-eval-playbook commit d2a6d547c5

Frequently asked questions

npx skillmds@latest add gasserane/ai-eval-playbook