Eval Guide

Use when writing eval code, configuring eval infrastructure, creating golden datasets, setting up PromptRegistry, authoring CI eval gates, or working with any eval tool: DeepEval, Ragas, Giskard OSS v3, Promptfoo, Langfuse, Arize Phoenix, adk eval, ADK User Simulation, Vertex GenAI Eval. Covers per-agent accuracy thresholds, CI tier structure (R1-R4), MCP eval suites, golden dataset structure, and PromptRegistry architecture. Also covers pytest harness configuration (asyncio_mode, InMemoryRunner, parametrize-over-golden).

kumaran-is 3a88457 12 files · 70.0 KB Updated

File contents

kumaran-is/claude-code-onboarding/tree/main/.claude/skills/eval-guide commit 3a8845753d

Frequently asked questions

npx skillmds@latest add kumaran-is/eval-guide