06 Evaluation

Use when evaluating an agent's response quality and safety before deployment. Covers running agent-evaluate, evaluation dataset format, built-in judges (relevance, groundedness, safety), interpreting results, and customizing eval datasets. Track A Step 6. Consumes a working agent with tools from Steps 1-5. Produces evaluation results and confidence to deploy.

databricks-solutions 51a7f13 11.8 KB Updated

File contents

databricks-solutions/vibe-coding-workshop-template/tree/main/genai-agents/tracks/A-custom-agent-apps/06-evaluation commit 51a7f13f34

Frequently asked questions

npx skillmds@latest add databricks-solutions/06-evaluation