Databricks Genai Evaluation Observability

Use this skill to review generative-AI evaluation, tracing, and observability design on Databricks: MLflow Tracing instrumentation and span design, trace storage and governance, `mlflow.genai.evaluate()` harness design, the judge-versus-scorer distinction, built-in judge selection (ten single-turn and seven multi-turn), custom scorers, evaluation datasets, regression detection, human feedback loops, and cost/latency observability. Treats every LLM judge as an instrument with error.

VincentChuWaiChow Updated

File contents

VincentChuWaiChow/vanguard-frontier-agentic/tree/main/skills/databricks/databricks-genai-evaluation-observability commit b341757ca4

Frequently asked questions

npx skillmds@latest add vincentchuwaichow/databricks-genai-evaluation-observability