Error Analysis

Run agent-assisted error analysis on a trace store. Help a human build a review interface over Langfuse, organize human observations into failure modes, validate one LLM judge per selected subjective mode, use DocETL to apply judges to trace batches, and emit a corrected failure report.

ai-evals-course 1a10039 2 files · 46.7 KB Updated

File contents

ai-evals-course/cartwheel-homeworks/tree/main/analysis/skill commit 1a10039a9e

Frequently asked questions

npx skillmds@latest add ai-evals-course/error-analysis