AI Evaluation

Define how to evaluate ML models and GenAI/LLM apps — golden datasets, offline metrics, RAG evaluation (faithfulness, relevance, context precision/recall) with ragas/deepeval/promptfoo, LLM-as-judge design and calibration, eval regression gates in CI, online A/B and human feedback. Logging/comparing runs → ml-experiment-tracking; building the app → llm-app-engineering; general experiment statistics → statistical-analysis.

SWEStash e07b19d 3 files · 13.5 KB Updated

File contents

SWEStash/swe-workflow-skills/tree/main/skills/ai-evaluation commit e07b19d0f4

Frequently asked questions

npx skillmds@latest add swestash/ai-evaluation