LLM Evals And Retrieval Quality

Reference-grade guide to evaluating LLM and RAG systems — golden/regression/adversarial eval sets, LLM-as-judge and its biases, retrieval metrics (recall@k, MRR, nDCG), grounding/faithfulness/attribution, RAGAS-style scoring, eval-set construction from production traces, and the CI gates that stop silent regressions.

jpoindexter Updated

File contents

jpoindexter/design-and-ai-skills/tree/main/ai-engineering-skills/llm-evals-and-retrieval-quality commit 55b86f32c4

Frequently asked questions

npx skillmds@latest add jpoindexter/llm-evals-and-retrieval-quality