RAG Evaluation Review

Reviews the retrieval and grounding evaluation for a RAG-based GenAI use case in a regulated financial-services firm. Confirms the corpus is fit for the intended purpose, retrieval quality is measured with named methods on a labelled set, grounding (faithfulness, citation precision, refusal on out-of-scope queries) is tested with documented metric semantics, failure modes are catalogued, mitigations are evidenced, and ongoing monitoring runs against a real ground-truth pipeline. Output is a second-line memo on whether the RAG implementation can be relied on for the use case's intended purpose, with named gaps, residual-risk framing, and owner actions. Best for: - A first-line owner has built a RAG system and second-line needs an evaluation review before pre-prod, expansion, or annual revalidation. - A foundation-model swap or a retrieval-stack change has happened and the grounding evaluation needs to be re-confirmed. - An incident on hallucination, off-corpus answer, or cross-scope leakage has surfaced and th

anotb 6521926 14 files · 156.0 KB Updated

File contents

anotb/second-line-financial-services/tree/main/plugins/capability-plugins/ai-governance-model-risk/skills/rag-evaluation-review commit 65219267c4

Frequently asked questions

npx skillmds@latest add anotb/rag-evaluation-review