RAG Eval Harness

Set up a controlled RAG evaluation harness. Use when starting a new RAG benchmarking project, adding a new retrieval architecture to an existing benchmark, or evaluating a retrieval change against a stable baseline. The skill enforces fixing all components except the architecture under test, which is the methodological prerequisite for a meaningful comparison.

surpradhan d0c3b6c 8.8 KB Updated

File contents

surpradhan/claude-code-for-ai-engineers/tree/main/skills/rag-eval-harness commit d0c3b6c325

Frequently asked questions

npx skillmds@latest add surpradhan/rag-eval-harness