Trec RAG Support Eval

Evaluates the ability of LLM judges versus human annotators to assess sentence-level grounding (support) in RAG-generated answers. It measures how well models cite relevant passages and whether the cited text actually supports the generated claims. Use when the user wants to benchmark on TREC 2024 RAG Track, or asks about evaluating this task. Reports weighted precision.

qhjqhj00 293ea9a 3.2 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/trec-rag-support-eval commit 293ea9a0f4

Frequently asked questions

npx skillmds add qhjqhj00/trec-rag-support-eval