Mtrag Un Eval

Evaluates multi-turn RAG systems on handling unanswerable, underspecified, and non-standalone questions. It probes both retrieval ranking quality and generation quality, including the model's ability to correctly refuse or request clarification when context is insufficient. Use when the user wants to benchmark on MTRAG-UN, or asks about evaluating this task. Reports RB_llm.

qhjqhj00 0f33389 3.2 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/mtrag-un-eval commit 0f333891bd

Frequently asked questions

npx skillmds add qhjqhj00/mtrag-un-eval