Retrieval Robustness Eval

This evaluation probes how consistently large language models maintain or improve their answer quality when provided with retrieved context, specifically measuring resilience to variations in retrieval size, document order, and the risk of performance degradation compared to non-retrieval baselines. Use when the user wants to benchmark on Wikipedia QA benchmark, or asks about evaluating this task. Reports No-Degradation Rate (NDR).

qhjqhj00 0bfba9c 4.0 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/retrieval-robustness-eval commit 0bfba9c4b3

Frequently asked questions

npx skillmds add qhjqhj00/retrieval-robustness-eval