Umbrela Relevance Assessment Eval

This protocol evaluates the reliability of automatically generated relevance judgments (via the UMBRELA tool) compared to human assessments across different workflow conditions. It measures how well LLM-generated qrels align with human qrels in ranking retrieval systems using standard IR metrics and rank correlation. Use when the user wants to benchmark on TREC 2024 RAG Track, or asks about evaluating this task. Reports Kendall's τ.

qhjqhj00 ac2cb99 3.9 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/umbrela-relevance-assessment-eval commit ac2cb99f7c

Frequently asked questions

npx skillmds add qhjqhj00/umbrela-relevance-assessment-eval