Trec Dl Passage Relevance Eval

This benchmark evaluates LLMs on their ability to assign relevance scores to document passages given a search query, comparing their outputs against human judgements. It specifically probes systematic biases such as query-term injection gullibility and instruction manipulation in information retrieval labelling tasks. Use when the user wants to benchmark on TREC DL21+DL22, or asks about evaluating this task. Reports MAE.

qhjqhj00 8f9de47 4.3 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/trec-dl-passage-relevance-eval commit 8f9de47f94

Frequently asked questions

npx skillmds add qhjqhj00/trec-dl-passage-relevance-eval