Relevance Judgment Eval

Evaluates whether pointwise re-rankers can function as binary relevance judges by predicting whether a document is relevant to a query. It probes the capability of adapted ranking models to perform direct relevance classification and compares their performance against LLM-based judges. Use when the user wants to benchmark on TREC-DL, or asks about evaluating this task. Reports binary accuracy.

qhjqhj00 af2d181 2.6 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/relevance-judgment-eval commit af2d18166d

Frequently asked questions

npx skillmds add qhjqhj00/relevance-judgment-eval