Birco Eval

This benchmark evaluates LLM-based information retrieval systems on complex, multi-faceted query objectives that go beyond simple lexical or semantic similarity. It probes whether models can correctly rank documents based on structured tasks like refuting claims, measuring drug effects, or identifying specific book details, often requiring explicit task understanding rather than just passage matching. Use when the user wants to benchmark on DORIS-MAE, ArguAna, WhatsThatBook, Clinical-Trial, RELIC, or asks about evaluating this task. Reports nDCG@10.

qhjqhj00 6eede77 3.6 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/birco-eval commit 6eede77bf2

Frequently asked questions

npx skillmds add qhjqhj00/birco-eval