Ifir Eval

This benchmark evaluates an information retrieval system's ability to follow complex, domain-specific instructions when retrieving relevant passages. It probes whether models can interpret nuanced constraints (e.g., patient demographics, legal case details, financial goals) rather than just matching keyword semantics. Use when the user wants to benchmark on IfIR, or asks about evaluating this task. Reports InstFol@20.

qhjqhj00 ec1879b 3.6 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/ifir-eval commit ec1879b8cb

Frequently asked questions

npx skillmds add qhjqhj00/ifir-eval