Browsecomp Plus Eval

Evaluates the end-to-end effectiveness of deep-research agents in retrieving evidence and answering complex queries, as well as the standalone effectiveness of various retrievers. It probes the interplay between retrieval quality, reasoning capability, and search efficiency in agentic workflows. Use when the user wants to benchmark on BrowseComp-Plus, or asks about evaluating this task. Reports Accuracy.

qhjqhj00 7ec74a6 3.9 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/browsecomp-plus-eval commit 7ec74a6750

Frequently asked questions

npx skillmds add qhjqhj00/browsecomp-plus-eval