Deepwidesearch Eval

Evaluates agentic systems' ability to perform wide-scale information collection and deep multi-hop reasoning simultaneously to fill structured result tables. It probes combinatorial search complexity, tool orchestration, reflection, and context management in real-world information-seeking tasks. Use when the user wants to benchmark on DeepWideBenchmark, or asks about evaluating this task. Reports Success Rate.

qhjqhj00 f1a443e 4.2 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/deepwidesearch-eval commit f1a443ef21

Frequently asked questions

npx skillmds add qhjqhj00/deepwidesearch-eval