Web Agent Benchmark Eval

Evaluates long-horizon web navigation and information-seeking capabilities. Probes the agent's ability to formulate search queries, browse multiple web pages, synthesize information from diverse sources, and answer complex multi-step questions. Use when the user wants to benchmark on BrowseComp-en, BrowseComp-zh, GAIA, WebWalkerQA, FRAMES, XBench-DeepSearch, HLE, or asks about evaluating this task. Reports Avg@4 Accuracy.

qhjqhj00 9e7d8cc 3.2 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/web-agent-benchmark-eval commit 9e7d8cc028

Frequently asked questions

npx skillmds add qhjqhj00/web-agent-benchmark-eval