Assistantbench Eval

Evaluates web agents' ability to perform realistic, time-consuming multi-hop navigation and information retrieval tasks across the open web. It probes planning, memory, dynamic interaction, and robustness against hallucinations and navigation failures. Use when the user wants to benchmark on AssistantBench, or asks about evaluating this task. Reports Acc..

qhjqhj00 53dd33b 3.7 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/assistantbench-eval commit 53dd33be74

Frequently asked questions

npx skillmds add qhjqhj00/assistantbench-eval