X Webagentbench Eval

Evaluates LLM-based agents' ability to comprehend multilingual shopping instructions and successfully navigate interactive web environments across 14 languages. Use when the user wants to benchmark on X-WebAgentBench, or asks about evaluating this task. Reports Task Score.

qhjqhj00 f387b49 2.3 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/x-webagentbench-eval commit f387b49ac7

Frequently asked questions

npx skillmds add qhjqhj00/x-webagentbench-eval