Odysseys Eval

Probes an agent's ability to perform realistic, long-horizon web navigation tasks that require sustained cross-site reasoning, context maintenance across multiple tabs, and efficient action execution. It evaluates whether models can complete complex, multi-step user journeys derived from real browsing behavior within strict step budgets. Use when the user wants to benchmark on Odysseys, or asks about evaluating this task. Reports Perfect Rubrics (%).

qhjqhj00 c7a9e72 3.1 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/odysseys-eval commit c7a9e729a4

Frequently asked questions

npx skillmds add qhjqhj00/odysseys-eval