Benchmark browser agents on a fixed stealth and task suite with browser-use benchmark

Compare browser-agent reliability on a repeatable task and anti-bot suite before choosing a stack or claiming progress.

agentskillexchange Updated 28 repo stars

File contents

Benchmark browser agents on a fixed stealth and task suite with browser-use benchmark

Compare browser-agent reliability on a repeatable task and anti-bot suite before choosing a stack or claiming progress.

Prerequisites

Python, uv, benchmark repository dependencies, required API keys for the judge model and selected browser provider, target browser agent configuration

Installation

Use the upstream install or setup path that matches your environment:

  • pip install uv
  • uv sync
  • uv run python run_eval.py --browser

Requirements and caveats from upstream:

  • python -c "

Basic usage or getting-started notes:

Documentation

Source

agentskillexchange/skills/tree/main/skills/benchmark-browser-agents-on-a-fixed-stealth-and-task-suite-with-browser-use-benchmark commit 95b4204412

Frequently asked questions

npx skillmds@latest add agentskillexchange/benchmark-browser-agents-on-a-fixed-stealth-and-task-suite-w