agent-browser — Automated E2E Browser Testing
agent-browser is a CLI-driven browser automation tool (by Vercel Labs) that the AI agent controls directly via shell commands. No Playwright/Puppeteer code required — the agent issues commands, reads the structured output, and makes decisions.
Installation
# Install globally
npm install -g agent-browser
# Download browser engine (Chrome for Testing)
agent-browser install
Platform notes:
- macOS/Linux: Works natively after install.
- Windows: Has a known issue with Unix domain sockets. Use WSL as a workaround, or run from a Linux container.
- Docker/CI: Install in the image with the two commands above.
Core workflow
- Navigate:
agent-browser open <url> - Snapshot:
agent-browser snapshot -i— returns interactive elements tagged with refs (@e1,@e2, …) - Interact using those refs
- Re-snapshot after navigation or significant DOM changes
- Assert state with
get,is, orwaitcommands - Screenshot for evidence:
agent-browser screenshot path.png
E2E Testing Protocol
When using agent-browser to validate a feature end-to-end:
- Start the dev/preview server if it is not already running (check CLAUDE.md for the project's start command).
- Navigate to the feature's entry URL.
- Snapshot to discover the interactive elements.
- Exercise the happy path — fill inputs, click buttons, submit forms, assert success state.
- Exercise key error/edge paths — missing required fields, invalid input, auth-required pages.
- Screenshot at each key moment — before interaction, after success, after error. Save to
screenshots/with descriptive names. - Check console errors with
agent-browser errorsat the end of the session. - Close the browser:
agent-browser close.
Commands
Navigation
agent-browser open <url> # Navigate to URL
agent-browser back # Go back
agent-browser forward # Go forward
agent-browser reload # Reload page
agent-browser close # Close browser
Snapshot (page analysis)
agent-browser snapshot # Full accessibility tree
agent-browser snapshot -i # Interactive elements only (recommended for testing)
agent-browser snapshot -c # Compact output
agent-browser snapshot -d 3 # Limit depth to 3
agent-browser snapshot -s "#main" # Scope to CSS selector
Interactions (use @refs from snapshot)
agent-browser click @e1 # Click
agent-browser dblclick @e1 # Double-click
agent-browser focus @e1 # Focus element
agent-browser fill @e2 "text" # Clear and type
agent-browser type @e2 "text" # Type without clearing
agent-browser press Enter # Press key
agent-browser press Control+a # Key combination
agent-browser hover @e1 # Hover
agent-browser check @e1 # Check checkbox
agent-browser uncheck @e1 # Uncheck checkbox
agent-browser select @e1 "value" # Select dropdown option
agent-browser scroll down 500 # Scroll page
agent-browser scrollintoview @e1 # Scroll element into view
agent-browser drag @e1 @e2 # Drag and drop
agent-browser upload @e1 file.pdf # Upload file
Get information / Assert state
agent-browser get text @e1 # Get element text
agent-browser get html @e1 # Get innerHTML
agent-browser get value @e1 # Get input value
agent-browser get attr @e1 href # Get attribute
agent-browser get title # Get page title
agent-browser get url # Get current URL
agent-browser get count ".item" # Count matching elements
agent-browser get box @e1 # Get bounding box
agent-browser is visible @e1 # Check if visible
agent-browser is enabled @e1 # Check if enabled
agent-browser is checked @e1 # Check if checked
Wait / Synchronization
agent-browser wait @e1 # Wait for element to appear
agent-browser wait 2000 # Wait milliseconds
agent-browser wait --text "Success" # Wait for text to appear
agent-browser wait --url "**/dashboard" # Wait for URL pattern
agent-browser wait --load networkidle # Wait for network idle
agent-browser wait --fn "window.ready" # Wait for JS condition
Screenshots & PDF
agent-browser screenshot # Screenshot to stdout
agent-browser screenshot screenshots/step1.png # Save to file
agent-browser screenshot --full # Full page screenshot
agent-browser pdf output.pdf # Save page as PDF
Semantic locators (alternative to @refs)
agent-browser find role button click --name "Submit"
agent-browser find text "Sign In" click
agent-browser find label "Email" fill "user@test.com"
agent-browser find first ".item" click
agent-browser find nth 2 "a" text
Debugging
agent-browser open <url> --headed # Show browser window (useful locally)
agent-browser console # View console messages
agent-browser errors # View page errors (run at end of test)
agent-browser highlight @e1 # Highlight element
agent-browser trace start # Start recording trace
agent-browser trace stop trace.zip # Stop and save trace
agent-browser record start ./demo.webm # Record video session
agent-browser record stop # Stop and save video
JavaScript
agent-browser eval "document.title" # Run JavaScript and get result
Network (mocking/interception)
agent-browser network route <url> # Intercept requests
agent-browser network route <url> --abort # Block requests
agent-browser network route <url> --body '{}' # Mock response body
agent-browser network unroute [url] # Remove routes
agent-browser network requests # View tracked requests
Sessions (parallel browsers)
agent-browser --session test1 open site-a.com
agent-browser --session test2 open site-b.com
agent-browser session list
Saved auth state
# Login once, save state
agent-browser open https://app.example.com/login
agent-browser fill @e1 "username"
agent-browser fill @e2 "password"
agent-browser click @e3
agent-browser wait --url "**/dashboard"
agent-browser state save auth.json
# Reuse in subsequent test runs
agent-browser state load auth.json
agent-browser open https://app.example.com/dashboard
JSON output (for programmatic checks)
agent-browser snapshot -i --json
agent-browser get text @e1 --json
Example: End-to-end form submission test
# 1. Start
agent-browser open http://localhost:3000/signup
# 2. Discover
agent-browser snapshot -i
# → textbox "Email" [ref=e1], textbox "Password" [ref=e2], button "Create account" [ref=e3]
# 3. Happy path
agent-browser fill @e1 "test@example.com"
agent-browser fill @e2 "SecurePass123!"
agent-browser screenshot screenshots/before-submit.png
agent-browser click @e3
agent-browser wait --url "**/dashboard"
agent-browser screenshot screenshots/after-signup.png
agent-browser get title
# → Expected: "Dashboard"
# 4. Error path — submit empty form
agent-browser open http://localhost:3000/signup
agent-browser snapshot -i
agent-browser click @e3 # Submit without filling in
agent-browser wait --text "required"
agent-browser screenshot screenshots/validation-errors.png
# 5. Check for console errors
agent-browser errors
# 6. Clean up
agent-browser close
Integration with plan-feature validation levels
When a plan's Level 5 (E2E / Browser Automation) is reached:
- Confirm the dev server is running.
- Follow the E2E Testing Protocol above.
- Save screenshots to
screenshots/<ticket-id>-<description>.png. - Paste the screenshot paths and a pass/fail summary into the plan's completion checklist.
Notes
- Always re-snapshot after a navigation or significant DOM mutation — refs are invalidated.
- Prefer
agent-browser wait --load networkidleafter form submissions before asserting state. - For CI environments, omit
--headed; for local debugging, add it so you can watch the browser. agent-browser errorsat the end of a session catches JS exceptions that wouldn't otherwise surface.