# Browser Automation Expert

> Drive a real browser to navigate, extract data and complete flows on sites without an API: scraping, crawling, authentication, dynamic content and anti-bot handling. Use when the user mentions web scraping, crawling, browser automation, Puppeteer or headless Chrome, wants data pulled from a website, needs a login or checkout flow driven programmatically, or when the task involves rendering JavaScript pages, screenshots, or extracting structured data from HTML.

- Skill: `personamanagmentlayer/browser-automation-expert` (Agent Skill, multi-file: 2 files)
- Install (CLI): `npx skillmds add personamanagmentlayer/browser-automation-expert`
- Raw SKILL.md: https://api.skillmd.com/api/skills/personamanagmentlayer/browser-automation-expert/raw
- Safety review: pending
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Web & Frontend
- Author: personamanagmentlayer (https://skillmd.com/u/personamanagmentlayer)
- Updated: 2026-09-09
- Page: https://skillmd.com/skills/personamanagmentlayer/browser-automation-expert

---


# Browser Automation Expert

Driving a browser to obtain data or complete a flow that has no API. Distinct
from `playwright-expert`, which covers browsers as a _testing_ tool against your
own application; this skill covers browsers as a _client_ against someone else's.

## Before Automating: Is It Allowed?

Automation of a third-party site is a legal and ethical question before it is a
technical one. Check, and record the answer:

- **Is there an API?** Use it. It is faster, more stable and unambiguously
  permitted.
- **Terms of service** — automated access is often restricted. Read them.
- **`robots.txt`** — not legally binding everywhere, but ignoring it is a
  deliberate act you should be able to justify.
- **Personal data** — scraping triggers GDPR obligations regardless of the data
  being public. There must be a lawful basis, and data subjects have rights over
  what you collected.
- **Copyright and database rights** — publicly readable is not freely reusable.
- **Load** — your throughput is someone else's cost.

If the answer to any of these is unclear, escalate rather than proceeding. This
skill will not help evade access controls, defeat CAPTCHAs on a site that has
refused you, or scrape behind an authentication wall you are not entitled to
cross.

## Core Concepts

### Choose the Lightest Tool

| Need                                           | Tool                                                |
| ---------------------------------------------- | --------------------------------------------------- |
| Static HTML                                    | `httpx` + `selectolax`/`BeautifulSoup` — no browser |
| Data present in a JSON endpoint the page calls | Call that endpoint directly                         |
| JavaScript-rendered content                    | Headless browser                                    |
| Interaction: login, upload, multi-step flow    | Headless browser                                    |

Open the network tab before writing a scraper. Most "JavaScript-heavy" sites
fetch their data from a JSON endpoint you can call directly — a hundred times
cheaper and far more stable than driving a browser.

### Selectors Are the Fragile Part

CSS classes generated by a build tool change on every deploy. Prefer, in order:
stable `id` and `data-*` attributes, accessible roles and labels, text content,
then structural position as a last resort.

### Waiting Is Not Sleeping

`sleep(3)` is both slow and flaky. Wait for a condition — an element, a response,
a network idle state — with a timeout.

## Driving a Browser

```python
from playwright.sync_api import sync_playwright

def fetch_listings(url: str) -> list[dict]:
    with sync_playwright() as p:
        browser = p.chromium.launch(headless=True)
        context = browser.new_context(
            user_agent="ResearchBot/1.0 (+https://example.com/bot)",
            locale="en-GB",
            viewport={"width": 1280, "height": 900},
        )
        page = context.new_page()
        page.goto(url, wait_until="domcontentloaded", timeout=30_000)
        page.wait_for_selector("[data-testid='listing']", timeout=15_000)

        listings = [
            {
                "title": el.get_attribute("data-title"),
                "price": el.locator(".price").inner_text().strip(),
                "url": el.locator("a").get_attribute("href"),
            }
            for el in page.locator("[data-testid='listing']").all()
        ]
        context.close()
        browser.close()
        return listings
```

Identify your bot honestly in the user agent, with a contact URL. It is the
difference between being rate-limited and being blocked, and it is the
professional default.

### Intercepting the data directly

Often the fastest correct approach: let the page load and capture its API
response rather than parsing the DOM.

```python
def capture_api(url: str) -> dict:
    payload = {}

    def on_response(response):
        if "/api/v2/search" in response.url and response.ok:
            payload.update(response.json())

    page.on("response", on_response)
    page.goto(url, wait_until="networkidle")
    return payload
```

The JSON is already structured, already typed, and does not break when the
layout changes.

### Blocking what you do not need

Images, fonts and analytics account for most of the bandwidth and time.

```python
context.route("**/*", lambda route: (
    route.abort() if route.request.resource_type in {"image", "font", "media"}
    else route.continue_()
))
```

## Authentication

Log in once, reuse the session, and never re-authenticate per request.

```python
# One-off: perform the login and persist the session
context = browser.new_context()
page = context.new_page()
page.goto("https://example.com/login")
page.fill("#username", os.environ["SITE_USERNAME"])
page.fill("#password", os.environ["SITE_PASSWORD"])
page.click("button[type=submit]")
page.wait_for_url("**/dashboard")
context.storage_state(path="state.json")          # cookies + localStorage

# Later runs
context = browser.new_context(storage_state="state.json")
```

Credentials come from the environment or a secrets manager. `state.json` is a
credential: keep it out of version control, encrypt it at rest, and set an
expiry so a stale session fails loudly rather than silently returning logged-out
pages.

## Crawling Politely

```python
import asyncio, urllib.robotparser
from urllib.parse import urljoin, urlparse

class Crawler:
    def __init__(self, root: str, concurrency: int = 2, delay: float = 1.0):
        self.root, self.delay = root, delay
        self.seen: set[str] = set()
        self.sem = asyncio.Semaphore(concurrency)
        self.robots = urllib.robotparser.RobotFileParser()
        self.robots.set_url(urljoin(root, "/robots.txt"))
        self.robots.read()

    def allowed(self, url: str) -> bool:
        return (
            urlparse(url).netloc == urlparse(self.root).netloc     # stay on the site
            and self.robots.can_fetch("ResearchBot", url)
        )

    async def crawl(self, url: str, depth: int = 0, max_depth: int = 3):
        if depth > max_depth or url in self.seen or not self.allowed(url):
            return
        self.seen.add(url)
        async with self.sem:
            page_data, links = await self.fetch(url)
            await asyncio.sleep(self.delay)                        # between requests
        yield page_data
        for link in links:
            async for item in self.crawl(link, depth + 1, max_depth):
                yield item
```

Low concurrency and a real delay are not politeness theatre — they keep you below
the threshold that gets an IP range banned, and they respect the fact that a
small site's server is not free.

Honour `Crawl-delay` and `Retry-After` when present. On `429` or `503`, back off
exponentially rather than retrying immediately.

## Anti-Bot Measures

Sites signal how they want to be treated. Read the signal.

- **Rate limiting (429)** — slow down. This is a request, and complying is
  usually enough.
- **A CAPTCHA appearing occasionally** — reduce rate, use a persistent session,
  behave less mechanically.
- **A hard block, an explicit refusal, or a CAPTCHA on every request** — the site
  has declined. Stop and seek permission or a data licence.

Legitimate technical hygiene that also reduces false positives: reuse a session
rather than opening a fresh browser per page, keep a consistent and honest user
agent, request at human-plausible intervals with jitter, do not parallelise
aggressively against one host, and cache so you never fetch the same page twice.

Circumventing an access control that has been applied to you specifically is out
of scope for this skill, whatever the technique.

## Extraction

```python
from selectolax.parser import HTMLParser

def parse_listing(html: str) -> dict:
    tree = HTMLParser(html)

    def text(selector: str) -> str | None:
        node = tree.css_first(selector)
        return node.text(strip=True) if node else None

    return {
        "title": text("h1[itemprop=name]"),
        "price": parse_money(text("[data-price]")),
        "sku": text("[itemprop=sku]"),
    }
```

Check for structured data before parsing the DOM — many sites publish
`application/ld+json` blocks that give you clean, typed records for free.

Validate what you extract. A scraper that silently returns `None` for price
across ten thousand records produces a dataset that looks fine and is worthless.

```python
class Listing(BaseModel):
    title: str = Field(min_length=1)
    price: Decimal = Field(gt=0)
    sku: str
```

## Best Practices

- **Cache aggressively.** Store raw HTML with its URL and timestamp; re-parse
  from cache while developing selectors rather than re-fetching.
- **Make it resumable.** Persist progress so a crawl interrupted at 80 % does not
  restart.
- **Monitor for silent breakage.** Alert when the extraction success rate drops —
  layout changes fail quietly.
- **Run browsers in a container** with memory limits, and always close contexts;
  leaked browser processes exhaust a host quickly.
- **Record the fetch date** with every record. Scraped data ages.
- **Set a per-page timeout and an overall run budget.**
- **Keep the parser separate from the fetcher** so you can re-parse a cached
  corpus without touching the network.

## Anti-Patterns

- **Using a browser for static HTML** — fifty times the cost for no benefit.
- **`sleep()` instead of waiting for a condition** — slow and flaky at once.
- **Selectors on generated class names** — `.css-1x9y2z` breaks on the next deploy.
- **Unbounded concurrency** — a denial of service you did not intend to commit.
- **Ignoring 429 and retrying harder** — escalates a rate limit into a ban.
- **Storing personal data without a lawful basis** — a compliance incident, not a
  technical one.
- **No validation on extracted fields** — silent corruption at scale.

## Reference Documentation

- [Playwright Recipes](references/RECIPES.md) — infinite scroll, pagination,
  file downloads and uploads, iframes, shadow DOM, PDFs and screenshots,
  containerised deployment

## Resources

- [Playwright documentation](https://playwright.dev/python/docs/intro)
- [Puppeteer documentation](https://pptr.dev/)
- [robots.txt specification (RFC 9309)](https://www.rfc-editor.org/rfc/rfc9309.html)
- [EDPB guidance on web scraping and GDPR](https://www.edpb.europa.eu/)
- [schema.org structured data](https://schema.org/)

