Web Testing Protocol (L2)
Exploratory, AI-driven validation of dashboard UI changes — not regression testing. Regression coverage is L1's job (smoke + e2e suites under the dashboard's test directories). L2 catches things L1 misses: layout bugs, mobile regressions, interaction flows that only fail in a real browser.
Adopting in another repo: the procedure (read diff → map to routes → run agent-browser at two viewports → screenshot → verdict) is repo-agnostic. The route table, test paths, and verdict schema below are examples from
onsager-ai/onsager. Fork the skill and replace those concrete bits for your own dashboard.
When to invoke
- A PR touches
apps/dashboard/** - L1 e2e fails and you need to know if it's a real regression, flaky, or env
- Someone says "validate the UI" / "dogfood this change"
The app under test
The CI pipeline builds crates/stiglab/deploy/Dockerfile — a single image bundling the Rust backends (stiglab + synodic) and the prebuilt dashboard SPA. It listens on http://localhost:3000.
Primary routes:
| Route | Page | Heading |
|---|---|---|
/ |
Factory overview | Factory |
/sessions |
Sessions list | Sessions |
/sessions/:id |
Session detail | — (dynamic) |
/nodes |
Nodes list | Nodes |
/artifacts |
Artifacts list | Artifacts |
/spine |
Event spine viewer | — (dynamic) |
/governance |
Governance | Governance |
/settings |
Settings + credentials | Settings |
Viewports (always test both)
- Desktop:
agent-browser set viewport 1280 720 - Mobile:
agent-browser set viewport 375 812
Mobile matters — the dashboard ships with a responsive layout (see the md: breakpoints throughout). Horizontal overflow and hidden nav are the top-two regression classes.
Procedure
- Read the diff (
git diff $DIFF_RANGE) — you will receiveDIFF_RANGEas an env var from CI. - Map changes to routes. A change in
src/pages/SessionsPage.tsx⇒/sessions. A change insrc/components/layout/**⇒ every route. - For each affected route, at each viewport:
agent-browser open http://localhost:3000<route>- Snapshot the page; verify the heading + key elements render.
- Actively exercise interactive elements — don't just check markup:
- Click buttons, submit forms, open dialogs.
- Verify the result — did the UI state change, did the dialog close, did new data appear? Presence of markup is not proof of working.
- Check for layout bugs. On mobile especially: horizontal scroll is a failure; a nav that blocks content is a failure.
agent-browser screenshot --screenshot-dir /tmp/l2-screenshotsthen rename the output to{route-slug}-{desktop|mobile}.png.
- Crystallize findings. When you validate new behavior or catch a bug whose fix you can describe, write a deterministic L1 test under
apps/dashboard/tests/smoke/(component-level) orapps/dashboard/tests/e2e/(browser-level). This is how L2 discoveries become permanent L1 coverage. - Emit the verdict. Return JSON matching
tests/l2-verdict-schema.json:PASSif every affected route passes at both viewports.FAILif any route fails; include the specific failure inviewports[].issues[].
Triage mode
When invoked after an L1 e2e failure, your job is different:
- Read the failing test file(s) under
apps/dashboard/tests/e2e/. - Reproduce against
http://localhost:3000with agent-browser. - For each failure, classify: regression (real bug), flaky (intermittent / timing), or environment (test harness or CI issue).
- Return JSON matching
tests/l2-triage-schema.jsonwith root cause and suggested fix.
Guardrails
- Scope to the diff. Don't re-test the whole app on a one-line change.
- Screenshots are required evidence — no screenshot, the viewport didn't run.
- Don't invent routes. If a new route was added in the diff, use that one; otherwise stick to the table above.
- Keep it cheap. One browser session per viewport is plenty.