# Website To CLI

> Guided workflow for converting repeatable browser or website interaction flows into CLI automations by pairing Chrome DevTools (port 9222) + Python Playwright with human confirmations. Use when Codex needs to reuse an already-open browser tab, inspect live DOM state, validate each step in the real page, and turn the approved flow into a deterministic CLI command reviewed step-by-step before execution.

- Skill: `picasso250/website-to-cli` (Agent Skill, multi-file: 9 files)
- Install (CLI): `npx skillmds@latest add picasso250/website-to-cli`
- Raw SKILL.md: https://api.skillmd.com/api/skills/picasso250/website-to-cli/raw
- Safety review: pending
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Productivity
- Author: picasso250 (https://skillmd.com/u/picasso250)
- Updated: 2026-09-17
- Page: https://skillmd.com/skills/picasso250/website-to-cli

---


# Website To Cli

## Overview
Guide the user through capturing the manual story, attaching Playwright to Chrome (port 9222), and translating each verified click/typing step into a CLI routine while pausing for approvals after every transformation.

## When to Use
- When the request is “Turn this browser flow into a repeatable CLI,” and Chrome/Edge can start with `--remote-debugging-port=9222`.
- When every site action must be understood, recorded, and confirmed before being codified into CLI arguments.
- When the deliverable should be runnable code (CLI + docs) instead of just prose.

## Workflow

核心迭代流程：
- 先写验证脚本，等用户确认。
- 用户确认后，再合并到主脚本。
- 每次获得用户确认后，立即执行一次 `git add`，把当前已确认的改动加入暂存区。
- 按这个节奏持续循环，不要一开始就把整条流程一次性写完。
- 自动化默认优先确定性、可复现、低噪声，不模拟人类随机行为。
- 对新的复杂页面步骤，验证脚本默认先“观察”，再“行动”；先把看到的目标元素、坐标、文本和结构打印清楚，确认后再写点击/输入动作。
- 在主循环里默认优先复用已经打开的目标 tab；只有明确需要隔离状态、避免干扰或当前页面不存在时，才新开 tab。
- `verify_*` 脚本是一次性验证件，不需要设计 `--help` 或长期兼容接口。
- `verify_*` 脚本默认不要写 `try/except`；应当让异常直接抛出，以便保留现场并暴露真实失败点。
- 当对应步骤已经稳定并合并进主脚本后，相关 `verify_*` 脚本最终都要删除，不保留为长期资产。
- 打开页面后的首次稳定等待，默认固定等待 `2s`。
- 任何物理点击之前，总是先固定等待 `0.2s`。
- 任何物理 hover 之前，也总是先固定等待 `0.2s`。
- 输入策略默认坚持模拟物理输入：先用物理点击聚焦输入框，再用键盘事件输入，不直接使用 `fill`、不直接改 DOM `value`。
- 模拟物理输入时，逐字输入前等待 `0.01s`，逐字输入后也等待 `0.01s`；这段节奏应抽成可复用函数。
- 对单行输入框，完成输入后默认再按一次 `Tab` 失焦，让页面校验和计数器刷新。
- 点击策略默认坚持模拟物理点击，优先使用鼠标坐标点击，而不是 DOM 级别的 `locator.click()`。
- 物理点击坐标默认取元素几何中心：`x = left + width / 2`，`y = top + height / 2`。
- 物理 hover 坐标也默认取元素几何中心：`x = left + width / 2`，`y = top + height / 2`。
- 这个技能现在也承担原先独立 `browser` skill 的 CDP 会话附着职责；涉及浏览器会话选择、tab 枚举、复用已打开页面时，统一使用本技能下的基础脚本，不再依赖独立 `browser` skill。
- 对 Chrome CDP，不要把 `http://127.0.0.1:9222` 直接传给 `connect_over_cdp()`；先读取 `/json/version`，再使用其中的 browser-level `webSocketDebuggerUrl`（下称 `browserWsEndpoint`），否则某些运行时会反复返回 HTTP 400。
- 区分两类 WebSocket：`browserWsEndpoint` 来自 `/json/version`，用于 Playwright `connect_over_cdp()`；`pageWsEndpoint` 来自 `/json/list` 的单个 page target，用于 `eval-tab-js.py`、`collect-console.py` 这类直接观察某个 tab 的 CDP 脚本。

开始任何基于该技能的新自动化之前，先阅读这五个脚本：
- `scripts/ls-tabs.py`
- `scripts/new-tab.py`
- `scripts/cdp_common.py`
- `scripts/eval-tab-js.py`
- `scripts/collect-console.py`

其中前三个脚本是基础设施：
- `ls-tabs.py` 用来确认 DevTools 连接信息、当前有哪些标签页可复用、目标页面是否已经存在，并同时打印 `browserWsEndpoint` 和各 tab 的 `pageWsEndpoint`。
- `new-tab.py` 用 DevTools HTTP `/json/new` 在现有会话里稳定打开一个新的目标标签页，避免 browser-level Playwright attach 卡住。
- `cdp_common.py` 提供 browser-level `webSocketDebuggerUrl` 解析、context/page 枚举、按精确 URL 选页等共用能力。

另外两个脚本用于现场观察：
- `eval-tab-js.py` 用目标 tab 的 `pageWsEndpoint` 直接执行 JavaScript，快速获取 DOM 结构、文本、坐标、属性和页面状态。
- `collect-console.py` 用来监听已打开目标 tab 后续产生的 `console.log`/`console.error` 等 console 消息。

### Step 1: Capture the story and acceptance criteria
- Read the user’s manual flow and list the pages, buttons, inputs, and outputs they care about.
- Confirm success criteria, required inputs (forms, files, credentials), and tolerable failure modes before touching the browser.
- Record the scenarios so that every subsequent CLI action maps back to a concrete story beat.

### Step 2: Attach to Chrome DevTools and pin the websocket
- Launch Chrome/Edge with `--remote-debugging-port=9222` and confirm you can reach `http://127.0.0.1:9222/json/version`.
- Run `python scripts/ls-tabs.py --port 9222` to print the browser-level endpoint plus available tabs and their page-level endpoints; the script uses `127.0.0.1` directly.
- Never pass the raw `http://127.0.0.1:9222` endpoint straight into Playwright; resolve `/json/version` first and keep the returned `browserWsEndpoint`.
- Ask the user which page/window should drive the CLI, note its title/URL, and save the matching page URL or `pageWsEndpoint` for direct tab inspection.
- If the target page is already open, prefer reusing that tab in the verification loop instead of opening a fresh one.

### Step 3: Observe the interaction with CDP and Playwright
- Start observation from `scripts/eval-tab-js.py`, using the confirmed exact page URL to inspect DOM structure, visible text, coordinates, attributes, storage, and page state.
- Create a site-specific main script only after the relevant observation/action snippets have been verified with the user; do not start from a generic long-lived template.
- Reuse the confirmed target tab whenever possible so each validation step stays attached to the same live page state.
- For a new or ambiguous UI state, split verification into an observe step first and an action step second.
- Before writing an action script for a new page state, prefer inspecting the live tab first with `python scripts/eval-tab-js.py --url <exact-url>` and pipe the JavaScript via `--code` or stdin.
- After opening a page, insert the default stabilized wait of `2s` before treating the page as ready.
- Before every physical click, insert the default pre-click wait of `0.2s`.
- Before every physical hover, insert the default pre-hover wait of `0.2s`.
- For text entry, focus the field via a physical click first, then use keyboard events to type; do not mutate the DOM value directly.
- For text entry, wait `0.01s` before each character and `0.01s` after each character; implement this as a reusable helper.
- For single-line fields, press `Tab` after typing so blur-driven validation can run.
- Prefer physical mouse clicks over DOM-triggered clicks unless the user explicitly approves a different strategy.
- Prefer physical mouse hovers over DOM-triggered hover helpers unless there is a specific reason not to.
- Prefer keyboard-driven typing over `fill()` or script-driven value injection unless the user explicitly approves a different strategy.
- Keep every physical click at the element's geometric center.
- Keep every physical hover at the element's geometric center.
- Step through the user’s clicks, typing, waits, and data captures while narrating: describe the selector, why it’s needed, and the expected result.
- Freeze after each step, summarize the intended CLI effect, and ask the user “May I treat this as CLI step X?” before committing it to the site-specific script.
- When anything depends on dynamic data (tokens, IDs, timestamps), capture how the data is derived and how the CLI will accept or compute it.

### Step 4: Assemble the CLI surface
- Map the approved Playwright steps into sequential CLI arguments or subcommands, keeping each flag tied to a single action.
- Document any required environment variables (e.g., `PLAYWRIGHT_BROWSERS_PATH`, credentials, `browserWsEndpoint` or target page URL), and expose them through `argparse`/`click`.
- Keep the user involved by reading each new CLI flag aloud, explaining how it changes the browser automation, and confirming that it matches their mental model.

### Step 5: Validate and document with the human in the loop
- Run the CLI scenarios, show the logs or screenshots to the user, and ask them to confirm the outputs match the original manual steps.
- If the behavior drifts, note the discrepancy, adjust the script, and replay the scenario until the user signs off.
- Record each confirmation, checkbox, and assumption (see `references/human-in-loop.md`) so the next agent or developer understands what “approved” meant.

## Resources
- Use `references/human-in-loop.md` whenever you need phrasing for check-ins, progress updates, or describing trade-offs and open questions.
- Read `scripts/ls-tabs.py`, `scripts/new-tab.py`, `scripts/cdp_common.py`, `scripts/eval-tab-js.py`, and `scripts/collect-console.py` before building a new site-specific script on top of this skill.
- Run `python scripts/ls-tabs.py --port 9222` before touching the CLI to capture the correct `browserWsEndpoint` for Playwright and identify the target tab for page-level inspection.
- Prefer reusing an existing matching tab discovered via `scripts/ls-tabs.py`; use `scripts/new-tab.py` only when no suitable page is already open.
- Use `python scripts/eval-tab-js.py --url <exact-page-url> --code "<js>"` or pipe JS via stdin when you need to inspect a live tab before deciding selectors or physical click coordinates.
- Use `python scripts/collect-console.py --url <exact-page-url> --seconds 10` when you need to capture new console output from a live tab.
- Run `python scripts/new-tab.py --url <target>` when the next verification step needs a fresh tab instead of reusing an existing page.
- Start each new automation by observing the live tab with `python scripts/eval-tab-js.py --url <exact-page-url> --code "<js>"`, then move only confirmed snippets into a site-specific CLI script.

