URL to Markdown
Fetches any URL via baoyu-fetch CLI (Chrome CDP + site-specific adapters) and converts it to clean markdown.
CLI Setup
Important: The CLI source is vendored in the scripts/vendor/baoyu-fetch/ subdirectory of this skill.
Agent Execution Instructions:
- Determine this SKILL.md file's directory path as
{baseDir}
- CLI entry point =
{baseDir}/scripts/vendor/baoyu-fetch/src/cli.ts
- Resolve
${BUN_X} runtime: if bun installed → bun; if npx available → npx -y bun; else suggest installing bun
${READER} = ${BUN_X} {baseDir}/scripts/vendor/baoyu-fetch/src/cli.ts
- Replace all
${READER} in this document with the resolved value
Preferences (EXTEND.md)
Check EXTEND.md existence (priority order):
# macOS, Linux, WSL, Git Bash
test -f .baoyu-skills/baoyu-url-to-markdown/EXTEND.md && echo "project"
test -f "${XDG_CONFIG_HOME:-$HOME/.config}/baoyu-skills/baoyu-url-to-markdown/EXTEND.md" && echo "xdg"
test -f "$HOME/.baoyu-skills/baoyu-url-to-markdown/EXTEND.md" && echo "user"
# PowerShell (Windows)
if (Test-Path .baoyu-skills/baoyu-url-to-markdown/EXTEND.md) { "project" }
$xdg = if ($env:XDG_CONFIG_HOME) { $env:XDG_CONFIG_HOME } else { "$HOME/.config" }
if (Test-Path "$xdg/baoyu-skills/baoyu-url-to-markdown/EXTEND.md") { "xdg" }
if (Test-Path "$HOME/.baoyu-skills/baoyu-url-to-markdown/EXTEND.md") { "user" }
| Path |
Location |
.baoyu-skills/baoyu-url-to-markdown/EXTEND.md |
Project directory |
$HOME/.baoyu-skills/baoyu-url-to-markdown/EXTEND.md |
User home |
| Result |
Action |
| Found |
Read, parse, apply settings |
| Not found |
MUST run first-time setup (see below) — do NOT silently create defaults |
EXTEND.md Supports: Download media by default | Default output directory
First-Time Setup (BLOCKING)
CRITICAL: When EXTEND.md is not found, you MUST use AskUserQuestion to ask the user for their preferences before creating EXTEND.md. NEVER create EXTEND.md with defaults without asking. This is a BLOCKING operation — do NOT proceed with any conversion until setup is complete.
Use AskUserQuestion with ALL questions in ONE call:
Question 1 — header: "Media", question: "How to handle images and videos in pages?"
- "Ask each time (Recommended)" — After saving markdown, ask whether to download media
- "Always download" — Always download media to local imgs/ and videos/ directories
- "Never download" — Keep original remote URLs in markdown
Question 2 — header: "Output", question: "Default output directory?"
- "url-to-markdown (Recommended)" — Save to ./url-to-markdown/{domain}/{slug}.md
- (User may choose "Other" to type a custom path)
Question 3 — header: "Save", question: "Where to save preferences?"
- "User (Recommended)" — ~/.baoyu-skills/ (all projects)
- "Project" — .baoyu-skills/ (this project only)
After user answers, create EXTEND.md at the chosen location, confirm "Preferences saved to [path]", then continue.
Full reference: references/config/first-time-setup.md
Supported Keys
| Key |
Default |
Values |
Description |
download_media |
ask |
ask / 1 / 0 |
ask = prompt each time, 1 = always download, 0 = never |
default_output_dir |
empty |
path or empty |
Default output directory (empty = ./url-to-markdown/) |
EXTEND.md → CLI mapping:
| EXTEND.md key |
CLI argument |
Notes |
download_media: 1 |
--download-media |
Requires --output to be set |
default_output_dir: ./posts/ |
Agent constructs --output ./posts/{domain}/{slug}.md |
Agent generates path, not a direct CLI flag |
Value priority:
- CLI arguments (
--download-media, --output)
- EXTEND.md
- Skill defaults
Features
- Chrome CDP for full JavaScript rendering via
baoyu-fetch CLI
- Site-specific adapters: X/Twitter, YouTube, Hacker News, generic (Defuddle)
- Automatic adapter selection based on URL, or force with
--adapter
- Interaction gate detection: Cloudflare, reCAPTCHA, hCAPTCHA, custom challenges
- Two capture modes: headless (default) or interactive with wait-for-interaction
- Clean markdown output with YAML front matter
- Structured JSON output available via
--format json
- X/Twitter: extracts tweets, threads, and X Articles with media
- YouTube: transcript/caption extraction, chapters, cover images
- Hacker News: threaded comment parsing with proper nesting
- Generic: Defuddle extraction with Readability fallback
- Download images and videos to local directories
- Chrome profile persistence for authenticated sessions
- Debug artifact output for troubleshooting
Usage
# Default: headless capture, markdown to stdout
${READER} <url>
# Save to file
${READER} <url> --output article.md
# Save with media download
${READER} <url> --output article.md --download-media
# Headless mode (explicit)
${READER} <url> --headless --output article.md
# Wait for interaction (login/CAPTCHA) — auto-detect and continue
${READER} <url> --wait-for interaction --output article.md
# Wait for interaction — manual control (Enter to continue)
${READER} <url> --wait-for force --output article.md
# JSON output
${READER} <url> --format json --output article.json
# Force specific adapter
${READER} <url> --adapter youtube --output transcript.md
# Connect to existing Chrome
${READER} <url> --cdp-url http://localhost:9222 --output article.md
# Debug artifacts
${READER} <url> --output article.md --debug-dir ./debug/
Options
| Option |
Description |
<url> |
URL to fetch |
--output <path> |
Output file path (default: stdout) |
--format <type> |
Output format: markdown (default) or json |
--json |
Shorthand for --format json |
--adapter <name> |
Force adapter: x, youtube, hn, or generic (default: auto-detect) |
--headless |
Force headless Chrome (no visible window) |
--wait-for <mode> |
Interaction wait mode: none (default), interaction, or force |
--wait-for-interaction |
Alias for --wait-for interaction |
--wait-for-login |
Alias for --wait-for interaction |
--timeout <ms> |
Page load timeout (default: 30000) |
--interaction-timeout <ms> |
Login/CAPTCHA wait timeout (default: 600000 = 10 min) |
--interaction-poll-interval <ms> |
Poll interval for interaction checks (default: 1500) |
--download-media |
Download images/videos to local imgs/ and videos/, rewrite markdown links. Requires --output |
--media-dir <dir> |
Base directory for downloaded media (default: same as --output directory) |
--cdp-url <url> |
Reuse existing Chrome DevTools Protocol endpoint |
--browser-path <path> |
Custom Chrome/Chromium binary path |
--chrome-profile-dir <path> |
Chrome user data directory (default: BAOYU_CHROME_PROFILE_DIR env or ./baoyu-skills/chrome-profile) |
--debug-dir <dir> |
Write debug artifacts (document.json, markdown.md, page.html, network.json) |
Capture Modes
| Mode |
Behavior |
Use When |
| Default |
Headless Chrome, auto-extract on network idle |
Public pages, static content |
--headless |
Explicit headless (same as default) |
Clarify intent |
--wait-for interaction |
Opens visible Chrome, auto-detects login/CAPTCHA gates, waits for them to clear, then continues |
Login-required, CAPTCHA-protected |
--wait-for force |
Opens visible Chrome, auto-detects OR accepts Enter keypress to continue |
Complex flows, lazy loading, paywalls |
Interaction gate auto-detection:
- Cloudflare Turnstile / "just a moment" pages
- Google reCAPTCHA
- hCaptcha
- Custom challenge / verification screens
Wait-for-interaction workflow:
- Run with
--wait-for interaction → Chrome opens visibly
- CLI auto-detects login/CAPTCHA gates
- User completes login or solves CAPTCHA in the browser
- CLI auto-detects gate cleared → captures page
- If
--wait-for force is used, user can also press Enter to trigger capture manually
Agent Quality Gate
CRITICAL: The agent must treat default headless capture as provisional. Some sites render differently in headless mode and can silently return low-quality content without causing the CLI to fail.
After every headless run, the agent MUST inspect the saved markdown output.
Quality checks the agent must perform
- Confirm the markdown title matches the target page, not a generic site shell
- Confirm the body contains the expected article or page content, not just navigation, footer, or a generic error
- Watch for obvious failure signs:
Application error
This page could not be found
- Login, signup, subscribe, or verification shells
- Extremely short markdown for a page that should be long-form
- Raw framework payloads or mostly boilerplate content
- If the result is low quality, incomplete, or clearly wrong, do not accept the run as successful just because the CLI exited with code 0
Tip: Use --format json to get structured output including status, login.state, and interaction fields for programmatic quality assessment. A "status": "needs_interaction" response means the page requires manual interaction.
Recovery workflow the agent must follow
- First run headless (default) unless there is already a clear reason to use interaction mode
- Review markdown quality immediately after the run
- If the content is low quality or indicates login/CAPTCHA:
--wait-for interaction for auto-detected gates (login, CAPTCHA, Cloudflare)
--wait-for force when the page needs manual browsing, scroll loading, or complex interaction
- If
--wait-for is used, tell the user exactly what to do:
- If login is required, ask them to sign in in the browser
- If CAPTCHA appears, ask them to solve it
- If the page needs time to load, ask them to wait until content is visible
- For
--wait-for force: tell them to press Enter when ready
- If JSON output shows
"status": "needs_interaction", switch to --wait-for interaction automatically
Output Path Generation
The agent must construct the output file path since baoyu-fetch does not auto-generate paths.
Algorithm:
- Determine base directory from EXTEND.md
default_output_dir or default ./url-to-markdown/
- Extract domain from URL (e.g.,
example.com)
- Generate slug from URL path or page title (kebab-case, 2-6 words)
- Construct:
{base_dir}/{domain}/{slug}/{slug}.md — each URL gets its own directory so media files stay isolated
- Conflict resolution: append timestamp
{slug}-YYYYMMDD-HHMMSS/{slug}-YYYYMMDD-HHMMSS.md
Pass the constructed path to --output. Media files (--download-media) are saved into subdirectories next to the markdown file, keeping each URL's assets self-contained.
Output Format
Markdown output to stdout (or file with --output) as clean markdown text.
JSON output (--format json) returns structured data including:
adapter — which adapter handled the URL
status — "ok" or "needs_interaction"
login — login state detection (logged_in, logged_out, unknown)
interaction — interaction gate details (kind, provider, prompt)
document — structured content (url, title, author, publishedAt, content blocks, metadata)
media — collected media assets with url, kind, role
markdown — converted markdown text
downloads — media download results (when --download-media used)
When --download-media is enabled:
- Images are saved to
imgs/ next to the output file (or in --media-dir)
- Videos are saved to
videos/ next to the output file (or in --media-dir)
- Markdown media links are rewritten to local relative paths
Built-in Adapters
| Adapter |
URLs |
Key Features |
x |
x.com, twitter.com |
Tweets, threads, X Articles, media, login detection |
youtube |
youtube.com, youtu.be |
Transcript/captions, chapters, cover image, metadata |
hn |
news.ycombinator.com |
Threaded comments, story metadata, nested replies |
generic |
Any URL (fallback) |
Defuddle extraction, Readability fallback, auto-scroll, network idle detection |
Adapter is auto-selected based on URL. Use --adapter <name> to override.
Media Download Workflow
Based on download_media setting in EXTEND.md:
| Setting |
Behavior |
1 (always) |
Run CLI with --download-media --output <path> |
0 (never) |
Run CLI with --output <path> (no media download) |
ask (default) |
Follow the ask-each-time flow below |
Ask-Each-Time Flow
- Run CLI without
--download-media with --output <path> → markdown saved
- Check saved markdown for remote media URLs (
https:// in image/video links)
- If no remote media found → done, no prompt needed
- If remote media found → use
AskUserQuestion:
- header: "Media", question: "Download N images/videos to local files?"
- "Yes" — Download to local directories
- "No" — Keep remote URLs
- If user confirms → run CLI again with
--download-media --output <same-path> (overwrites markdown with localized links)
Environment Variables
| Variable |
Description |
BAOYU_CHROME_PROFILE_DIR |
Chrome user data directory (can also use --chrome-profile-dir) |
Troubleshooting: Chrome not found → use --browser-path. Timeout → increase --timeout. Login/CAPTCHA pages → use --wait-for interaction. Debug → use --debug-dir to inspect captured HTML and network logs.
YouTube Notes
- YouTube adapter extracts transcripts/captions automatically when available
- Transcript format:
[MM:SS] Text segment with chapter headings
- Transcript availability depends on YouTube exposing a caption track. Videos with captions disabled or restricted playback may produce description-only output
- Use
--wait-for force if the page needs time to finish loading player metadata
X/Twitter Notes
- Extracts single tweets, threads, and X Articles
- Auto-detects login state; if logged out and content requires auth, JSON output will show
"status": "needs_interaction"
- Use
--wait-for interaction for login-protected content
Hacker News Notes
- Parses threaded comments with proper nesting and reply hierarchy
- Includes story metadata (title, URL, author, score, comment count)
- Shows comment deletion/dead status
Extension Support
Custom configurations via EXTEND.md. See Preferences section for paths and supported options.
1---2name: baoyu-url-to-markdown3description: Fetch any URL and convert to markdown using baoyu-fetch CLI (Chrome CDP with site-specific adapters). Built-in adapters for X/Twitter, YouTube transcripts, Hacker News threads, and generic pages via Defuddle. Handles login/CAPTCHA via interaction wait modes. Use when user wants to save a webpage as markdown.4---56# URL to Markdown78Fetches any URL via `baoyu-fetch` CLI (Chrome CDP + site-specific adapters) and converts it to clean markdown.910## CLI Setup1112**Important**: The CLI source is vendored in the `scripts/vendor/baoyu-fetch/` subdirectory of this skill.1314**Agent Execution Instructions**:151. Determine this SKILL.md file's directory path as `{baseDir}`162. CLI entry point = `{baseDir}/scripts/vendor/baoyu-fetch/src/cli.ts`173. Resolve `${BUN_X}` runtime: if `bun` installed → `bun`; if `npx` available → `npx -y bun`; else suggest installing bun184. `${READER}` = `${BUN_X} {baseDir}/scripts/vendor/baoyu-fetch/src/cli.ts`195. Replace all `${READER}` in this document with the resolved value2021## Preferences (EXTEND.md)2223Check EXTEND.md existence (priority order):2425```bash26# macOS, Linux, WSL, Git Bash27test -f .baoyu-skills/baoyu-url-to-markdown/EXTEND.md && echo "project"28test -f "${XDG_CONFIG_HOME:-$HOME/.config}/baoyu-skills/baoyu-url-to-markdown/EXTEND.md" && echo "xdg"29test -f "$HOME/.baoyu-skills/baoyu-url-to-markdown/EXTEND.md" && echo "user"30```3132```powershell33# PowerShell (Windows)34if (Test-Path .baoyu-skills/baoyu-url-to-markdown/EXTEND.md) { "project" }35$xdg = if ($env:XDG_CONFIG_HOME) { $env:XDG_CONFIG_HOME } else { "$HOME/.config" }36if (Test-Path "$xdg/baoyu-skills/baoyu-url-to-markdown/EXTEND.md") { "xdg" }37if (Test-Path "$HOME/.baoyu-skills/baoyu-url-to-markdown/EXTEND.md") { "user" }38```3940| Path | Location |41|------|----------|42| `.baoyu-skills/baoyu-url-to-markdown/EXTEND.md` | Project directory |43| `$HOME/.baoyu-skills/baoyu-url-to-markdown/EXTEND.md` | User home |4445| Result | Action |46|--------|--------|47| Found | Read, parse, apply settings |48| Not found | **MUST** run first-time setup (see below) — do NOT silently create defaults |4950**EXTEND.md Supports**: Download media by default | Default output directory5152### First-Time Setup (BLOCKING)5354**CRITICAL**: When EXTEND.md is not found, you **MUST use `AskUserQuestion`** to ask the user for their preferences before creating EXTEND.md. **NEVER** create EXTEND.md with defaults without asking. This is a **BLOCKING** operation — do NOT proceed with any conversion until setup is complete.5556Use `AskUserQuestion` with ALL questions in ONE call:5758**Question 1** — header: "Media", question: "How to handle images and videos in pages?"59- "Ask each time (Recommended)" — After saving markdown, ask whether to download media60- "Always download" — Always download media to local imgs/ and videos/ directories61- "Never download" — Keep original remote URLs in markdown6263**Question 2** — header: "Output", question: "Default output directory?"64- "url-to-markdown (Recommended)" — Save to ./url-to-markdown/{domain}/{slug}.md65- (User may choose "Other" to type a custom path)6667**Question 3** — header: "Save", question: "Where to save preferences?"68- "User (Recommended)" — ~/.baoyu-skills/ (all projects)69- "Project" — .baoyu-skills/ (this project only)7071After user answers, create EXTEND.md at the chosen location, confirm "Preferences saved to [path]", then continue.7273Full reference: [references/config/first-time-setup.md](references/config/first-time-setup.md)7475### Supported Keys7677| Key | Default | Values | Description |78|-----|---------|--------|-------------|79| `download_media` | `ask` | `ask` / `1` / `0` | `ask` = prompt each time, `1` = always download, `0` = never |80| `default_output_dir` | empty | path or empty | Default output directory (empty = `./url-to-markdown/`) |8182**EXTEND.md → CLI mapping**:83| EXTEND.md key | CLI argument | Notes |84|---------------|-------------|-------|85| `download_media: 1` | `--download-media` | Requires `--output` to be set |86| `default_output_dir: ./posts/` | Agent constructs `--output ./posts/{domain}/{slug}.md` | Agent generates path, not a direct CLI flag |8788**Value priority**:891. CLI arguments (`--download-media`, `--output`)902. EXTEND.md913. Skill defaults9293## Features9495- Chrome CDP for full JavaScript rendering via `baoyu-fetch` CLI96- Site-specific adapters: X/Twitter, YouTube, Hacker News, generic (Defuddle)97- Automatic adapter selection based on URL, or force with `--adapter`98- Interaction gate detection: Cloudflare, reCAPTCHA, hCAPTCHA, custom challenges99- Two capture modes: headless (default) or interactive with wait-for-interaction100- Clean markdown output with YAML front matter101- Structured JSON output available via `--format json`102- X/Twitter: extracts tweets, threads, and X Articles with media103- YouTube: transcript/caption extraction, chapters, cover images104- Hacker News: threaded comment parsing with proper nesting105- Generic: Defuddle extraction with Readability fallback106- Download images and videos to local directories107- Chrome profile persistence for authenticated sessions108- Debug artifact output for troubleshooting109110## Usage111112```bash113# Default: headless capture, markdown to stdout114${READER} <url>115116# Save to file117${READER} <url> --output article.md118119# Save with media download120${READER} <url> --output article.md --download-media121122# Headless mode (explicit)123${READER} <url> --headless --output article.md124125# Wait for interaction (login/CAPTCHA) — auto-detect and continue126${READER} <url> --wait-for interaction --output article.md127128# Wait for interaction — manual control (Enter to continue)129${READER} <url> --wait-for force --output article.md130131# JSON output132${READER} <url> --format json --output article.json133134# Force specific adapter135${READER} <url> --adapter youtube --output transcript.md136137# Connect to existing Chrome138${READER} <url> --cdp-url http://localhost:9222 --output article.md139140# Debug artifacts141${READER} <url> --output article.md --debug-dir ./debug/142```143144## Options145146| Option | Description |147|--------|-------------|148| `<url>` | URL to fetch |149| `--output <path>` | Output file path (default: stdout) |150| `--format <type>` | Output format: `markdown` (default) or `json` |151| `--json` | Shorthand for `--format json` |152| `--adapter <name>` | Force adapter: `x`, `youtube`, `hn`, or `generic` (default: auto-detect) |153| `--headless` | Force headless Chrome (no visible window) |154| `--wait-for <mode>` | Interaction wait mode: `none` (default), `interaction`, or `force` |155| `--wait-for-interaction` | Alias for `--wait-for interaction` |156| `--wait-for-login` | Alias for `--wait-for interaction` |157| `--timeout <ms>` | Page load timeout (default: 30000) |158| `--interaction-timeout <ms>` | Login/CAPTCHA wait timeout (default: 600000 = 10 min) |159| `--interaction-poll-interval <ms>` | Poll interval for interaction checks (default: 1500) |160| `--download-media` | Download images/videos to local `imgs/` and `videos/`, rewrite markdown links. Requires `--output` |161| `--media-dir <dir>` | Base directory for downloaded media (default: same as `--output` directory) |162| `--cdp-url <url>` | Reuse existing Chrome DevTools Protocol endpoint |163| `--browser-path <path>` | Custom Chrome/Chromium binary path |164| `--chrome-profile-dir <path>` | Chrome user data directory (default: `BAOYU_CHROME_PROFILE_DIR` env or `./baoyu-skills/chrome-profile`) |165| `--debug-dir <dir>` | Write debug artifacts (document.json, markdown.md, page.html, network.json) |166167## Capture Modes168169| Mode | Behavior | Use When |170|------|----------|----------|171| Default | Headless Chrome, auto-extract on network idle | Public pages, static content |172| `--headless` | Explicit headless (same as default) | Clarify intent |173| `--wait-for interaction` | Opens visible Chrome, auto-detects login/CAPTCHA gates, waits for them to clear, then continues | Login-required, CAPTCHA-protected |174| `--wait-for force` | Opens visible Chrome, auto-detects OR accepts Enter keypress to continue | Complex flows, lazy loading, paywalls |175176**Interaction gate auto-detection**:177- Cloudflare Turnstile / "just a moment" pages178- Google reCAPTCHA179- hCaptcha180- Custom challenge / verification screens181182**Wait-for-interaction workflow**:1831. Run with `--wait-for interaction` → Chrome opens visibly1842. CLI auto-detects login/CAPTCHA gates1853. User completes login or solves CAPTCHA in the browser1864. CLI auto-detects gate cleared → captures page1875. If `--wait-for force` is used, user can also press Enter to trigger capture manually188189## Agent Quality Gate190191**CRITICAL**: The agent must treat default headless capture as provisional. Some sites render differently in headless mode and can silently return low-quality content without causing the CLI to fail.192193After every headless run, the agent **MUST** inspect the saved markdown output.194195### Quality checks the agent must perform1961971. Confirm the markdown title matches the target page, not a generic site shell1982. Confirm the body contains the expected article or page content, not just navigation, footer, or a generic error1993. Watch for obvious failure signs:200 - `Application error`201 - `This page could not be found`202 - Login, signup, subscribe, or verification shells203 - Extremely short markdown for a page that should be long-form204 - Raw framework payloads or mostly boilerplate content2054. If the result is low quality, incomplete, or clearly wrong, do **not** accept the run as successful just because the CLI exited with code 0206207**Tip**: Use `--format json` to get structured output including `status`, `login.state`, and `interaction` fields for programmatic quality assessment. A `"status": "needs_interaction"` response means the page requires manual interaction.208209### Recovery workflow the agent must follow2102111. First run headless (default) unless there is already a clear reason to use interaction mode2122. Review markdown quality immediately after the run2133. If the content is low quality or indicates login/CAPTCHA:214 - `--wait-for interaction` for auto-detected gates (login, CAPTCHA, Cloudflare)215 - `--wait-for force` when the page needs manual browsing, scroll loading, or complex interaction2164. If `--wait-for` is used, tell the user exactly what to do:217 - If login is required, ask them to sign in in the browser218 - If CAPTCHA appears, ask them to solve it219 - If the page needs time to load, ask them to wait until content is visible220 - For `--wait-for force`: tell them to press Enter when ready2215. If JSON output shows `"status": "needs_interaction"`, switch to `--wait-for interaction` automatically222223## Output Path Generation224225The agent must construct the output file path since `baoyu-fetch` does not auto-generate paths.226227**Algorithm**:2281. Determine base directory from EXTEND.md `default_output_dir` or default `./url-to-markdown/`2292. Extract domain from URL (e.g., `example.com`)2303. Generate slug from URL path or page title (kebab-case, 2-6 words)2314. Construct: `{base_dir}/{domain}/{slug}/{slug}.md` — each URL gets its own directory so media files stay isolated2325. Conflict resolution: append timestamp `{slug}-YYYYMMDD-HHMMSS/{slug}-YYYYMMDD-HHMMSS.md`233234Pass the constructed path to `--output`. Media files (`--download-media`) are saved into subdirectories next to the markdown file, keeping each URL's assets self-contained.235236## Output Format237238Markdown output to stdout (or file with `--output`) as clean markdown text.239240JSON output (`--format json`) returns structured data including:241- `adapter` — which adapter handled the URL242- `status` — `"ok"` or `"needs_interaction"`243- `login` — login state detection (`logged_in`, `logged_out`, `unknown`)244- `interaction` — interaction gate details (kind, provider, prompt)245- `document` — structured content (url, title, author, publishedAt, content blocks, metadata)246- `media` — collected media assets with url, kind, role247- `markdown` — converted markdown text248- `downloads` — media download results (when `--download-media` used)249250When `--download-media` is enabled:251- Images are saved to `imgs/` next to the output file (or in `--media-dir`)252- Videos are saved to `videos/` next to the output file (or in `--media-dir`)253- Markdown media links are rewritten to local relative paths254255## Built-in Adapters256257| Adapter | URLs | Key Features |258|---------|------|-------------|259| `x` | x.com, twitter.com | Tweets, threads, X Articles, media, login detection |260| `youtube` | youtube.com, youtu.be | Transcript/captions, chapters, cover image, metadata |261| `hn` | news.ycombinator.com | Threaded comments, story metadata, nested replies |262| `generic` | Any URL (fallback) | Defuddle extraction, Readability fallback, auto-scroll, network idle detection |263264Adapter is auto-selected based on URL. Use `--adapter <name>` to override.265266## Media Download Workflow267268Based on `download_media` setting in EXTEND.md:269270| Setting | Behavior |271|---------|----------|272| `1` (always) | Run CLI with `--download-media --output <path>` |273| `0` (never) | Run CLI with `--output <path>` (no media download) |274| `ask` (default) | Follow the ask-each-time flow below |275276### Ask-Each-Time Flow2772781. Run CLI **without** `--download-media` with `--output <path>` → markdown saved2792. Check saved markdown for remote media URLs (`https://` in image/video links)2803. **If no remote media found** → done, no prompt needed2814. **If remote media found** → use `AskUserQuestion`:282 - header: "Media", question: "Download N images/videos to local files?"283 - "Yes" — Download to local directories284 - "No" — Keep remote URLs2855. If user confirms → run CLI **again** with `--download-media --output <same-path>` (overwrites markdown with localized links)286287## Environment Variables288289| Variable | Description |290|----------|-------------|291| `BAOYU_CHROME_PROFILE_DIR` | Chrome user data directory (can also use `--chrome-profile-dir`) |292293**Troubleshooting**: Chrome not found → use `--browser-path`. Timeout → increase `--timeout`. Login/CAPTCHA pages → use `--wait-for interaction`. Debug → use `--debug-dir` to inspect captured HTML and network logs.294295### YouTube Notes296297- YouTube adapter extracts transcripts/captions automatically when available298- Transcript format: `[MM:SS] Text segment` with chapter headings299- Transcript availability depends on YouTube exposing a caption track. Videos with captions disabled or restricted playback may produce description-only output300- Use `--wait-for force` if the page needs time to finish loading player metadata301302### X/Twitter Notes303304- Extracts single tweets, threads, and X Articles305- Auto-detects login state; if logged out and content requires auth, JSON output will show `"status": "needs_interaction"`306- Use `--wait-for interaction` for login-protected content307308### Hacker News Notes309310- Parses threaded comments with proper nesting and reply hierarchy311- Includes story metadata (title, URL, author, score, comment count)312- Shows comment deletion/dead status313314## Extension Support315316Custom configurations via EXTEND.md. See **Preferences** section for paths and supported options.