Visible Text Extractor
Use this skill to turn a webpage article, URL, screenshot set, long image set, or local image collection into complete, readable, reusable text.
Core workflow
- Extract visible body text from the main source.
- Discover ordered images and GIF-like assets.
- OCR image content when needed.
- Preserve a raw/audit layer.
- Run a human-first cleanup pass.
- Classify image-like content by likely information type.
- Reconstruct image content into human-readable supplements instead of raw OCR dumps.
- Output polished markdown first; keep raw OCR as JSON or appendix data.
What this skill is good at
- General webpage article extraction
- WeChat / 公众号 article extraction with special handling
- News pages, blogs, tutorials, explainers, and image-heavy articles
- Screenshots and long-image OCR
- Image directory OCR in display order
- GIF frame extraction plus OCR when
ffmpeg is available
- Rebuilding noisy OCR into a cleaner reading version
- Producing either reader-friendly clean output or full transcript-style output
Main script
scripts/extract_visible_text.py
Supporting resources
scripts/postprocess_ocr_text.py — clean OCR output, merge broken spacing, remove obvious garbage, and regroup into readable sections
scripts/extract_with_browser.js — browser-rendered fallback for JS-heavy pages
scripts/extract_gif_frames.sh — GIF frame extraction via ffmpeg
scripts/build_deliverable_docx.js — convert cleaned markdown into a Word document
scripts/build_transcript_docx.js — convert transcript-style markdown into a Word document
scripts/build_authorized_capture_docx.py — one-step pipeline for already-authorized browser pages, saved HTML, screenshots, and mixed inputs into clean markdown + JSON + Word deliverable
scripts/extract_visible_text_deliverable.py — one-step pipeline from source input to clean markdown + JSON + Word deliverable
scripts/extract_visible_text_transcript_deliverable.py — one-step pipeline for transcript-style full extraction output
scripts/extract_visible_text_reading_order_deliverable.py — one-step pipeline for reading-order transcript output
scripts/build_wechat_interleaved_docx.py — reconstruct WeChat article reading order by interleaving extracted body blocks and image OCR text in original flow order
scripts/ocr_high_accuracy.py — higher-accuracy OCR with preprocessing variants and segmented long-image handling
references/output-schema.md — target output structure and cleanup rules
references/deliverable-workflow.md — one-step deliverable workflow guidance
references/troubleshooting.md — failure patterns, environment limits, and how to respond cleanly
references/product-positioning.md — what mature deliverable quality means for this skill
references/generalization-plan.md — how to evolve the skill across travel deals, rule pages, event posters, and tutorial long images
references/universal-article-extractor-spec.md — generalized capability contract for article, mixed-media, and screenshot-heavy extraction
Required behavior
When raw OCR is noisy, do not stop at extraction.
- Keep the raw candidate layer for traceability.
- Prefer readability over raw OCR score when two candidates are close.
- Remove decorative fragments, isolated symbols, repeated garbage, and near-duplicate lines from the polished result.
- Keep uncertainty visible instead of pretending confidence.
- Never silently drop a major section when partial reconstruction is possible.
- Never present raw OCR dump as the final answer if a cleaner reconstruction can be produced.
- Preserve article structure when available: title, subtitle, author/source/time, heading levels, paragraphs, lists, captions, table-like rows, and appended notes.
- Treat information-bearing images as first-class content rather than an appendix afterthought.
- For image-heavy pages, support transcript-style and reading-order outputs in addition to clean article outputs.
WeChat / 公众号 handling
For mp.weixin.qq.com URLs:
- Try dedicated article extraction first when available.
- Fall back to static HTML parsing.
- Fall back again to browser rendering if needed.
- When the user cares about article readability, prefer reconstructing the final Word output in original reading order instead of appending all image OCR at the end.
- Use
scripts/build_wechat_interleaved_docx.py when the task is specifically “keep original article order” for WeChat posts.
- If the page is blocked / validation-gated, report
blocked: true clearly instead of pretending success.
Typical commands
Extract URL to markdown:
python3 {baseDir}/scripts/extract_visible_text.py \
--url 'https://example.com/post' \
--format markdown \
--output result.md
Extract URL to JSON:
python3 {baseDir}/scripts/extract_visible_text.py \
--url 'https://example.com/post' \
--format json \
--output result.json
Extract WeChat article with fallbacks:
python3 {baseDir}/scripts/extract_visible_text.py \
--url 'https://mp.weixin.qq.com/s/xxxx' \
--browser-fallback \
--page-screenshot-ocr \
--format markdown \
--output wechat.md
Extract local screenshot or long image:
python3 {baseDir}/scripts/extract_visible_text.py \
--image ./screenshot.png \
--ocr-images \
--format markdown \
--output image-result.md
Run OCR post-processing:
python3 {baseDir}/scripts/postprocess_ocr_text.py \
--input-json ./ocr-result.json \
--title 'Clean Result' \
--body-text 'Optional summary or body text' \
--output-json ./clean.json \
--output-markdown ./clean.md
Run the one-step deliverable pipeline:
python3 {baseDir}/scripts/extract_visible_text_deliverable.py \
--url 'https://mp.weixin.qq.com/s/xxxx' \
--browser-fallback \
--page-screenshot-ocr \
--ocr-images \
--dedupe \
--output-prefix ./deliverable/result
This should emit:
result.raw.json
result.clean.json
result.clean.md
result.docx
Run the already-authorized capture pipeline when the page can be opened in a browser or exported/saved first:
python3 {baseDir}/scripts/build_authorized_capture_docx.py \
--url 'https://example.com/page' \
--browser-capture \
--ocr-images \
--dedupe \
--output-prefix ./deliverable/captured
Useful cases:
- browser can open the page but direct fetch is incomplete
- user provides a saved HTML page plus screenshots
- user wants one command that turns visible page content into a Word document
- user wants status visibility instead of silent long waits
Operational expectations for this pipeline:
- print stage logs so long OCR jobs do not look stuck
- fail loudly if expected outputs are not created
- detect obvious WeChat validation/interstitial text early
- optionally send the generated docx back to Feishu in one run
- when a source is blocked, stop pretending and switch to authorized-input workflows: saved HTML, screenshots, long images, copied text
Practical optimization rule:
- do not keep hammering a blocked source in the same mode
- if browser/direct fetch returns validation text, pivot immediately to the best authorized artifact path
- prioritize delivery quality: visible content captured by the user is better than repeated blocked fetch attempts
Key options
--url webpage URL
--text-file local plain text / markdown input
--html-file local saved HTML page
--image PATH add one local image or GIF; repeat as needed
--image-dir DIR OCR all supported images / GIFs in a directory
--format markdown|json output format
--output PATH output file path
--ocr-images OCR discovered or provided images
--dedupe deduplicate repeated merged lines
--browser-fallback use browser-rendered fallback for incomplete pages
--page-screenshot-ocr OCR the browser full-page screenshot as a last resort
--gif-mode none|placeholder conservative GIF handling mode
Quality standard
Default target: produce something a human can read comfortably and share without cleanup.
Release-quality target for article deliverables:
- preserve the article's original reading order whenever the source structure allows it
- avoid dumping all image OCR at the end when images belong in the middle of the article
- prefer a comfortable reading experience over a mechanically grouped OCR appendix
- keep English-heavy charts, dashboards, and mixed Chinese-English figures readable enough that key labels, axes, legends, and result summaries survive extraction
The skill should increasingly treat extraction as a full article understanding and recovery problem, not only a body scrape plus OCR problem:
- recover visible article structure from normal webpages, WeChat posts, blogs, tutorials, and mixed-media articles
- infer whether an image is mainly a price/product page, rules page, poster/event page, course outline, scenery/introduction card, or table-like detail page
- pull out high-value facts first when the user wants a clean readable result
- preserve near-complete text when the user wants transcript completeness
- avoid raw OCR dumps as the main deliverable unless the user explicitly wants audit output
When the user explicitly wants completeness, the skill must support a fuller extraction mode:
- treat each discovered image as a first-class source
- prefer segmented OCR for tall or dense images
- preserve near-complete per-image text blocks before compressing into summaries
- keep summary and full-text layers separate instead of replacing one with the other
- support reading-order transcript output so text and image-derived content can be followed from start to finish
For clean article outputs, prefer a structure like:
- Title
- Metadata (author/source/time) when meaningful
- Main sections in order
- Integrated image-derived supplements where needed
- Uncertainty notes only when necessary
For transcript outputs, prefer a structure like:
- Title
- Intro/body chunks in order
- Image text blocks in order or reading order
- Tail matter / credits / appended notes
Mature-skill rule:
- default users toward the clean markdown / docx outputs unless they ask for transcript completeness
- keep raw JSON for audit, not as the main deliverable
- degrade honestly when the source is blocked or image quality is poor
- do not optimize only for one article family; keep checking travel-deal posts, rule/scoring posts, event posters, news/blog/tutorial pages, and course-outline long images
Read these references when needed:
references/output-schema.md
references/deliverable-workflow.md
references/troubleshooting.md
references/product-positioning.md
references/generalization-plan.md
references/universal-article-extractor-spec.md
Environment notes
- OCR depends on the local
ocr-local skill or compatible Tesseract.js setup.
- Browser fallback depends on real browser availability plus
playwright-core support.
- GIF frame extraction depends on
ffmpeg.
- Some pages remain partially inaccessible due to login, anti-bot, or validation flows; mark those limits explicitly.
1---2name: visible-text-extractor3description: Extract and reconstruct as much visible text as possible from webpage URLs, article pages, screenshots, long images, image directories, and GIFs. Use when the goal is not just raw OCR, but a clean, human-readable result with section grouping, OCR cleanup, deduplication, structured JSON, original reading-order reconstruction, and explicit uncertainty notes. Especially useful for WeChat articles, event posters, long screenshots, mixed text-plus-image pages, and cases where visible information must be preserved without dumping noisy OCR into the final answer.4---56# Visible Text Extractor78Use this skill to turn a webpage article, URL, screenshot set, long image set, or local image collection into complete, readable, reusable text.910## Core workflow11121. Extract visible body text from the main source.132. Discover ordered images and GIF-like assets.143. OCR image content when needed.154. Preserve a raw/audit layer.165. Run a human-first cleanup pass.176. Classify image-like content by likely information type.187. Reconstruct image content into human-readable supplements instead of raw OCR dumps.198. Output polished markdown first; keep raw OCR as JSON or appendix data.2021## What this skill is good at2223- General webpage article extraction24- WeChat / 公众号 article extraction with special handling25- News pages, blogs, tutorials, explainers, and image-heavy articles26- Screenshots and long-image OCR27- Image directory OCR in display order28- GIF frame extraction plus OCR when `ffmpeg` is available29- Rebuilding noisy OCR into a cleaner reading version30- Producing either reader-friendly clean output or full transcript-style output3132## Main script3334- `scripts/extract_visible_text.py`3536## Supporting resources3738- `scripts/postprocess_ocr_text.py` — clean OCR output, merge broken spacing, remove obvious garbage, and regroup into readable sections39- `scripts/extract_with_browser.js` — browser-rendered fallback for JS-heavy pages40- `scripts/extract_gif_frames.sh` — GIF frame extraction via `ffmpeg`41- `scripts/build_deliverable_docx.js` — convert cleaned markdown into a Word document42- `scripts/build_transcript_docx.js` — convert transcript-style markdown into a Word document43- `scripts/build_authorized_capture_docx.py` — one-step pipeline for already-authorized browser pages, saved HTML, screenshots, and mixed inputs into clean markdown + JSON + Word deliverable44- `scripts/extract_visible_text_deliverable.py` — one-step pipeline from source input to clean markdown + JSON + Word deliverable45- `scripts/extract_visible_text_transcript_deliverable.py` — one-step pipeline for transcript-style full extraction output46- `scripts/extract_visible_text_reading_order_deliverable.py` — one-step pipeline for reading-order transcript output47- `scripts/build_wechat_interleaved_docx.py` — reconstruct WeChat article reading order by interleaving extracted body blocks and image OCR text in original flow order48- `scripts/ocr_high_accuracy.py` — higher-accuracy OCR with preprocessing variants and segmented long-image handling49- `references/output-schema.md` — target output structure and cleanup rules50- `references/deliverable-workflow.md` — one-step deliverable workflow guidance51- `references/troubleshooting.md` — failure patterns, environment limits, and how to respond cleanly52- `references/product-positioning.md` — what mature deliverable quality means for this skill53- `references/generalization-plan.md` — how to evolve the skill across travel deals, rule pages, event posters, and tutorial long images54- `references/universal-article-extractor-spec.md` — generalized capability contract for article, mixed-media, and screenshot-heavy extraction5556## Required behavior5758When raw OCR is noisy, do not stop at extraction.5960- Keep the raw candidate layer for traceability.61- Prefer readability over raw OCR score when two candidates are close.62- Remove decorative fragments, isolated symbols, repeated garbage, and near-duplicate lines from the polished result.63- Keep uncertainty visible instead of pretending confidence.64- Never silently drop a major section when partial reconstruction is possible.65- Never present raw OCR dump as the final answer if a cleaner reconstruction can be produced.66- Preserve article structure when available: title, subtitle, author/source/time, heading levels, paragraphs, lists, captions, table-like rows, and appended notes.67- Treat information-bearing images as first-class content rather than an appendix afterthought.68- For image-heavy pages, support transcript-style and reading-order outputs in addition to clean article outputs.6970## WeChat / 公众号 handling7172For `mp.weixin.qq.com` URLs:7374- Try dedicated article extraction first when available.75- Fall back to static HTML parsing.76- Fall back again to browser rendering if needed.77- When the user cares about article readability, prefer reconstructing the final Word output in original reading order instead of appending all image OCR at the end.78- Use `scripts/build_wechat_interleaved_docx.py` when the task is specifically “keep original article order” for WeChat posts.79- If the page is blocked / validation-gated, report `blocked: true` clearly instead of pretending success.8081## Typical commands8283Extract URL to markdown:8485```bash86python3 {baseDir}/scripts/extract_visible_text.py \87 --url 'https://example.com/post' \88 --format markdown \89 --output result.md90```9192Extract URL to JSON:9394```bash95python3 {baseDir}/scripts/extract_visible_text.py \96 --url 'https://example.com/post' \97 --format json \98 --output result.json99```100101Extract WeChat article with fallbacks:102103```bash104python3 {baseDir}/scripts/extract_visible_text.py \105 --url 'https://mp.weixin.qq.com/s/xxxx' \106 --browser-fallback \107 --page-screenshot-ocr \108 --format markdown \109 --output wechat.md110```111112Extract local screenshot or long image:113114```bash115python3 {baseDir}/scripts/extract_visible_text.py \116 --image ./screenshot.png \117 --ocr-images \118 --format markdown \119 --output image-result.md120```121122Run OCR post-processing:123124```bash125python3 {baseDir}/scripts/postprocess_ocr_text.py \126 --input-json ./ocr-result.json \127 --title 'Clean Result' \128 --body-text 'Optional summary or body text' \129 --output-json ./clean.json \130 --output-markdown ./clean.md131```132133Run the one-step deliverable pipeline:134135```bash136python3 {baseDir}/scripts/extract_visible_text_deliverable.py \137 --url 'https://mp.weixin.qq.com/s/xxxx' \138 --browser-fallback \139 --page-screenshot-ocr \140 --ocr-images \141 --dedupe \142 --output-prefix ./deliverable/result143```144145This should emit:146- `result.raw.json`147- `result.clean.json`148- `result.clean.md`149- `result.docx`150151Run the already-authorized capture pipeline when the page can be opened in a browser or exported/saved first:152153```bash154python3 {baseDir}/scripts/build_authorized_capture_docx.py \155 --url 'https://example.com/page' \156 --browser-capture \157 --ocr-images \158 --dedupe \159 --output-prefix ./deliverable/captured160```161162Useful cases:163- browser can open the page but direct fetch is incomplete164- user provides a saved HTML page plus screenshots165- user wants one command that turns visible page content into a Word document166- user wants status visibility instead of silent long waits167168Operational expectations for this pipeline:169- print stage logs so long OCR jobs do not look stuck170- fail loudly if expected outputs are not created171- detect obvious WeChat validation/interstitial text early172- optionally send the generated docx back to Feishu in one run173- when a source is blocked, stop pretending and switch to authorized-input workflows: saved HTML, screenshots, long images, copied text174175Practical optimization rule:176- do not keep hammering a blocked source in the same mode177- if browser/direct fetch returns validation text, pivot immediately to the best authorized artifact path178- prioritize delivery quality: visible content captured by the user is better than repeated blocked fetch attempts179180## Key options181182- `--url` webpage URL183- `--text-file` local plain text / markdown input184- `--html-file` local saved HTML page185- `--image PATH` add one local image or GIF; repeat as needed186- `--image-dir DIR` OCR all supported images / GIFs in a directory187- `--format markdown|json` output format188- `--output PATH` output file path189- `--ocr-images` OCR discovered or provided images190- `--dedupe` deduplicate repeated merged lines191- `--browser-fallback` use browser-rendered fallback for incomplete pages192- `--page-screenshot-ocr` OCR the browser full-page screenshot as a last resort193- `--gif-mode none|placeholder` conservative GIF handling mode194195## Quality standard196197Default target: produce something a human can read comfortably and share without cleanup.198199Release-quality target for article deliverables:200- preserve the article's original reading order whenever the source structure allows it201- avoid dumping all image OCR at the end when images belong in the middle of the article202- prefer a comfortable reading experience over a mechanically grouped OCR appendix203- keep English-heavy charts, dashboards, and mixed Chinese-English figures readable enough that key labels, axes, legends, and result summaries survive extraction204205The skill should increasingly treat extraction as a full article understanding and recovery problem, not only a body scrape plus OCR problem:206- recover visible article structure from normal webpages, WeChat posts, blogs, tutorials, and mixed-media articles207- infer whether an image is mainly a price/product page, rules page, poster/event page, course outline, scenery/introduction card, or table-like detail page208- pull out high-value facts first when the user wants a clean readable result209- preserve near-complete text when the user wants transcript completeness210- avoid raw OCR dumps as the main deliverable unless the user explicitly wants audit output211212When the user explicitly wants completeness, the skill must support a fuller extraction mode:213- treat each discovered image as a first-class source214- prefer segmented OCR for tall or dense images215- preserve near-complete per-image text blocks before compressing into summaries216- keep summary and full-text layers separate instead of replacing one with the other217- support reading-order transcript output so text and image-derived content can be followed from start to finish218219For clean article outputs, prefer a structure like:2202211. Title2222. Metadata (author/source/time) when meaningful2233. Main sections in order2244. Integrated image-derived supplements where needed2255. Uncertainty notes only when necessary226227For transcript outputs, prefer a structure like:2282291. Title2302. Intro/body chunks in order2313. Image text blocks in order or reading order2324. Tail matter / credits / appended notes233234Mature-skill rule:235- default users toward the clean markdown / docx outputs unless they ask for transcript completeness236- keep raw JSON for audit, not as the main deliverable237- degrade honestly when the source is blocked or image quality is poor238- do not optimize only for one article family; keep checking travel-deal posts, rule/scoring posts, event posters, news/blog/tutorial pages, and course-outline long images239240Read these references when needed:241- `references/output-schema.md`242- `references/deliverable-workflow.md`243- `references/troubleshooting.md`244- `references/product-positioning.md`245- `references/generalization-plan.md`246- `references/universal-article-extractor-spec.md`247248## Environment notes249250- OCR depends on the local `ocr-local` skill or compatible Tesseract.js setup.251- Browser fallback depends on real browser availability plus `playwright-core` support.252- GIF frame extraction depends on `ffmpeg`.253- Some pages remain partially inaccessible due to login, anti-bot, or validation flows; mark those limits explicitly.