html-to-markdown
html-to-markdown is a high-performance HTML→Markdown converter with a Rust core and 12 native language bindings. It converts HTML to CommonMark Markdown, Djot, or plain text in a single pass, optionally extracting metadata, tables, inline images, and a structured document tree.
Use this skill when writing code that:
- Converts HTML strings, files, or live URLs to Markdown, Djot, or plain text
- Extracts metadata (title, OG tags, JSON-LD/Microdata/RDFa, headers, links, images, language) from HTML
- Extracts structured table data (GFM markdown + cell grids) from HTML
- Extracts a structured document-structure tree
- Extracts inline images (data URIs, SVGs) from HTML
- Uses preprocessing to clean noisy HTML (ads, navigation, forms) before conversion
Capability map
| Capability |
CLI |
SDKs |
| HTML→Markdown / Djot / plain text |
html-to-markdown FILE |
convert(html, options) |
| Read HTML from stdin / file / URL |
cat f | …, FILE, --url URL |
convert(htmlString, …) |
| ~30 config options (headings, code blocks, lists, escaping, wrapping…) |
flags |
ConversionOptions |
| Metadata extraction |
--json (default-extracted) |
result.metadata |
| Table extraction |
--json → tables[] |
result.tables |
| Document structure tree |
--json --include-structure |
include_document_structure=true |
| Inline image extraction |
--json --extract-inline-images |
extract_images=true |
| HTML preprocessing |
--preprocess [--preset …] |
PreprocessingOptions |
| Extraction-only (no Markdown body) |
--json --no-content |
options + read fields |
Installation
CLI
# (Homebrew 6.0+ requires explicit trust for third-party taps)
brew trust xberg-io/tap
brew install xberg-io/tap/html-to-markdown
# or run without a persistent install (the CLI proxy package self-installs the binary):
npx @xberg-io/html-to-markdown-cli --help
uvx --from html-to-markdown-cli html-to-markdown --help
# or download a prebuilt binary from the latest GitHub release:
# https://github.com/xberg-io/html-to-markdown/releases/latest
# or build from source:
cargo install html-to-markdown-cli
Language SDKs
pip install html-to-markdown # Python
npm install @xberg-io/html-to-markdown # TypeScript / Node.js
cargo add html-to-markdown-rs # Rust (features: metadata default; full = all)
gem install html-to-markdown # Ruby
composer require xberg-io/html-to-markdown # PHP
go get github.com/xberg-io/html-to-markdown/packages/go/v3 # Go
dotnet add package XbergIo.HtmlToMarkdown # C#
npm install @xberg-io/html-to-markdown-wasm # WASM
- Java (Maven):
io.xberg:html-to-markdown
- Elixir:
{:html_to_markdown, "~> 3.8"} in mix.exs
- R:
install.packages("htmltomarkdown", repos = "https://xberg-io.r-universe.dev")
- C (FFI): pre-built
.so / .dll / .dylib from GitHub releases
CLI vs SDK — which to use
- CLI — one-shot conversions, shell pipelines, fetching a single URL, ad-hoc metadata/table extraction via
--json | jq. Flags only for conversion (FILE is positional; omit or use - for stdin); the only subcommand is mcp.
- SDK — embedding conversion in application code, batch processing, custom element conversion (visitor pattern, Rust), and tight loops where process spawn overhead matters.
- MCP server —
html-to-markdown mcp exposes convert_html and extract_metadata as agent tools, so an MCP client can convert an HTML string directly with no shell-out. This plugin auto-registers it; see the using-the-mcp-server skill.
Both share the same ConversionResult shape, so output is interchangeable.
When to use html-to-markdown vs xberg vs crawlberg
- html-to-markdown — you already have HTML (a string, a file, or a single URL) and want clean Markdown plus structured metadata/tables. No OCR, no document parsing, no crawling.
- xberg — you have documents (PDF, Office, images, email, archives) and need full text/table/metadata extraction with optional OCR. Use it when the input is not already HTML.
- crawlberg — you need to crawl or scrape many pages, follow links, and handle JS-rendered sites with a headless-Chrome fallback. It uses html-to-markdown internally for the HTML→Markdown step.
Rule of thumb: single HTML in → Markdown out = html-to-markdown. Many URLs / a site = crawlberg. Non-HTML documents = xberg.
CLI quick start
# Convert a file to stdout
html-to-markdown input.html
# Convert and save
html-to-markdown input.html -o output.md
# Read from stdin
cat page.html | html-to-markdown
# Fetch and convert a URL
html-to-markdown --url https://example.com > out.md
# Full ConversionResult as JSON (content, tables, metadata, images, warnings)
html-to-markdown --json input.html
# JSON with document structure tree
html-to-markdown --json --include-structure input.html
# Extraction-only (no Markdown body)
html-to-markdown --json --no-content input.html
# Aggressive web-page cleanup
html-to-markdown input.html --preprocess --preset aggressive
SDK quick start
Rust
use html_to_markdown_rs::convert;
let result = convert("<h1>Hello World</h1><p>A paragraph.</p>", None)?;
println!("{}", result.content.unwrap_or_default());
Python
from html_to_markdown import convert
result = convert("<h1>Hello World</h1><p>A paragraph.</p>")
print(result.content) # # Hello World\n\nA paragraph.
print(result.metadata) # title, links, headers, …
TypeScript / Node.js
import { convert } from "@xberg-io/html-to-markdown";
// Node's convert() returns a ConversionResult object directly.
const result = convert("<h1>Hello World</h1><p>A paragraph.</p>");
console.log(result.content);
ConversionResult fields
All languages return the same structure (dict, object, or struct).
| Field |
Description |
content |
Converted text (Markdown/Djot/plain). null only in extraction-only mode. |
metadata |
Title, OG, headers, links, images, structured data. |
tables |
Tables with grid (structured cells) and markdown fields. |
images |
Extracted inline images (requires inline-image extraction). |
document |
Structured document tree when structure extraction is enabled. |
warnings |
Non-fatal processing warnings (message, kind). |
Configuration
All languages expose the same ~30 options. See references/configuration.md for the complete table. Common ones:
| Option |
Values |
Default |
heading_style |
atx, underlined, atx-closed |
atx |
code_block_style |
backticks, indented, tildes |
backticks |
output_format |
markdown, djot, plain |
markdown |
wrap / wrap_width |
bool / 20–500 |
off / 80 |
autolinks (SDK) / --no-autolinks (CLI) |
bool / flag |
true (on); disable in CLI with --no-autolinks |
| preprocessing |
minimal / standard / aggressive |
off |
Rust (builder)
use html_to_markdown_rs::{convert, ConversionOptions, HeadingStyle, OutputFormat};
let options = ConversionOptions::builder()
.heading_style(HeadingStyle::Atx)
.output_format(OutputFormat::Markdown)
.wrap(true)
.wrap_width(100)
.build();
let result = convert(html, Some(options))?;
Python (dataclass)
from html_to_markdown import convert, ConversionOptions, PreprocessingOptions
html = "<h1>Title</h1><p>Body text.</p>"
result = convert(
html,
ConversionOptions(
heading_style="atx",
wrap=True,
wrap_width=100,
preprocessing=PreprocessingOptions(enabled=True, preset="aggressive"),
),
)
Metadata extraction
The library convert() extracts metadata by default; the CLI needs --json --extract-metadata (see the extracting-metadata skill). Fields include document (title, description, language, canonical_url, open_graph), headers, links (with link_type), images, and structured_data (JSON-LD/Microdata/RDFa).
Table extraction
Tables appear in result.tables, each with a pre-rendered markdown string and a structured cell grid. Markdown tables also appear inline in content. See the extracting-tables skill.
Document structure extraction
Enable structure extraction (--include-structure on the CLI, include_document_structure=true in SDKs) to get a semantic node tree under document. Node types include heading, paragraph, list, list_item, table, image, code, quote, group, metadata_block.
Common pitfalls
convert() returns a result object, not a string. Access .content for the Markdown text. This holds for Node.js too — convert() returns a ConversionResult object directly; do not JSON.parse() it.
--json outputs JSON, not Markdown. Omit --json for plain Markdown.
--include-structure, --extract-inline-images, and --no-content require --json.
- The conversion CLI is flags-only.
FILE is positional; the only subcommand is mcp (starts the MCP server).
--preset, --keep-navigation, --keep-forms require --preprocess.
Additional resources
- CLI Reference — every flag, JSON shape, exit codes
- Configuration Reference — all 30+ options with defaults
- Rust API Reference — signatures, builder, feature flags
- Python API Reference — functions, dataclasses, type hints
- TypeScript API Reference — functions, interfaces, Buffer support
- Other Bindings — Go, Ruby, PHP, Java, C#, Elixir, R, WASM, C FFI
GitHub: https://github.com/xberg-io/html-to-markdown
1---2name: html-to-markdown-23description: Convert HTML to Markdown, Djot, or plain text with structured extraction. Use when writing code that calls html-to-markdown APIs in Rust, Python, TypeScript, Go, Ruby, PHP, Java, C#, Elixir, R, C, or WASM. Covers installation, conversion, configuration, metadata extraction, tables, document structure, inline images, URL fetching, and CLI usage.4license: MIT5---67# html-to-markdown89html-to-markdown is a high-performance HTML→Markdown converter with a Rust core and 12 native language bindings. It converts HTML to CommonMark Markdown, Djot, or plain text in a single pass, optionally extracting metadata, tables, inline images, and a structured document tree.1011Use this skill when writing code that:1213- Converts HTML strings, files, or live URLs to Markdown, Djot, or plain text14- Extracts metadata (title, OG tags, JSON-LD/Microdata/RDFa, headers, links, images, language) from HTML15- Extracts structured table data (GFM markdown + cell grids) from HTML16- Extracts a structured document-structure tree17- Extracts inline images (data URIs, SVGs) from HTML18- Uses preprocessing to clean noisy HTML (ads, navigation, forms) before conversion1920## Capability map2122| Capability | CLI | SDKs |23| ---------- | --- | ---- |24| HTML→Markdown / Djot / plain text | `html-to-markdown FILE` | `convert(html, options)` |25| Read HTML from stdin / file / URL | `cat f \| …`, `FILE`, `--url URL` | `convert(htmlString, …)` |26| ~30 config options (headings, code blocks, lists, escaping, wrapping…) | flags | `ConversionOptions` |27| Metadata extraction | `--json` (default-extracted) | `result.metadata` |28| Table extraction | `--json` → `tables[]` | `result.tables` |29| Document structure tree | `--json --include-structure` | `include_document_structure=true` |30| Inline image extraction | `--json --extract-inline-images` | `extract_images=true` |31| HTML preprocessing | `--preprocess [--preset …]` | `PreprocessingOptions` |32| Extraction-only (no Markdown body) | `--json --no-content` | options + read fields |3334## Installation3536### CLI3738```bash39# (Homebrew 6.0+ requires explicit trust for third-party taps)40brew trust xberg-io/tap41brew install xberg-io/tap/html-to-markdown42# or run without a persistent install (the CLI proxy package self-installs the binary):43npx @xberg-io/html-to-markdown-cli --help44uvx --from html-to-markdown-cli html-to-markdown --help45# or download a prebuilt binary from the latest GitHub release:46# https://github.com/xberg-io/html-to-markdown/releases/latest47# or build from source:48cargo install html-to-markdown-cli49```5051### Language SDKs5253```bash54pip install html-to-markdown # Python55npm install @xberg-io/html-to-markdown # TypeScript / Node.js56cargo add html-to-markdown-rs # Rust (features: metadata default; full = all)57gem install html-to-markdown # Ruby58composer require xberg-io/html-to-markdown # PHP59go get github.com/xberg-io/html-to-markdown/packages/go/v3 # Go60dotnet add package XbergIo.HtmlToMarkdown # C#61npm install @xberg-io/html-to-markdown-wasm # WASM62```6364- Java (Maven): `io.xberg:html-to-markdown`65- Elixir: `{:html_to_markdown, "~> 3.8"}` in `mix.exs`66- R: `install.packages("htmltomarkdown", repos = "https://xberg-io.r-universe.dev")`67- C (FFI): pre-built `.so` / `.dll` / `.dylib` from GitHub releases6869## CLI vs SDK — which to use7071- **CLI** — one-shot conversions, shell pipelines, fetching a single URL, ad-hoc metadata/table extraction via `--json | jq`. Flags only for conversion (`FILE` is positional; omit or use `-` for stdin); the only subcommand is `mcp`.72- **SDK** — embedding conversion in application code, batch processing, custom element conversion (visitor pattern, Rust), and tight loops where process spawn overhead matters.73- **MCP server** — `html-to-markdown mcp` exposes `convert_html` and `extract_metadata` as agent tools, so an MCP client can convert an HTML string directly with no shell-out. This plugin auto-registers it; see the **using-the-mcp-server** skill.7475Both share the same `ConversionResult` shape, so output is interchangeable.7677## When to use html-to-markdown vs xberg vs crawlberg7879- **html-to-markdown** — you already have HTML (a string, a file, or a single URL) and want clean Markdown plus structured metadata/tables. No OCR, no document parsing, no crawling.80- **xberg** — you have *documents* (PDF, Office, images, email, archives) and need full text/table/metadata extraction with optional OCR. Use it when the input is not already HTML.81- **crawlberg** — you need to *crawl or scrape many pages*, follow links, and handle JS-rendered sites with a headless-Chrome fallback. It uses html-to-markdown internally for the HTML→Markdown step.8283Rule of thumb: single HTML in → Markdown out = html-to-markdown. Many URLs / a site = crawlberg. Non-HTML documents = xberg.8485## CLI quick start8687```bash88# Convert a file to stdout89html-to-markdown input.html9091# Convert and save92html-to-markdown input.html -o output.md9394# Read from stdin95cat page.html | html-to-markdown9697# Fetch and convert a URL98html-to-markdown --url https://example.com > out.md99100# Full ConversionResult as JSON (content, tables, metadata, images, warnings)101html-to-markdown --json input.html102103# JSON with document structure tree104html-to-markdown --json --include-structure input.html105106# Extraction-only (no Markdown body)107html-to-markdown --json --no-content input.html108109# Aggressive web-page cleanup110html-to-markdown input.html --preprocess --preset aggressive111```112113## SDK quick start114115### Rust116117```rust118use html_to_markdown_rs::convert;119120let result = convert("<h1>Hello World</h1><p>A paragraph.</p>", None)?;121println!("{}", result.content.unwrap_or_default());122```123124### Python125126```python127from html_to_markdown import convert128129result = convert("<h1>Hello World</h1><p>A paragraph.</p>")130print(result.content) # # Hello World\n\nA paragraph.131print(result.metadata) # title, links, headers, …132```133134### TypeScript / Node.js135136```typescript137import { convert } from "@xberg-io/html-to-markdown";138139// Node's convert() returns a ConversionResult object directly.140const result = convert("<h1>Hello World</h1><p>A paragraph.</p>");141console.log(result.content);142```143144## ConversionResult fields145146All languages return the same structure (dict, object, or struct).147148| Field | Description |149| ----- | ----------- |150| `content` | Converted text (Markdown/Djot/plain). `null` only in extraction-only mode. |151| `metadata` | Title, OG, headers, links, images, structured data. |152| `tables` | Tables with `grid` (structured cells) and `markdown` fields. |153| `images` | Extracted inline images (requires inline-image extraction). |154| `document` | Structured document tree when structure extraction is enabled. |155| `warnings` | Non-fatal processing warnings (`message`, `kind`). |156157## Configuration158159All languages expose the same ~30 options. See [references/configuration.md](references/configuration.md) for the complete table. Common ones:160161| Option | Values | Default |162| ------ | ------ | ------- |163| `heading_style` | `atx`, `underlined`, `atx-closed` | `atx` |164| `code_block_style` | `backticks`, `indented`, `tildes` | `backticks` |165| `output_format` | `markdown`, `djot`, `plain` | `markdown` |166| `wrap` / `wrap_width` | bool / 20–500 | off / `80` |167| `autolinks` (SDK) / `--no-autolinks` (CLI) | bool / flag | `true` (on); disable in CLI with `--no-autolinks` |168| preprocessing | `minimal` / `standard` / `aggressive` | off |169170### Rust (builder)171172```rust173use html_to_markdown_rs::{convert, ConversionOptions, HeadingStyle, OutputFormat};174175let options = ConversionOptions::builder()176 .heading_style(HeadingStyle::Atx)177 .output_format(OutputFormat::Markdown)178 .wrap(true)179 .wrap_width(100)180 .build();181let result = convert(html, Some(options))?;182```183184### Python (dataclass)185186```python187from html_to_markdown import convert, ConversionOptions, PreprocessingOptions188189html = "<h1>Title</h1><p>Body text.</p>"190result = convert(191 html,192 ConversionOptions(193 heading_style="atx",194 wrap=True,195 wrap_width=100,196 preprocessing=PreprocessingOptions(enabled=True, preset="aggressive"),197 ),198)199```200201## Metadata extraction202203The library `convert()` extracts metadata by default; the CLI needs `--json --extract-metadata` (see the **extracting-metadata** skill). Fields include `document` (title, description, language, canonical_url, open_graph), `headers`, `links` (with `link_type`), `images`, and `structured_data` (JSON-LD/Microdata/RDFa).204205## Table extraction206207Tables appear in `result.tables`, each with a pre-rendered `markdown` string and a structured cell `grid`. Markdown tables also appear inline in `content`. See the **extracting-tables** skill.208209## Document structure extraction210211Enable structure extraction (`--include-structure` on the CLI, `include_document_structure=true` in SDKs) to get a semantic node tree under `document`. Node types include `heading`, `paragraph`, `list`, `list_item`, `table`, `image`, `code`, `quote`, `group`, `metadata_block`.212213## Common pitfalls2142151. **`convert()` returns a result object, not a string.** Access `.content` for the Markdown text. This holds for Node.js too — `convert()` returns a `ConversionResult` object directly; do **not** `JSON.parse()` it.2162. **`--json` outputs JSON, not Markdown.** Omit `--json` for plain Markdown.2173. **`--include-structure`, `--extract-inline-images`, and `--no-content` require `--json`.**2184. **The conversion CLI is flags-only.** `FILE` is positional; the only subcommand is `mcp` (starts the MCP server).2195. **`--preset`, `--keep-navigation`, `--keep-forms` require `--preprocess`.**220221## Additional resources222223- **[CLI Reference](references/cli-reference.md)** — every flag, JSON shape, exit codes224- **[Configuration Reference](references/configuration.md)** — all 30+ options with defaults225- **[Rust API Reference](references/rust-api.md)** — signatures, builder, feature flags226- **[Python API Reference](references/python-api.md)** — functions, dataclasses, type hints227- **[TypeScript API Reference](references/typescript-api.md)** — functions, interfaces, Buffer support228- **[Other Bindings](references/other-bindings.md)** — Go, Ruby, PHP, Java, C#, Elixir, R, WASM, C FFI229230GitHub: <https://github.com/xberg-io/html-to-markdown>