Markdown Converter
Convert files to Markdown using uvx markitdown — no installation required.
Basic Usage
# Convert to stdout
uvx markitdown input.pdf
# Save to file
uvx markitdown input.pdf -o output.md
uvx markitdown input.docx > output.md
# From stdin
cat input.pdf | uvx markitdown
Supported Formats
- Documents: PDF, Word (.docx), PowerPoint (.pptx), Excel (.xlsx, .xls)
- Web/Data: HTML, CSV, JSON, XML
- Media: Images (EXIF + OCR), Audio (EXIF + transcription)
- Other: ZIP (iterates contents), YouTube URLs, EPub
Options
-o OUTPUT # Output file
-x EXTENSION # Hint file extension (for stdin)
-m MIME_TYPE # Hint MIME type
-c CHARSET # Hint charset (e.g., UTF-8)
-d # Use Azure Document Intelligence
-e ENDPOINT # Document Intelligence endpoint
--use-plugins # Enable 3rd-party plugins
--list-plugins # Show installed plugins
Examples
# Convert Word document
uvx markitdown report.docx -o report.md
# Convert Excel spreadsheet
uvx markitdown data.xlsx > data.md
# Convert PowerPoint presentation
uvx markitdown slides.pptx -o slides.md
# Convert with file type hint (for stdin)
cat document | uvx markitdown -x .pdf > output.md
# Use Azure Document Intelligence for better PDF extraction
uvx markitdown scan.pdf -d -e "https://your-resource.cognitiveservices.azure.com/"
Notes
- Output preserves document structure: headings, tables, lists, links
- First run caches dependencies; subsequent runs are faster
- For complex PDFs with poor extraction, use
-d with Azure Document Intelligence
When To Use
- When converting PDF, Word, PowerPoint, Excel, HTML, or other supported formats to Markdown
- When preparing documents for LLM processing or text analysis pipelines
- When extracting text content from images (OCR), audio (transcription), or ZIP archives
- When converting YouTube URLs or EPubs to readable Markdown
Boundaries
- Not for Markdown-to-other-format conversion (e.g., Markdown to PDF)
- Not for editing or reformatting existing Markdown files
- Not for web scraping or crawling — operates on local files, stdin, or single URLs
- Skip Azure Document Intelligence (
-d) unless standard PDF extraction is insufficient
Output
- Markdown text preserving document structure: headings, tables, lists, and links
- Written to stdout by default, or to a file with
-o flag
- One Markdown output per input file or URL
Verification
- Output contains recognizable document structure (headings, tables) matching the source
uvx is available in the environment (no pre-installation required)
- For Azure Document Intelligence, endpoint is reachable and credentials are valid
- Output file is non-empty when
-o is used
Sibling skills
fetchmd — sister format-to-markdown tool. Use this skill for file inputs (PDF, .docx, .pptx, .xlsx, image OCR); use fetchmd for web inputs (URLs or HTML).
1---2name: markdown-converter3description: Convert documents and files to Markdown using markitdown. Use when converting PDF, Word (.docx), PowerPoint (.pptx), Excel (.xlsx, .xls), HTML, CSV, JSON, XML, images (with EXIF/OCR), audio (with transcription), ZIP archives, YouTube URLs, or EPubs to Markdown format for LLM processing or text analysis.4---56# Markdown Converter78Convert files to Markdown using `uvx markitdown` — no installation required.910## Basic Usage1112```bash13# Convert to stdout14uvx markitdown input.pdf1516# Save to file17uvx markitdown input.pdf -o output.md18uvx markitdown input.docx > output.md1920# From stdin21cat input.pdf | uvx markitdown22```2324## Supported Formats2526- **Documents**: PDF, Word (.docx), PowerPoint (.pptx), Excel (.xlsx, .xls)27- **Web/Data**: HTML, CSV, JSON, XML28- **Media**: Images (EXIF + OCR), Audio (EXIF + transcription)29- **Other**: ZIP (iterates contents), YouTube URLs, EPub3031## Options3233```bash34-o OUTPUT # Output file35-x EXTENSION # Hint file extension (for stdin)36-m MIME_TYPE # Hint MIME type37-c CHARSET # Hint charset (e.g., UTF-8)38-d # Use Azure Document Intelligence39-e ENDPOINT # Document Intelligence endpoint40--use-plugins # Enable 3rd-party plugins41--list-plugins # Show installed plugins42```4344## Examples4546```bash47# Convert Word document48uvx markitdown report.docx -o report.md4950# Convert Excel spreadsheet51uvx markitdown data.xlsx > data.md5253# Convert PowerPoint presentation54uvx markitdown slides.pptx -o slides.md5556# Convert with file type hint (for stdin)57cat document | uvx markitdown -x .pdf > output.md5859# Use Azure Document Intelligence for better PDF extraction60uvx markitdown scan.pdf -d -e "https://your-resource.cognitiveservices.azure.com/"61```6263## Notes6465- Output preserves document structure: headings, tables, lists, links66- First run caches dependencies; subsequent runs are faster67- For complex PDFs with poor extraction, use `-d` with Azure Document Intelligence6869## When To Use7071- When converting PDF, Word, PowerPoint, Excel, HTML, or other supported formats to Markdown72- When preparing documents for LLM processing or text analysis pipelines73- When extracting text content from images (OCR), audio (transcription), or ZIP archives74- When converting YouTube URLs or EPubs to readable Markdown7576## Boundaries7778- Not for Markdown-to-other-format conversion (e.g., Markdown to PDF)79- Not for editing or reformatting existing Markdown files80- Not for web scraping or crawling — operates on local files, stdin, or single URLs81- Skip Azure Document Intelligence (`-d`) unless standard PDF extraction is insufficient8283## Output8485- Markdown text preserving document structure: headings, tables, lists, and links86- Written to stdout by default, or to a file with `-o` flag87- One Markdown output per input file or URL8889## Verification9091- Output contains recognizable document structure (headings, tables) matching the source92- `uvx` is available in the environment (no pre-installation required)93- For Azure Document Intelligence, endpoint is reachable and credentials are valid94- Output file is non-empty when `-o` is used9596## Sibling skills9798- `fetchmd` — sister format-to-markdown tool. Use this skill for *file* inputs (PDF, .docx, .pptx, .xlsx, image OCR); use `fetchmd` for *web* inputs (URLs or HTML).