PDF to Markdown Converter
Faithfully convert PDF content to markdown format using visual PDF capability or image OCR.
Process
- Read every page of the PDF using visual PDF understanding or image OCR
- Transcribe all content faithfully — this is conversion, not summarisation
- Preserve document structure: headings, lists, tables, formatting
- Output as markdown file matching the PDF filename
Conversion Rules
Complete transcription: Include ALL text content. Do not summarise or omit.
Preserve structure:
- Headings →
#,##,###(match hierarchy from source) - Bold text →
**bold** - Italic text →
*italic* - Bullet lists →
-or* - Numbered lists →
1.,2., etc.
Tables: Convert to markdown table format:
| Column 1 | Column 2 |
|----------|----------|
| Data | Data |
Equations/formulas: Use LaTeX notation:
- Inline:
$E = mc^2$ - Block:
$$F = ma$$
Diagrams/images: Describe in a blockquote with [DIAGRAM] prefix:
> [DIAGRAM]: Description of what the diagram shows, including labels and key information.
Remove repeated footers: Omit recurring footer content such as brand names, website links, copyright notices, page numbers, and other boilerplate that appears on multiple pages.
Page breaks: Optionally insert --- between major sections if helpful for navigation.
Output
Markdown file named to match the source PDF (e.g., Topic Name.md).
Quality Checklist
- Every page processed
- All text content included (no omissions)
- Headings hierarchy preserved
- Tables correctly formatted
- Equations in LaTeX notation
- Diagrams described with key details
- Document structure matches original