Document DOCX Skill - Quick Reference
This skill enables creation, editing, and analysis of .docx files for reports, contracts, proposals, documentation, and template-driven outputs.
Modern best practices (2026):
- Prefer templates + styles over manual formatting.
- Treat
.docx as the editable source; treat PDF as a release artifact.
- If distributing externally, include basic accessibility hygiene (headings, table headers, alt text).
Quick Reference
| Task |
Tool/Library |
Language |
When to Use |
| Create DOCX |
python-docx |
Python |
Reports, contracts, proposals |
| Create DOCX |
docx |
Node.js |
Server-side document generation |
| Convert to HTML |
mammoth.js |
Node.js |
Web display, content extraction |
| Parse DOCX |
python-docx |
Python |
Extract text, tables, metadata |
| Template fill |
docxtpl |
Python |
Mail merge, template-based generation |
| Review workflow |
Word compare, comments/highlights |
Any |
Human review without OOXML surgery |
| Tracked changes |
OOXML inspection, docx4j/OpenXML SDK/Aspose |
Any |
True redlines or parsing tracked changes |
Tool Selection
- Prefer
docxtpl when non-developers must edit layout/design in Word.
- Prefer
python-docx for structural edits (paragraphs/tables/headers/footers) when formatting complexity is moderate.
- Prefer
docx (Node.js) for server-side generation in TypeScript-heavy stacks.
- Prefer
mammoth for text-first extraction or DOCX-to-HTML (best effort; may drop some layout fidelity).
Known Limits (Plan Around These)
.doc (legacy) is not supported by these libraries; convert to .docx first (e.g., LibreOffice).
python-docx cannot reliably create true tracked changes; use Word compare or specialized OOXML tooling.
- Tables of Contents and many fields are placeholders until opened/updated in Word.
Core Operations
Create Document (Python - python-docx)
from docx import Document
from docx.shared import Inches, Pt
from docx.enum.text import WD_ALIGN_PARAGRAPH
doc = Document()
# Title
title = doc.add_heading('Document Title', 0)
title.alignment = WD_ALIGN_PARAGRAPH.CENTER
# Paragraph with formatting
para = doc.add_paragraph()
run = para.add_run('Bold and ')
run.bold = True
run = para.add_run('italic text.')
run.italic = True
# Table
table = doc.add_table(rows=3, cols=3)
table.style = 'Table Grid'
for i, row in enumerate(table.rows):
for j, cell in enumerate(row.cells):
cell.text = f'Row {i+1}, Col {j+1}'
# Image
doc.add_picture('image.png', width=Inches(4))
# Save
doc.save('output.docx')
Create Document (Node.js - docx)
import { Document, Packer, Paragraph, TextRun, Table, TableRow, TableCell } from 'docx';
import * as fs from 'fs';
const doc = new Document({
sections: [{
properties: {},
children: [
new Paragraph({
children: [
new TextRun({ text: 'Bold text', bold: true }),
new TextRun({ text: ' and normal text.' }),
],
}),
new Table({
rows: [
new TableRow({
children: [
new TableCell({ children: [new Paragraph('Cell 1')] }),
new TableCell({ children: [new Paragraph('Cell 2')] }),
],
}),
],
}),
],
}],
});
Packer.toBuffer(doc).then((buffer) => {
fs.writeFileSync('output.docx', buffer);
});
Template-Based Generation (Python - docxtpl)
from docxtpl import DocxTemplate
doc = DocxTemplate('template.docx')
context = {
'company_name': 'Acme Corp',
'date': '2025-01-15',
'items': [
{'name': 'Widget A', 'price': 100},
{'name': 'Widget B', 'price': 200},
]
}
doc.render(context)
doc.save('filled_template.docx')
Extract Content (Python - python-docx)
from docx import Document
doc = Document('input.docx')
# Extract all text
full_text = []
for para in doc.paragraphs:
full_text.append(para.text)
# Extract tables
for table in doc.tables:
for row in table.rows:
row_data = [cell.text for cell in row.cells]
print(row_data)
Styling Reference
| Element |
Python Method |
Node.js Class |
| Heading 1 |
add_heading(text, 1) |
HeadingLevel.HEADING_1 |
| Bold |
run.bold = True |
TextRun({ bold: true }) |
| Italic |
run.italic = True |
TextRun({ italics: true }) |
| Font size |
run.font.size = Pt(12) |
TextRun({ size: 24 }) (half-points) |
| Alignment |
WD_ALIGN_PARAGRAPH.CENTER |
AlignmentType.CENTER |
| Page break |
doc.add_page_break() |
new PageBreak() |
Do / Avoid (Dec 2025)
Do
- Use consistent heading levels and a table of contents for long docs.
- Capture decisions and action items with owners and due dates.
- Store docs in a versioned, searchable system.
Avoid
- Manual formatting instead of styles (breaks consistency).
- Docs with no owner or review cadence (stale quickly).
- Copy/pasting without updating definitions and links.
Output Quality Checklist
- Structure: consistent heading hierarchy, styles, and (when needed) an auto-generated table of contents.
- Decisions: decisions/actions captured with owner + due date (not buried in prose).
- Versioning: doc ID + version + change summary; review cadence defined.
- Accessibility hygiene: headings/reading order are correct; table headers are marked; alt text for non-decorative images.
- Reuse: use
assets/doc-template-pack.md for decision logs and recurring doc types.
Optional: AI / Automation
Use only when explicitly requested and policy-compliant.
- Summarize meeting notes into decisions/actions; humans verify accuracy.
- Draft first-pass docs from outlines; do not invent facts or quotes.
Navigation
Resources
- references/docx-patterns.md - Advanced formatting, styles, headers/footers
- references/template-workflows.md - Mail merge, batch generation
- references/tracked-changes.md - Tracked changes: what is feasible, and what is not
- data/sources.json - Library documentation links
Scripts
scripts/docx_inspect_ooxml.py - Dependency-free OOXML inspection (including tracked changes signals)
scripts/docx_extract.py - Extract text/tables to JSON (requires python-docx)
scripts/docx_render_template.py - Render a docxtpl template (requires docxtpl)
scripts/docx_to_html.mjs - Convert .docx to HTML (requires mammoth)
Templates
- assets/report-template.md - Standard report structure
- assets/contract-template.md - Legal document structure
- assets/doc-template-pack.md - Decision log, meeting notes, changelog templates
Related Skills
1---2name: document-docx3description: Create, edit, and analyze Microsoft Word .docx files (reports, contracts, proposals) with styles, tables, headers/footers, template filling, content extraction, and conversion to HTML; support review workflows (comments/highlights) and inspect tracked changes via OOXML when needed using Python/Node.js (python-docx, docxtpl, mammoth.js, docx).4---5
6# Document DOCX Skill - Quick Reference
7
8This skill enables creation, editing, and analysis of `.docx` files for reports, contracts, proposals, documentation, and template-driven outputs.
9
10Modern best practices (2026):
11- Prefer templates + styles over manual formatting.
12- Treat `.docx` as the editable source; treat PDF as a release artifact.
13- If distributing externally, include basic accessibility hygiene (headings, table headers, alt text).
14
15## Quick Reference
16
17| Task | Tool/Library | Language | When to Use |
18|------|--------------|----------|-------------|
19| Create DOCX | python-docx | Python | Reports, contracts, proposals |
20| Create DOCX | docx | Node.js | Server-side document generation |
21| Convert to HTML | mammoth.js | Node.js | Web display, content extraction |
22| Parse DOCX | python-docx | Python | Extract text, tables, metadata |
23| Template fill | docxtpl | Python | Mail merge, template-based generation |
24| Review workflow | Word compare, comments/highlights | Any | Human review without OOXML surgery |
25| Tracked changes | OOXML inspection, docx4j/OpenXML SDK/Aspose | Any | True redlines or parsing tracked changes |
26
27## Tool Selection
28
29- Prefer `docxtpl` when non-developers must edit layout/design in Word.
30- Prefer `python-docx` for structural edits (paragraphs/tables/headers/footers) when formatting complexity is moderate.
31- Prefer `docx` (Node.js) for server-side generation in TypeScript-heavy stacks.
32- Prefer `mammoth` for text-first extraction or DOCX-to-HTML (best effort; may drop some layout fidelity).
33
34## Known Limits (Plan Around These)
35
36- `.doc` (legacy) is not supported by these libraries; convert to `.docx` first (e.g., LibreOffice).
37- `python-docx` cannot reliably create true tracked changes; use Word compare or specialized OOXML tooling.
38- Tables of Contents and many fields are placeholders until opened/updated in Word.
39
40## Core Operations
41
42### Create Document (Python - python-docx)
43
44```python
45from docx import Document
46from docx.shared import Inches, Pt
47from docx.enum.text import WD_ALIGN_PARAGRAPH
48
49doc = Document()
50
51# Title
52title = doc.add_heading('Document Title', 0)
53title.alignment = WD_ALIGN_PARAGRAPH.CENTER
54
55# Paragraph with formatting
56para = doc.add_paragraph()
57run = para.add_run('Bold and ')
58run.bold = True
59run = para.add_run('italic text.')
60run.italic = True
61
62# Table
63table = doc.add_table(rows=3, cols=3)
64table.style = 'Table Grid'
65for i, row in enumerate(table.rows):
66 for j, cell in enumerate(row.cells):
67 cell.text = f'Row {i+1}, Col {j+1}'
68
69# Image
70doc.add_picture('image.png', width=Inches(4))
71
72# Save
73doc.save('output.docx')
74```
75
76### Create Document (Node.js - docx)
77
78```typescript
79import { Document, Packer, Paragraph, TextRun, Table, TableRow, TableCell } from 'docx';
80import * as fs from 'fs';
81
82const doc = new Document({
83 sections: [{
84 properties: {},
85 children: [
86 new Paragraph({
87 children: [
88 new TextRun({ text: 'Bold text', bold: true }),
89 new TextRun({ text: ' and normal text.' }),
90 ],
91 }),
92 new Table({
93 rows: [
94 new TableRow({
95 children: [
96 new TableCell({ children: [new Paragraph('Cell 1')] }),
97 new TableCell({ children: [new Paragraph('Cell 2')] }),
98 ],
99 }),
100 ],
101 }),
102 ],
103 }],
104});
105
106Packer.toBuffer(doc).then((buffer) => {
107 fs.writeFileSync('output.docx', buffer);
108});
109```
110
111### Template-Based Generation (Python - docxtpl)
112
113```python
114from docxtpl import DocxTemplate
115
116doc = DocxTemplate('template.docx')
117context = {
118 'company_name': 'Acme Corp',
119 'date': '2025-01-15',
120 'items': [
121 {'name': 'Widget A', 'price': 100},
122 {'name': 'Widget B', 'price': 200},
123 ]
124}
125doc.render(context)
126doc.save('filled_template.docx')
127```
128
129### Extract Content (Python - python-docx)
130
131```python
132from docx import Document
133
134doc = Document('input.docx')
135
136# Extract all text
137full_text = []
138for para in doc.paragraphs:
139 full_text.append(para.text)
140
141# Extract tables
142for table in doc.tables:
143 for row in table.rows:
144 row_data = [cell.text for cell in row.cells]
145 print(row_data)
146```
147
148## Styling Reference
149
150| Element | Python Method | Node.js Class |
151|---------|---------------|---------------|
152| Heading 1 | `add_heading(text, 1)` | `HeadingLevel.HEADING_1` |
153| Bold | `run.bold = True` | `TextRun({ bold: true })` |
154| Italic | `run.italic = True` | `TextRun({ italics: true })` |
155| Font size | `run.font.size = Pt(12)` | `TextRun({ size: 24 })` (half-points) |
156| Alignment | `WD_ALIGN_PARAGRAPH.CENTER` | `AlignmentType.CENTER` |
157| Page break | `doc.add_page_break()` | `new PageBreak()` |
158
159## Do / Avoid (Dec 2025)
160
161### Do
162
163- Use consistent heading levels and a table of contents for long docs.
164- Capture decisions and action items with owners and due dates.
165- Store docs in a versioned, searchable system.
166
167### Avoid
168
169- Manual formatting instead of styles (breaks consistency).
170- Docs with no owner or review cadence (stale quickly).
171- Copy/pasting without updating definitions and links.
172
173## Output Quality Checklist
174
175- Structure: consistent heading hierarchy, styles, and (when needed) an auto-generated table of contents.
176- Decisions: decisions/actions captured with owner + due date (not buried in prose).
177- Versioning: doc ID + version + change summary; review cadence defined.
178- Accessibility hygiene: headings/reading order are correct; table headers are marked; alt text for non-decorative images.
179- Reuse: use `assets/doc-template-pack.md` for decision logs and recurring doc types.
180
181## Optional: AI / Automation
182
183Use only when explicitly requested and policy-compliant.
184
185- Summarize meeting notes into decisions/actions; humans verify accuracy.
186- Draft first-pass docs from outlines; do not invent facts or quotes.
187
188## Navigation
189
190**Resources**
191- [references/docx-patterns.md](references/docx-patterns.md) - Advanced formatting, styles, headers/footers
192- [references/template-workflows.md](references/template-workflows.md) - Mail merge, batch generation
193- [references/tracked-changes.md](references/tracked-changes.md) - Tracked changes: what is feasible, and what is not
194- [data/sources.json](data/sources.json) - Library documentation links
195
196**Scripts**
197- `scripts/docx_inspect_ooxml.py` - Dependency-free OOXML inspection (including tracked changes signals)
198- `scripts/docx_extract.py` - Extract text/tables to JSON (requires `python-docx`)
199- `scripts/docx_render_template.py` - Render a `docxtpl` template (requires `docxtpl`)
200- `scripts/docx_to_html.mjs` - Convert `.docx` to HTML (requires `mammoth`)
201
202**Templates**
203- [assets/report-template.md](assets/report-template.md) - Standard report structure
204- [assets/contract-template.md](assets/contract-template.md) - Legal document structure
205- [assets/doc-template-pack.md](assets/doc-template-pack.md) - Decision log, meeting notes, changelog templates
206
207**Related Skills**
208- [../document-pdf/SKILL.md](../document-pdf/SKILL.md) - PDF generation and conversion
209- [../docs-codebase/SKILL.md](../docs-codebase/SKILL.md) - Technical writing patterns