PPTX creation, editing, and analysis
Overview
Create, edit, or analyze the contents of .pptx files when requested. A .pptx file is essentially a ZIP archive containing XML files and other resources. Different tools and workflows are available for different tasks.
Reading and analyzing content
Text extraction
To read just the text content of a presentation, convert the document to markdown:
# Convert document to markdown
python -m markitdown path-to-file.pptx
Raw XML access
Use raw XML access for: comments, speaker notes, slide layouts, animations, design elements, and complex formatting. To access these features, unpack a presentation and read its raw XML contents.
Unpacking a file
python ooxml/scripts/unpack.py <office_file> <output_dir>
Note: The unpack.py script is located at skills/pptx/ooxml/scripts/unpack.py relative to the project root. If the script doesn't exist at this path, use find . -name "unpack.py" to locate it.
Key file structures
ppt/presentation.xml - Main presentation metadata and slide references
ppt/slides/slide{N}.xml - Individual slide contents (slide1.xml, slide2.xml, etc.)
ppt/notesSlides/notesSlide{N}.xml - Speaker notes for each slide
ppt/comments/modernComment_*.xml - Comments for specific slides
ppt/slideLayouts/ - Layout templates for slides
ppt/slideMasters/ - Master slide templates
ppt/theme/ - Theme and styling information
ppt/media/ - Images and other media files
Typography and color extraction
To emulate example designs, analyze the presentation's typography and colors first using the methods below:
- Read theme file: Check
ppt/theme/theme1.xml for colors (<a:clrScheme>) and fonts (<a:fontScheme>)
- Sample slide content: Examine
ppt/slides/slide1.xml for actual font usage (<a:rPr>) and colors
- Search for patterns: Use grep to find color (
<a:solidFill>, <a:srgbClr>) and font references across all XML files
Creating a new PowerPoint presentation without a template
When creating a new PowerPoint presentation from scratch, use the html2pptx workflow to convert HTML slides to PowerPoint with accurate positioning.
Workflow
MANDATORY - READ ENTIRE FILE NOW: Read html2pptx.md completely from start to finish. NEVER set any range limits when reading this file. Read the full file content for detailed syntax, critical formatting rules, and best practices before proceeding with presentation creation.
PREREQUISITE - Install html2pptx library:
- Check and install if needed:
npm list -g @ant/html2pptx || npm install -g skills/pptx/html2pptx.tgz
- Note: If you see "Cannot find module '@ant/html2pptx'" error later, the package isn't installed
CRITICAL: Plan the presentation
- Plan the shared aspects of the presentation. Describe the tone of the presentation's content and the colors and typography that should be used in the presentation.
- Write a DETAILED outline of the presentation
- For each slide, describe the slide's layout and contents
- For each slide, write presenter notes (1 to 3 sentences per slide)
CRITICAL: Set CSS variables
- In a shared
.css file, override CSS variables to use on each slide for colors, typography, and spacing. DO NOT create classes in this file.
Create an HTML file for each slide with proper dimensions (e.g., 960px × 540px for 16:9)
- Recall the outline, layout/content description, and speaker notes you wrote for this slide in Step 3. Think out loud how to best apply them to this slide.
- Embed the contents of the shared
.css file in a <style> element
- Use
<p>, <h1>-<h6>, <ul>, <ol> for all text content
- IMPORTANT: Use CSS variables for colors, typography, and spacing
- IMPORTANT: Use
row col and fit classes for layout INSTEAD OF flexbox
- Use
class="placeholder" for areas where charts/tables will be added (render with gray background for visibility)
- CSS gradients: Use
linear-gradient() or radial-gradient() in CSS on block element backgrounds - automatically converted to PowerPoint
- Background images: Use
background-image: url(...) CSS property on block elements
- Block elements: Use
<div>, <section>, <header>, <footer>, <main>, <article>, <nav>, <aside> for containers with styling (all behave identically)
- Icons: Use inline SVG format or reference SVG files - SVG elements are automatically converted to images in PowerPoint
- Text balancing:
<h1> and <h2> elements are automatically balanced. Use data-balance attribute on other elements to auto-balance line lengths for better typography
- Layout: For slides with charts/tables/images, use either full-slide layout or two-column layout for better readability
Create and run a JavaScript file using the html2pptx library to convert HTML slides to PowerPoint and save the presentation
Run with: NODE_PATH="$(npm root -g)" node your-script.js 2>&1
Use the html2pptx function to process each HTML file
Add charts and tables to placeholder areas using PptxGenJS API
Save the presentation using pptx.writeFile()
⚠️ CRITICAL: Your script MUST follow this example structure. Think aloud before writing the script to make sure that you correctly use the APIs. Do NOT call pptx.addSlide.
const pptxgen = require("pptxgenjs");
const { html2pptx } = require("@ant/html2pptx");
// Create a new pptx presentation
const pptx = new pptxgen();
pptx.layout = "LAYOUT_16x9"; // Must match HTML body dimensions
// Add an HTML-only slide
await html2pptx("slide1.html", pptx);
// Add a HTML slide with chart placeholders
const { slide: slide2, placeholders } = await html2pptx("slide2.html", pptx);
slide.addChart(pptx.charts.LINE, chartData, placeholders[0]);
// Save the presentation
await pptx.writeFile("output.pptx");
Visual validation: Generate thumbnails and inspect for layout issues
- Create thumbnail grid:
python scripts/thumbnail.py output.pptx workspace/thumbnails --cols 4
- Read and carefully examine the thumbnail image for:
- Text cutoff: Text being cut off by header bars, shapes, or slide edges
- Text overlap: Text overlapping with other text or shapes
- Positioning issues: Content too close to slide boundaries or other elements
- Contrast issues: Insufficient contrast between text and backgrounds
- If issues found, adjust HTML margins/spacing/colors and regenerate the presentation
- Repeat until all slides are visually correct
Editing an existing PowerPoint presentation
To edit slides in an existing PowerPoint presentation, work with the raw Office Open XML (OOXML) format. This involves unpacking the .pptx file, editing the XML content, and repacking it.
Workflow
- MANDATORY - READ ENTIRE FILE: Read
ooxml.md (~500 lines) completely from start to finish. NEVER set any range limits when reading this file. Read the full file content for detailed guidance on OOXML structure and editing workflows before any presentation editing.
- Unpack the presentation:
python ooxml/scripts/unpack.py <office_file> <output_dir>
- Edit the XML files (primarily
ppt/slides/slide{N}.xml and related files)
- CRITICAL: Validate immediately after each edit and fix any validation errors before proceeding:
python ooxml/scripts/validate.py <dir> --original <file>
- Pack the final presentation:
python ooxml/scripts/pack.py <input_directory> <office_file>
Creating a new PowerPoint presentation using a template
To create a presentation that follows an existing template's design, duplicate and re-arrange template slides before replacing placeholder content.
Workflow
Extract template text AND create visual thumbnail grid:
- Extract text:
python -m markitdown template.pptx > template-content.md
- Read
template-content.md: Read the entire file to understand the contents of the template presentation. NEVER set any range limits when reading this file.
- Create thumbnail grids:
python scripts/thumbnail.py template.pptx
- See Creating Thumbnail Grids section for more details
Analyze template and save inventory to a file:
Visual Analysis: Review thumbnail grid(s) to understand slide layouts, design patterns, and visual structure
Create and save a template inventory file at template-inventory.md containing:
# Template Inventory Analysis
**Total Slides: [count]**
**IMPORTANT: Slides are 0-indexed (first slide = 0, last slide = count-1)**
## [Category Name]
- Slide 0: [Layout code if available] - Description/purpose
- Slide 1: [Layout code] - Description/purpose
- Slide 2: [Layout code] - Description/purpose
[... EVERY slide must be listed individually with its index ...]
Using the thumbnail grid: Reference the visual thumbnails to identify:
- Layout patterns (title slides, content layouts, section dividers)
- Image placeholder locations and counts
- Design consistency across slide groups
- Visual hierarchy and structure
This inventory file is REQUIRED for selecting appropriate templates in the next step
Create presentation outline based on template inventory:
- Review available templates from step 2.
- Choose an intro or title template for the first slide. This should be one of the first templates.
- Choose safe, text-based layouts for the other slides.
- CRITICAL: Match layout structure to actual content:
- Single-column layouts: Use for unified narrative or single topic
- Two-column layouts: Use ONLY when there are exactly 2 distinct items/concepts
- Three-column layouts: Use ONLY when there are exactly 3 distinct items/concepts
- Image + text layouts: Use ONLY when there are actual images to insert
- Quote layouts: Use ONLY for actual quotes from people (with attribution), never for emphasis
- Never use layouts with more placeholders than available content
- With 2 items, avoid forcing them into a 3-column layout
- With 4+ items, consider breaking into multiple slides or using a list format
- Count actual content pieces BEFORE selecting the layout
- Verify each placeholder in the chosen layout will be filled with meaningful content
- Select one option representing the best layout for each content section.
- Save
outline.md with content AND template mapping that leverages available designs
- Example template mapping:
# Template slides to use (0-based indexing)
# WARNING: Verify indices are within range! Template with 73 slides has indices 0-72
# Mapping: slide numbers from outline -> template slide indices
template_mapping = [
0, # Use slide 0 (Title/Cover)
34, # Use slide 34 (B1: Title and body)
34, # Use slide 34 again (duplicate for second B1)
50, # Use slide 50 (E1: Quote)
54, # Use slide 54 (F2: Closing + Text)
]
Duplicate, reorder, and delete slides using rearrange.py:
- Use the
scripts/rearrange.py script to create a new presentation with slides in the desired order:python scripts/rearrange.py template.pptx working.pptx 0,34,34,50,52
- The script handles duplicating repeated slides, deleting unused slides, and reordering automatically
- Slide indices are 0-based (first slide is 0, second is 1, etc.)
- The same slide index can appear multiple times to duplicate that slide
Extract ALL text using the inventory.py script:
Run inventory extraction:
python scripts/inventory.py working.pptx text-inventory.json
Read text-inventory.json: Read the entire text-inventory.json file to understand all shapes and their properties. NEVER set any range limits when reading this file.
The inventory JSON structure:
{
"slide-0": {
"shape-0": {
"placeholder_type": "TITLE", // or null for non-placeholders
"left": 1.5, // position in inches
"top": 2.0,
"width": 7.5,
"height": 1.2,
"paragraphs": [
{
"text": "Paragraph text",
// Optional properties (only included when non-default):
"bullet": true, // explicit bullet detected
"level": 0, // only included when bullet is true
"alignment": "CENTER", // CENTER, RIGHT (not LEFT)
"space_before": 10.0, // space before paragraph in points
"space_after": 6.0, // space after paragraph in points
"line_spacing": 22.4, // line spacing in points
"font_name": "Arial", // from first run
"font_size": 14.0, // in points
"bold": true,
"italic": false,
"underline": false,
"color": "FF0000" // RGB color
}
]
}
}
}
Key features:
- Slides: Named as "slide-0", "slide-1", etc.
- Shapes: Ordered by visual position (top-to-bottom, left-to-right) as "shape-0", "shape-1", etc.
- Placeholder types: TITLE, CENTER_TITLE, SUBTITLE, BODY, OBJECT, or null
- Default font size:
default_font_size in points extracted from layout placeholders (when available)
- Slide numbers are filtered: Shapes with SLIDE_NUMBER placeholder type are automatically excluded from inventory
- Bullets: When
bullet: true, level is always included (even if 0)
- Spacing:
space_before, space_after, and line_spacing in points (only included when set)
- Colors:
color for RGB (e.g., "FF0000"), theme_color for theme colors (e.g., "DARK_1")
- Properties: Only non-default values are included in the output
Generate replacement text and save the data to a JSON file
Based on the text inventory from the previous step:
- CRITICAL: First verify which shapes exist in the inventory - only reference shapes that are actually present
- VALIDATION: The replace.py script validates that all shapes in the replacement JSON exist in the inventory
- Referencing a non-existent shape produces an error showing available shapes
- Referencing a non-existent slide produces an error indicating the slide doesn't exist
- All validation errors are shown at once before the script exits
- IMPORTANT: The replace.py script uses inventory.py internally to identify ALL text shapes
- AUTOMATIC CLEARING: ALL text shapes from the inventory are cleared unless "paragraphs" are provided for them
- Add a "paragraphs" field to shapes that need content (not "replacement_paragraphs")
- Shapes without "paragraphs" in the replacement JSON have their text cleared automatically
- Paragraphs with bullets are automatically left aligned. Avoid setting the
alignment property when "bullet": true
- Generate appropriate replacement content for placeholder text
- Use shape size to determine appropriate content length
- CRITICAL: Include paragraph properties from the original inventory - don't just provide text
- IMPORTANT: When bullet: true, do NOT include bullet symbols (•, -, *) in text - they're added automatically
- ESSENTIAL FORMATTING RULES:
- Headers/titles should typically have
"bold": true
- List items should have
"bullet": true, "level": 0 (level is required when bullet is true)
- Preserve any alignment properties (e.g.,
"alignment": "CENTER" for centered text)
- Include font properties when different from default (e.g.,
"font_size": 14.0, "font_name": "Lora")
- Colors: Use
"color": "FF0000" for RGB or "theme_color": "DARK_1" for theme colors
- The replacement script expects properly formatted paragraphs, not just text strings
- Overlapping shapes: Prefer shapes with larger default_font_size or more appropriate placeholder_type
- Save the updated inventory with replacements to
replacement-text.json
- WARNING: Different template layouts have different shape counts - always check the actual inventory before creating replacements
Example paragraphs field showing proper formatting:
"paragraphs": [
{
"text": "New presentation title text",
"alignment": "CENTER",
"bold": true
},
{
"text": "Section Header",
"bold": true
},
{
"text": "First bullet point without bullet symbol",
"bullet": true,
"level": 0
},
{
"text": "Red colored text",
"color": "FF0000"
},
{
"text": "Theme colored text",
"theme_color": "DARK_1"
},
{
"text": "Regular paragraph text without special formatting"
}
]
Shapes not listed in the replacement JSON are automatically cleared:
{
"slide-0": {
"shape-0": {
"paragraphs": [...] // This shape gets new text
}
// shape-1 and shape-2 from inventory will be cleared automatically
}
}
Common formatting patterns for presentations:
- Title slides: Bold text, sometimes centered
- Section headers within slides: Bold text
- Bullet lists: Each item needs
"bullet": true, "level": 0
- Body text: Usually no special properties needed
- Quotes: May have special alignment or font properties
Apply replacements using the replace.py script
python scripts/replace.py working.pptx replacement-text.json output.pptx
The script will:
- First extract the inventory of ALL text shapes using functions from inventory.py
- Validate that all shapes in the replacement JSON exist in the inventory
- Clear text from ALL shapes identified in the inventory
- Apply new text only to shapes with "paragraphs" defined in the replacement JSON
- Preserve formatting by applying paragraph properties from the JSON
- Handle bullets, alignment, font properties, and colors automatically
- Save the updated presentation
Example validation errors:
ERROR: Invalid shapes in replacement JSON:
- Shape 'shape-99' not found on 'slide-0'. Available shapes: shape-0, shape-1, shape-4
- Slide 'slide-999' not found in inventory
ERROR: Replacement text made overflow worse in these shapes:
- slide-0/shape-2: overflow worsened by 1.25" (was 0.00", now 1.25")
Creating Thumbnail Grids
To create visual thumbnail grids of PowerPoint slides for quick analysis and reference:
python scripts/thumbnail.py template.pptx [output_prefix]
Features:
- Creates:
thumbnails.jpg (or thumbnails-1.jpg, thumbnails-2.jpg, etc. for large decks)
- Default: 5 columns, max 30 slides per grid (5×6)
- Custom prefix:
python scripts/thumbnail.py template.pptx my-grid
- Note: The output prefix should include the path if you want output in a specific directory (e.g.,
workspace/my-grid)
- Adjust columns:
--cols 4 (range: 3-6, affects slides per grid)
- Grid limits: 3 cols = 12 slides/grid, 4 cols = 20, 5 cols = 30, 6 cols = 42
- Slides are zero-indexed (Slide 0, Slide 1, etc.)
Use cases:
- Template analysis: Quickly understand slide layouts and design patterns
- Content review: Visual overview of entire presentation
- Navigation reference: Find specific slides by their visual appearance
- Quality check: Verify all slides are properly formatted
Examples:
# Basic usage
python scripts/thumbnail.py presentation.pptx
# Combine options: custom name, columns
python scripts/thumbnail.py template.pptx analysis --cols 4
Converting Slides to Images
To visually analyze PowerPoint slides, convert them to images using a two-step process:
Convert PPTX to PDF:
soffice --headless --convert-to pdf template.pptx
Convert PDF pages to JPEG images:
pdftoppm -jpeg -r 150 template.pdf slide
This creates files like slide-1.jpg, slide-2.jpg, etc.
Options:
-r 150: Sets resolution to 150 DPI (adjust for quality/size balance)
-jpeg: Output JPEG format (use -png for PNG if preferred)
-f N: First page to convert (e.g., -f 2 starts from page 2)
-l N: Last page to convert (e.g., -l 5 stops at page 5)
slide: Prefix for output files
Example for specific range:
pdftoppm -jpeg -r 150 -f 2 -l 5 template.pdf slide # Converts only pages 2-5
Code Style Guidelines
IMPORTANT: When generating code for PPTX operations:
- Write concise code
- Avoid verbose variable names and redundant operations
- Avoid unnecessary print statements
Dependencies
Required dependencies (should already be installed):
- markitdown:
pip install "markitdown[pptx]" (for text extraction from presentations)
- pptxgenjs:
npm install -g pptxgenjs (for creating presentations via html2pptx)
- playwright:
npm install -g playwright (for HTML rendering in html2pptx)
- react-icons:
npm install -g react-icons react react-dom (for icons in SVG format)
- LibreOffice:
sudo apt-get install libreoffice (for PDF conversion)
- Poppler:
sudo apt-get install poppler-utils (for pdftoppm to convert PDF to images)
- defusedxml:
pip install defusedxml (for secure XML parsing)
1---2name: pptx-43description: Presentation creation, editing, and analysis. When Claude needs to work with presentations (.pptx files) for: (1) Creating new presentations, (2) Modifying or editing content, (3) Working with layouts, (4) Adding comments or speaker notes, or any other presentation tasks4license: Proprietary. LICENSE.txt has complete terms5---6
7# PPTX creation, editing, and analysis
8
9## Overview
10
11Create, edit, or analyze the contents of .pptx files when requested. A .pptx file is essentially a ZIP archive containing XML files and other resources. Different tools and workflows are available for different tasks.
12
13## Reading and analyzing content
14
15### Text extraction
16
17To read just the text content of a presentation, convert the document to markdown:
18
19```bash
20# Convert document to markdown
21python -m markitdown path-to-file.pptx
22```
23
24### Raw XML access
25
26Use raw XML access for: comments, speaker notes, slide layouts, animations, design elements, and complex formatting. To access these features, unpack a presentation and read its raw XML contents.
27
28#### Unpacking a file
29
30`python ooxml/scripts/unpack.py <office_file> <output_dir>`
31
32**Note**: The unpack.py script is located at `skills/pptx/ooxml/scripts/unpack.py` relative to the project root. If the script doesn't exist at this path, use `find . -name "unpack.py"` to locate it.
33
34#### Key file structures
35
36- `ppt/presentation.xml` - Main presentation metadata and slide references
37- `ppt/slides/slide{N}.xml` - Individual slide contents (slide1.xml, slide2.xml, etc.)
38- `ppt/notesSlides/notesSlide{N}.xml` - Speaker notes for each slide
39- `ppt/comments/modernComment_*.xml` - Comments for specific slides
40- `ppt/slideLayouts/` - Layout templates for slides
41- `ppt/slideMasters/` - Master slide templates
42- `ppt/theme/` - Theme and styling information
43- `ppt/media/` - Images and other media files
44
45#### Typography and color extraction
46
47**To emulate example designs**, analyze the presentation's typography and colors first using the methods below:
48
491. **Read theme file**: Check `ppt/theme/theme1.xml` for colors (`<a:clrScheme>`) and fonts (`<a:fontScheme>`)
502. **Sample slide content**: Examine `ppt/slides/slide1.xml` for actual font usage (`<a:rPr>`) and colors
513. **Search for patterns**: Use grep to find color (`<a:solidFill>`, `<a:srgbClr>`) and font references across all XML files
52
53## Creating a new PowerPoint presentation **without a template**
54
55When creating a new PowerPoint presentation from scratch, use the **html2pptx** workflow to convert HTML slides to PowerPoint with accurate positioning.
56
57### Workflow
58
591. **MANDATORY - READ ENTIRE FILE NOW**: Read [`html2pptx.md`](html2pptx.md) completely from start to finish. **NEVER set any range limits when reading this file.** Read the full file content for detailed syntax, critical formatting rules, and best practices before proceeding with presentation creation.
602. **PREREQUISITE - Install html2pptx library**:
61 - Check and install if needed: `npm list -g @ant/html2pptx || npm install -g skills/pptx/html2pptx.tgz`
62 - **Note**: If you see "Cannot find module '@ant/html2pptx'" error later, the package isn't installed
633. **CRITICAL**: Plan the presentation
64 - Plan the shared aspects of the presentation. Describe the tone of the presentation's content and the colors and typography that should be used in the presentation.
65 - Write a DETAILED outline of the presentation
66 - For each slide, describe the slide's layout and contents
67 - For each slide, write presenter notes (1 to 3 sentences per slide)
684. **CRITICAL**: Set CSS variables
69 - In a shared `.css` file, override CSS variables to use on each slide for colors, typography, and spacing. DO NOT create classes in this file.
705. Create an HTML file for each slide with proper dimensions (e.g., 960px × 540px for 16:9)
71 - Recall the outline, layout/content description, and speaker notes you wrote for this slide in Step 3. Think out loud how to best apply them to this slide.
72 - Embed the contents of the shared `.css` file in a `<style>` element
73 - Use `<p>`, `<h1>`-`<h6>`, `<ul>`, `<ol>` for all text content
74 - **IMPORTANT:** Use CSS variables for colors, typography, and spacing
75 - **IMPORTANT:** Use `row` `col` and `fit` classes for layout INSTEAD OF flexbox
76 - Use `class="placeholder"` for areas where charts/tables will be added (render with gray background for visibility)
77 - **CSS gradients**: Use `linear-gradient()` or `radial-gradient()` in CSS on block element backgrounds - automatically converted to PowerPoint
78 - **Background images**: Use `background-image: url(...)` CSS property on block elements
79 - **Block elements**: Use `<div>`, `<section>`, `<header>`, `<footer>`, `<main>`, `<article>`, `<nav>`, `<aside>` for containers with styling (all behave identically)
80 - **Icons**: Use inline SVG format or reference SVG files - SVG elements are automatically converted to images in PowerPoint
81 - **Text balancing**: `<h1>` and `<h2>` elements are automatically balanced. Use `data-balance` attribute on other elements to auto-balance line lengths for better typography
82 - **Layout**: For slides with charts/tables/images, use either full-slide layout or two-column layout for better readability
836. Create and run a JavaScript file using the [`html2pptx`](./html2pptx) library to convert HTML slides to PowerPoint and save the presentation
84
85 - Run with: `NODE_PATH="$(npm root -g)" node your-script.js 2>&1`
86 - Use the `html2pptx` function to process each HTML file
87 - Add charts and tables to placeholder areas using PptxGenJS API
88 - Save the presentation using `pptx.writeFile()`
89
90 - **⚠️ CRITICAL:** Your script MUST follow this example structure. Think aloud before writing the script to make sure that you correctly use the APIs. Do NOT call `pptx.addSlide`.
91
92 ```javascript
93 const pptxgen = require("pptxgenjs");
94 const { html2pptx } = require("@ant/html2pptx");
95
96 // Create a new pptx presentation
97 const pptx = new pptxgen();
98 pptx.layout = "LAYOUT_16x9"; // Must match HTML body dimensions
99
100 // Add an HTML-only slide
101 await html2pptx("slide1.html", pptx);
102
103 // Add a HTML slide with chart placeholders
104 const { slide: slide2, placeholders } = await html2pptx("slide2.html", pptx);
105 slide.addChart(pptx.charts.LINE, chartData, placeholders[0]);
106
107 // Save the presentation
108 await pptx.writeFile("output.pptx");
109 ```
110
1117. **Visual validation**: Generate thumbnails and inspect for layout issues
112 - Create thumbnail grid: `python scripts/thumbnail.py output.pptx workspace/thumbnails --cols 4`
113 - Read and carefully examine the thumbnail image for:
114 - **Text cutoff**: Text being cut off by header bars, shapes, or slide edges
115 - **Text overlap**: Text overlapping with other text or shapes
116 - **Positioning issues**: Content too close to slide boundaries or other elements
117 - **Contrast issues**: Insufficient contrast between text and backgrounds
118 - If issues found, adjust HTML margins/spacing/colors and regenerate the presentation
119 - Repeat until all slides are visually correct
120
121## Editing an existing PowerPoint presentation
122
123To edit slides in an existing PowerPoint presentation, work with the raw Office Open XML (OOXML) format. This involves unpacking the .pptx file, editing the XML content, and repacking it.
124
125### Workflow
126
1271. **MANDATORY - READ ENTIRE FILE**: Read [`ooxml.md`](ooxml.md) (~500 lines) completely from start to finish. **NEVER set any range limits when reading this file.** Read the full file content for detailed guidance on OOXML structure and editing workflows before any presentation editing.
1282. Unpack the presentation: `python ooxml/scripts/unpack.py <office_file> <output_dir>`
1293. Edit the XML files (primarily `ppt/slides/slide{N}.xml` and related files)
1304. **CRITICAL**: Validate immediately after each edit and fix any validation errors before proceeding: `python ooxml/scripts/validate.py <dir> --original <file>`
1315. Pack the final presentation: `python ooxml/scripts/pack.py <input_directory> <office_file>`
132
133## Creating a new PowerPoint presentation **using a template**
134
135To create a presentation that follows an existing template's design, duplicate and re-arrange template slides before replacing placeholder content.
136
137### Workflow
138
1391. **Extract template text AND create visual thumbnail grid**:
140
141 - Extract text: `python -m markitdown template.pptx > template-content.md`
142 - Read `template-content.md`: Read the entire file to understand the contents of the template presentation. **NEVER set any range limits when reading this file.**
143 - Create thumbnail grids: `python scripts/thumbnail.py template.pptx`
144 - See [Creating Thumbnail Grids](#creating-thumbnail-grids) section for more details
145
1462. **Analyze template and save inventory to a file**:
147
148 - **Visual Analysis**: Review thumbnail grid(s) to understand slide layouts, design patterns, and visual structure
149 - Create and save a template inventory file at `template-inventory.md` containing:
150
151 ```markdown
152 # Template Inventory Analysis
153
154 **Total Slides: [count]**
155 **IMPORTANT: Slides are 0-indexed (first slide = 0, last slide = count-1)**
156
157 ## [Category Name]
158
159 - Slide 0: [Layout code if available] - Description/purpose
160 - Slide 1: [Layout code] - Description/purpose
161 - Slide 2: [Layout code] - Description/purpose
162 [... EVERY slide must be listed individually with its index ...]
163 ```
164
165 - **Using the thumbnail grid**: Reference the visual thumbnails to identify:
166 - Layout patterns (title slides, content layouts, section dividers)
167 - Image placeholder locations and counts
168 - Design consistency across slide groups
169 - Visual hierarchy and structure
170 - This inventory file is REQUIRED for selecting appropriate templates in the next step
171
1723. **Create presentation outline based on template inventory**:
173
174 - Review available templates from step 2.
175 - Choose an intro or title template for the first slide. This should be one of the first templates.
176 - Choose safe, text-based layouts for the other slides.
177 - **CRITICAL: Match layout structure to actual content**:
178 - Single-column layouts: Use for unified narrative or single topic
179 - Two-column layouts: Use ONLY when there are exactly 2 distinct items/concepts
180 - Three-column layouts: Use ONLY when there are exactly 3 distinct items/concepts
181 - Image + text layouts: Use ONLY when there are actual images to insert
182 - Quote layouts: Use ONLY for actual quotes from people (with attribution), never for emphasis
183 - Never use layouts with more placeholders than available content
184 - With 2 items, avoid forcing them into a 3-column layout
185 - With 4+ items, consider breaking into multiple slides or using a list format
186 - Count actual content pieces BEFORE selecting the layout
187 - Verify each placeholder in the chosen layout will be filled with meaningful content
188 - Select one option representing the **best** layout for each content section.
189 - Save `outline.md` with content AND template mapping that leverages available designs
190 - Example template mapping:
191 ```
192 # Template slides to use (0-based indexing)
193 # WARNING: Verify indices are within range! Template with 73 slides has indices 0-72
194 # Mapping: slide numbers from outline -> template slide indices
195 template_mapping = [
196 0, # Use slide 0 (Title/Cover)
197 34, # Use slide 34 (B1: Title and body)
198 34, # Use slide 34 again (duplicate for second B1)
199 50, # Use slide 50 (E1: Quote)
200 54, # Use slide 54 (F2: Closing + Text)
201 ]
202 ```
203
2044. **Duplicate, reorder, and delete slides using `rearrange.py`**:
205
206 - Use the `scripts/rearrange.py` script to create a new presentation with slides in the desired order:
207 ```bash
208 python scripts/rearrange.py template.pptx working.pptx 0,34,34,50,52
209 ```
210 - The script handles duplicating repeated slides, deleting unused slides, and reordering automatically
211 - Slide indices are 0-based (first slide is 0, second is 1, etc.)
212 - The same slide index can appear multiple times to duplicate that slide
213
2145. **Extract ALL text using the `inventory.py` script**:
215
216 - **Run inventory extraction**:
217 ```bash
218 python scripts/inventory.py working.pptx text-inventory.json
219 ```
220 - **Read text-inventory.json**: Read the entire text-inventory.json file to understand all shapes and their properties. **NEVER set any range limits when reading this file.**
221
222 - The inventory JSON structure:
223
224 ```json
225 {
226 "slide-0": {
227 "shape-0": {
228 "placeholder_type": "TITLE", // or null for non-placeholders
229 "left": 1.5, // position in inches
230 "top": 2.0,
231 "width": 7.5,
232 "height": 1.2,
233 "paragraphs": [
234 {
235 "text": "Paragraph text",
236 // Optional properties (only included when non-default):
237 "bullet": true, // explicit bullet detected
238 "level": 0, // only included when bullet is true
239 "alignment": "CENTER", // CENTER, RIGHT (not LEFT)
240 "space_before": 10.0, // space before paragraph in points
241 "space_after": 6.0, // space after paragraph in points
242 "line_spacing": 22.4, // line spacing in points
243 "font_name": "Arial", // from first run
244 "font_size": 14.0, // in points
245 "bold": true,
246 "italic": false,
247 "underline": false,
248 "color": "FF0000" // RGB color
249 }
250 ]
251 }
252 }
253 }
254 ```
255
256 - Key features:
257 - **Slides**: Named as "slide-0", "slide-1", etc.
258 - **Shapes**: Ordered by visual position (top-to-bottom, left-to-right) as "shape-0", "shape-1", etc.
259 - **Placeholder types**: TITLE, CENTER_TITLE, SUBTITLE, BODY, OBJECT, or null
260 - **Default font size**: `default_font_size` in points extracted from layout placeholders (when available)
261 - **Slide numbers are filtered**: Shapes with SLIDE_NUMBER placeholder type are automatically excluded from inventory
262 - **Bullets**: When `bullet: true`, `level` is always included (even if 0)
263 - **Spacing**: `space_before`, `space_after`, and `line_spacing` in points (only included when set)
264 - **Colors**: `color` for RGB (e.g., "FF0000"), `theme_color` for theme colors (e.g., "DARK_1")
265 - **Properties**: Only non-default values are included in the output
266
2676. **Generate replacement text and save the data to a JSON file**
268 Based on the text inventory from the previous step:
269
270 - **CRITICAL**: First verify which shapes exist in the inventory - only reference shapes that are actually present
271 - **VALIDATION**: The replace.py script validates that all shapes in the replacement JSON exist in the inventory
272 - Referencing a non-existent shape produces an error showing available shapes
273 - Referencing a non-existent slide produces an error indicating the slide doesn't exist
274 - All validation errors are shown at once before the script exits
275 - **IMPORTANT**: The replace.py script uses inventory.py internally to identify ALL text shapes
276 - **AUTOMATIC CLEARING**: ALL text shapes from the inventory are cleared unless "paragraphs" are provided for them
277 - Add a "paragraphs" field to shapes that need content (not "replacement_paragraphs")
278 - Shapes without "paragraphs" in the replacement JSON have their text cleared automatically
279 - Paragraphs with bullets are automatically left aligned. Avoid setting the `alignment` property when `"bullet": true`
280 - Generate appropriate replacement content for placeholder text
281 - Use shape size to determine appropriate content length
282 - **CRITICAL**: Include paragraph properties from the original inventory - don't just provide text
283 - **IMPORTANT**: When bullet: true, do NOT include bullet symbols (•, -, \*) in text - they're added automatically
284 - **ESSENTIAL FORMATTING RULES**:
285 - Headers/titles should typically have `"bold": true`
286 - List items should have `"bullet": true, "level": 0` (level is required when bullet is true)
287 - Preserve any alignment properties (e.g., `"alignment": "CENTER"` for centered text)
288 - Include font properties when different from default (e.g., `"font_size": 14.0`, `"font_name": "Lora"`)
289 - Colors: Use `"color": "FF0000"` for RGB or `"theme_color": "DARK_1"` for theme colors
290 - The replacement script expects **properly formatted paragraphs**, not just text strings
291 - **Overlapping shapes**: Prefer shapes with larger default_font_size or more appropriate placeholder_type
292 - Save the updated inventory with replacements to `replacement-text.json`
293 - **WARNING**: Different template layouts have different shape counts - always check the actual inventory before creating replacements
294
295 Example paragraphs field showing proper formatting:
296
297 ```json
298 "paragraphs": [
299 {
300 "text": "New presentation title text",
301 "alignment": "CENTER",
302 "bold": true
303 },
304 {
305 "text": "Section Header",
306 "bold": true
307 },
308 {
309 "text": "First bullet point without bullet symbol",
310 "bullet": true,
311 "level": 0
312 },
313 {
314 "text": "Red colored text",
315 "color": "FF0000"
316 },
317 {
318 "text": "Theme colored text",
319 "theme_color": "DARK_1"
320 },
321 {
322 "text": "Regular paragraph text without special formatting"
323 }
324 ]
325 ```
326
327 **Shapes not listed in the replacement JSON are automatically cleared**:
328
329 ```json
330 {
331 "slide-0": {
332 "shape-0": {
333 "paragraphs": [...] // This shape gets new text
334 }
335 // shape-1 and shape-2 from inventory will be cleared automatically
336 }
337 }
338 ```
339
340 **Common formatting patterns for presentations**:
341
342 - Title slides: Bold text, sometimes centered
343 - Section headers within slides: Bold text
344 - Bullet lists: Each item needs `"bullet": true, "level": 0`
345 - Body text: Usually no special properties needed
346 - Quotes: May have special alignment or font properties
347
3487. **Apply replacements using the `replace.py` script**
349
350 ```bash
351 python scripts/replace.py working.pptx replacement-text.json output.pptx
352 ```
353
354 The script will:
355
356 - First extract the inventory of ALL text shapes using functions from inventory.py
357 - Validate that all shapes in the replacement JSON exist in the inventory
358 - Clear text from ALL shapes identified in the inventory
359 - Apply new text only to shapes with "paragraphs" defined in the replacement JSON
360 - Preserve formatting by applying paragraph properties from the JSON
361 - Handle bullets, alignment, font properties, and colors automatically
362 - Save the updated presentation
363
364 Example validation errors:
365
366 ```
367 ERROR: Invalid shapes in replacement JSON:
368 - Shape 'shape-99' not found on 'slide-0'. Available shapes: shape-0, shape-1, shape-4
369 - Slide 'slide-999' not found in inventory
370 ```
371
372 ```
373 ERROR: Replacement text made overflow worse in these shapes:
374 - slide-0/shape-2: overflow worsened by 1.25" (was 0.00", now 1.25")
375 ```
376
377## Creating Thumbnail Grids
378
379To create visual thumbnail grids of PowerPoint slides for quick analysis and reference:
380
381```bash
382python scripts/thumbnail.py template.pptx [output_prefix]
383```
384
385**Features**:
386
387- Creates: `thumbnails.jpg` (or `thumbnails-1.jpg`, `thumbnails-2.jpg`, etc. for large decks)
388- Default: 5 columns, max 30 slides per grid (5×6)
389- Custom prefix: `python scripts/thumbnail.py template.pptx my-grid`
390 - Note: The output prefix should include the path if you want output in a specific directory (e.g., `workspace/my-grid`)
391- Adjust columns: `--cols 4` (range: 3-6, affects slides per grid)
392- Grid limits: 3 cols = 12 slides/grid, 4 cols = 20, 5 cols = 30, 6 cols = 42
393- Slides are zero-indexed (Slide 0, Slide 1, etc.)
394
395**Use cases**:
396
397- Template analysis: Quickly understand slide layouts and design patterns
398- Content review: Visual overview of entire presentation
399- Navigation reference: Find specific slides by their visual appearance
400- Quality check: Verify all slides are properly formatted
401
402**Examples**:
403
404```bash
405# Basic usage
406python scripts/thumbnail.py presentation.pptx
407
408# Combine options: custom name, columns
409python scripts/thumbnail.py template.pptx analysis --cols 4
410```
411
412## Converting Slides to Images
413
414To visually analyze PowerPoint slides, convert them to images using a two-step process:
415
4161. **Convert PPTX to PDF**:
417
418 ```bash
419 soffice --headless --convert-to pdf template.pptx
420 ```
421
4222. **Convert PDF pages to JPEG images**:
423 ```bash
424 pdftoppm -jpeg -r 150 template.pdf slide
425 ```
426 This creates files like `slide-1.jpg`, `slide-2.jpg`, etc.
427
428Options:
429
430- `-r 150`: Sets resolution to 150 DPI (adjust for quality/size balance)
431- `-jpeg`: Output JPEG format (use `-png` for PNG if preferred)
432- `-f N`: First page to convert (e.g., `-f 2` starts from page 2)
433- `-l N`: Last page to convert (e.g., `-l 5` stops at page 5)
434- `slide`: Prefix for output files
435
436Example for specific range:
437
438```bash
439pdftoppm -jpeg -r 150 -f 2 -l 5 template.pdf slide # Converts only pages 2-5
440```
441
442## Code Style Guidelines
443
444**IMPORTANT**: When generating code for PPTX operations:
445
446- Write concise code
447- Avoid verbose variable names and redundant operations
448- Avoid unnecessary print statements
449
450## Dependencies
451
452Required dependencies (should already be installed):
453
454- **markitdown**: `pip install "markitdown[pptx]"` (for text extraction from presentations)
455- **pptxgenjs**: `npm install -g pptxgenjs` (for creating presentations via html2pptx)
456- **playwright**: `npm install -g playwright` (for HTML rendering in html2pptx)
457- **react-icons**: `npm install -g react-icons react react-dom` (for icons in SVG format)
458- **LibreOffice**: `sudo apt-get install libreoffice` (for PDF conversion)
459- **Poppler**: `sudo apt-get install poppler-utils` (for pdftoppm to convert PDF to images)
460- **defusedxml**: `pip install defusedxml` (for secure XML parsing)