# PPTX

> Use this skill any time a .pptx file is involved in any way — as input, output, or both. This includes: creating slide decks, pitch decks, or presentations; reading, parsing, or extracting text from any .pptx file (even if the extracted content will be used elsewhere, like in an email or summary); editing, modifying, or updating existing presentations; combining or splitting slide files; working with templates, layouts, speaker notes, or comments. Trigger whenever the user mentions "deck," "slides," "presentation," or references a .pptx filename, regardless of what they plan to do with the content afterward. If a .pptx file needs to be opened, created, or touched, use this skill.

- Skill: `wakeeys/pptx` (Agent Skill, multi-file: 60 files)
- Install (CLI): `npx skillmds@latest add wakeeys/pptx`
- Raw SKILL.md: https://api.skillmd.com/api/skills/wakeeys/pptx/raw
- Safety review: pending
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Docs & Writing
- License: Proprietary. LICENSE.txt has complete terms
- Author: wakeeys (https://skillmd.com/u/wakeeys)
- Updated: 2026-09-22
- Page: https://skillmd.com/skills/wakeeys/pptx

---


# PPTX Skill

## Quick Reference

| Task | Guide |
|------|-------|
| Read/analyze content | `python -m markitdown presentation.pptx` |
| Edit or create from template | Read [editing.md](editing.md) |
| Create from scratch | Read [pptxgenjs.md](pptxgenjs.md) |
| Crop a figure/table from a PDF or image | `python scripts/crop_image.py` — measure the edge with `--profile`, then cut (see [Cropping Figures](#cropping-figures-for-slides)) |

---

## Reading Content

```bash
# Text extraction
python -m markitdown presentation.pptx

# Visual overview
python scripts/thumbnail.py presentation.pptx

# Raw XML
python scripts/office/unpack.py presentation.pptx unpacked/
```

---

## Editing Workflow

**Read [editing.md](editing.md) for full details.**

1. Analyze template with `thumbnail.py`
2. Unpack → manipulate slides → edit content → clean → pack

---

## Creating from Scratch

**Read [pptxgenjs.md](pptxgenjs.md) for full details.**

Use when no template or reference presentation is available.

---

## Design Ideas

**Don't create boring slides.** Plain bullets on a white background won't impress anyone. Consider ideas from this list for each slide.

### Before Starting

- **Pick a bold, content-informed color palette**: The palette should feel designed for THIS topic. If swapping your colors into a completely different presentation would still "work," you haven't made specific enough choices.
- **Dominance over equality**: One color should dominate (60-70% visual weight), with 1-2 supporting tones and one sharp accent. Never give all colors equal weight.
- **Dark/light contrast**: Dark backgrounds for title + conclusion slides, light for content ("sandwich" structure). Or commit to dark throughout for a premium feel.
- **Commit to a visual motif**: Pick ONE distinctive element and repeat it — rounded image frames, icons in colored circles, thick single-side borders. Carry it across every slide.

### Color Palettes

Choose colors that match your topic — don't default to generic blue. Use these palettes as inspiration:

| Theme | Primary | Secondary | Accent |
|-------|---------|-----------|--------|
| **Midnight Executive** | `1E2761` (navy) | `CADCFC` (ice blue) | `FFFFFF` (white) |
| **Forest & Moss** | `2C5F2D` (forest) | `97BC62` (moss) | `F5F5F5` (cream) |
| **Coral Energy** | `F96167` (coral) | `F9E795` (gold) | `2F3C7E` (navy) |
| **Warm Terracotta** | `B85042` (terracotta) | `E7E8D1` (sand) | `A7BEAE` (sage) |
| **Ocean Gradient** | `065A82` (deep blue) | `1C7293` (teal) | `21295C` (midnight) |
| **Charcoal Minimal** | `36454F` (charcoal) | `F2F2F2` (off-white) | `212121` (black) |
| **Teal Trust** | `028090` (teal) | `00A896` (seafoam) | `02C39A` (mint) |
| **Berry & Cream** | `6D2E46` (berry) | `A26769` (dusty rose) | `ECE2D0` (cream) |
| **Sage Calm** | `84B59F` (sage) | `69A297` (eucalyptus) | `50808E` (slate) |
| **Cherry Bold** | `990011` (cherry) | `FCF6F5` (off-white) | `2F3C7E` (navy) |

### For Each Slide

**Every slide needs a visual element** — image, chart, icon, or shape. Text-only slides are forgettable.

**Layout options:**
- Two-column (text left, illustration on right)
- Icon + text rows (icon in colored circle, bold header, description below)
- 2x2 or 2x3 grid (image on one side, grid of content blocks on other)
- Half-bleed image (full left or right side) with content overlay

**Data display:**
- Large stat callouts (big numbers 60-72pt with small labels below)
- Comparison columns (before/after, pros/cons, side-by-side options)
- Timeline or process flow (numbered steps, arrows)

**Visual polish:**
- Icons in small colored circles next to section headers
- Italic accent text for key stats or taglines

### Typography

**Choose an interesting font pairing** — don't default to Arial. Pick a header font with personality and pair it with a clean body font.

| Header Font | Body Font |
|-------------|-----------|
| Georgia | Calibri |
| Arial Black | Arial |
| Calibri | Calibri Light |
| Cambria | Calibri |
| Trebuchet MS | Calibri |
| Impact | Arial |
| Palatino | Garamond |
| Consolas | Calibri |

| Element | Size |
|---------|------|
| Slide title | 40-48pt bold |
| Section header | 26-30pt bold |
| Body text | 18pt (16pt only if a slide is genuinely text-dense) |
| Captions | 12-14pt muted |

Body text should be **18pt** by default — 14-16pt reads as cramped on a projected
slide. Scale the other levels up from there so the hierarchy stays obvious: a
title around 2.5× the body (≈44pt), section headers around 1.5× (≈28pt). If a
slide is so dense that 18pt overflows, prefer **splitting it into two slides**
over shrinking the body below 16pt.

### Spacing

- 0.5" minimum margins
- 0.3-0.5" between content blocks
- **Line spacing 1.15-1.25** for body text — single spacing packs lines too
  tightly to scan comfortably at a distance.
- **Bigger gap before a section header than after it.** A header needs clear
  air above it to separate it from the previous block, but should sit close to
  the body it introduces. Concretely: space-before ≈ 14-20pt on a header, and a
  small space-after (≈0-4pt) so the header "owns" the bullets beneath it. This
  asymmetry is what makes the structure read as grouped sections rather than an
  evenly-spaced wall of text.
- Leave breathing room—don't fill every inch.

### Avoid (Common Mistakes)

- **Don't repeat the same layout** — vary columns, cards, and callouts across slides
- **Don't center body text** — left-align paragraphs and lists; center only titles
- **Don't skimp on size contrast** — titles need ~2.5× the body size (≈44pt over 18pt) to stand out
- **Don't default to blue** — pick colors that reflect the specific topic
- **Don't mix spacing randomly** — choose 0.3" or 0.5" gaps and use consistently
- **Don't style one slide and leave the rest plain** — commit fully or keep it simple throughout
- **Don't create text-only slides** — add images, icons, charts, or visual elements; avoid plain title + bullets
- **Don't forget text box padding** — when aligning lines or shapes with text edges, set `margin: 0` on the text box or offset the shape to account for padding
- **Don't use low-contrast elements** — icons AND text need strong contrast against the background; avoid light text on light backgrounds or dark text on dark backgrounds
- **Don't put gray borders on images** — inserted pictures should be borderless; omit the `<a:ln>` outline on the picture unless the design explicitly calls for a frame
- **Don't add caption text boxes under figures/tables** — a separate italic "Fig. N — ..." / "Table N — ..." caption below an image is usually unnecessary clutter; rely on the section header and let the figure speak for itself. Only add a caption if the user explicitly asks.
- **NEVER use accent lines under titles** — these are a hallmark of AI-generated slides; use whitespace or background color instead

---

## QA (Required)

**Assume there are problems. Your job is to find them.**

Your first render is almost never correct. Approach QA as a bug hunt, not a confirmation step. If you found zero issues on first inspection, you weren't looking hard enough.

### Content QA

```bash
python -m markitdown output.pptx
```

Check for missing content, typos, wrong order.

**When using templates, check for leftover placeholder text:**

```bash
python -m markitdown output.pptx | grep -iE "xxxx|lorem|ipsum|this.*(page|slide).*layout"
```

If grep returns results, fix them before declaring success.

### Visual QA

**⚠️ USE SUBAGENTS** — even for 2-3 slides. You've been staring at the code and will see what you expect, not what's there. Subagents have fresh eyes.

Convert slides to images (see [Converting to Images](#converting-to-images)), then use this prompt:

```
Visually inspect these slides. Assume there are issues — find them.

Look for:
- Overlapping elements (text through shapes, lines through words, stacked elements)
- Text overflow or cut off at edges/box boundaries
- Decorative lines positioned for single-line text but title wrapped to two lines
- Source citations or footers colliding with content above
- Elements too close (< 0.3" gaps) or cards/sections nearly touching
- Uneven gaps (large empty area in one place, cramped in another)
- Insufficient margin from slide edges (< 0.5")
- Columns or similar elements not aligned consistently
- Low-contrast text (e.g., light gray text on cream-colored background)
- Low-contrast icons (e.g., dark icons on dark backgrounds without a contrasting circle)
- Text boxes too narrow causing excessive wrapping
- Leftover placeholder content

For each slide, list issues or areas of concern, even if minor.

Read and analyze these images:
1. /path/to/slide-01.jpg (Expected: [brief description])
2. /path/to/slide-02.jpg (Expected: [brief description])

Report ALL issues found, including minor ones.
```

### Verification Loop

1. Generate slides → Convert to images → Inspect
2. **List issues found** (if none found, look again more critically)
3. Fix issues
4. **Re-verify affected slides** — one fix often creates another problem
5. Repeat until a full pass reveals no new issues

**Do not declare success until you've completed at least one fix-and-verify cycle.**

---

## Cropping Figures for Slides

When a slide needs a figure, table, or algorithm pulled from a **PDF paper** or
an existing image, use `scripts/crop_image.py`. Do **not** reuse low-resolution
pre-cropped images — they often have missing edges or cut-off content. Always
re-extract from the highest-quality source.

### Measure, then cut (fewest re-crops)

Eyeballing the crop box is the #1 cause of crop-guess-recrop loops — you misjudge
an edge by a few pixels and silently slice off content (the classic "the bottom
got cut off"). Don't guess edges; **measure them** with `--profile`, then cut once.

```bash
# Step 1 — LOCATE: render the whole page and Read it to see roughly where the
# figure sits (no crop args = full page out). Use 400 DPI for slide-bound figures.
python scripts/crop_image.py paper.pdf page13.png --page 13 --dpi 400

# Step 2 — MEASURE the tight edge deterministically (don't eyeball it). Scan ink
# density along the edge's axis, bounded to the figure's column. 'h' = find a
# TOP/BOTTOM row; 'v' = find a LEFT/RIGHT column.
python scripts/crop_image.py paper.pdf /dev/null --page 13 --dpi 400 \
    --profile h --region 0.05,0.21,0.49,0.28
#   → prints e.g.  24.2% ink=86 (tick numbers) / 24.7-25.6% ink=0 (BLANK GAP)
#                  / 25.9% ink=193 (caption). Put the bottom edge at ~25.1%.

# Step 3 — CUT with the measured box. --trim cleans residual margin, --pad adds room.
python scripts/crop_image.py paper.pdf fig7.png --page 13 --dpi 400 \
    --frac 0.05,0.075,0.49,0.251 --trim --pad 12

# Existing image: crop by exact pixels, or just trim its white margins:
python scripts/crop_image.py fig.png out.png --box 100,80,900,1200
python scripts/crop_image.py fig.png out.png --trim --pad 16
```

### Why this avoids the "missing content" problem
- **Render at 300+ DPI (400 for slides).** The default is 300; low DPI looks fine
  in a thumbnail but is visibly soft when projected. This is the #1 reason a crop
  looks worse than the original.
- `--profile` makes the exact edge **deterministic** (a printed blank-gap, not a
  visual estimate of where content ends).
- `--frac L,T,R,B` then cuts to the measured box; `--trim` shrinks off residual
  uniform margin (it never removes real content), `--pad N` adds breathing room.

### Edge-finding rules (this is where crops actually go wrong)
- **A profile tells you where content IS; the edge belongs in the blank gap
  AFTER it**, not on the last inked line — and that gap is often <1% of page
  height, which is why eyeballing it shaves content off.
- **BOTTOM of a line plot is the riskiest edge.** The x-axis tick-number row +
  "Steps"/axis title is a thin band right above the "Figure N:" caption. Include
  the ticks, land the edge in the blank gap, stop before the caption.
- **Tables and algorithm boxes are bounded by horizontal rules, and the rule is
  PART of the figure.** A profile shows each rule as a tall ink spike; put the
  edge in the blank *outside* the rule, and **do NOT use `--trim` on these** —
  trim hugs the content and shaves the top/bottom rule off (this is exactly how a
  table ends up looking "cut in half"). After cropping, confirm the top rule,
  bottom rule, and the first AND last data row are all present.
- **Read the source table/figure first to learn its true extent.** Don't trust
  the user's row list as the boundary — e.g. a table the user describes as "three
  blocks" may actually have a fourth. Crop the real figure, and flag the
  discrepancy rather than silently dropping rows.
- **Two-column papers:** a figure in one column has the OTHER column's text at
  the same height. Keep L/R inside the column (≈`0.05–0.49` or `0.50–0.96`),
  never span the gutter. Bound `--profile` with `--region` to the column so
  neighbouring text doesn't skew the ink count.
- **Always exclude the caption** ("Figure N:" / "Table N:") unless asked to keep it.
- **`--trim` is a safety net for borderless plots, not a fix for a bad edge.** It
  removes blank margin so a slightly *loose* box is safe, but it will NOT rescue a
  box whose edge already cuts *through* content. Measure tight edges with `--profile`.
- **`--pad` adds white margin AFTER cropping — it does NOT pull in more of the
  source.** If your `--frac` edge lands a hair short of a table's bounding rule,
  the rule is already gone; `--pad` just paints white outside the clipped box. So
  the rule must be INSIDE the `--frac` box: set the bottom/top fraction *past* the
  rule (into the blank gap beyond it), e.g. rule at 42.69% → bottom edge ≈ 0.434,
  not 0.426. A rule clipped by ~0.1% is the classic "bottom line / top line is
  missing" bug.
- After cropping, **Read the output** and check all four edges — for multi-row
  tables / multi-panel figures confirm the first AND last row/panel, the bounding
  rules, and all axes are present.

Run `python scripts/crop_image.py -h` for all options.

**Note:** `crop_image.py` does not add any border. If your slide-insertion code
draws an image outline (e.g. a gray `<a:ln>`), that border comes from the slide
XML, not the image — remove it there if unwanted.

---

## Converting to Images

Convert presentations to individual slide images for visual inspection:

```bash
python scripts/office/soffice.py --headless --convert-to pdf output.pptx
pdftoppm -jpeg -r 150 output.pdf slide
```

This creates `slide-01.jpg`, `slide-02.jpg`, etc.

To re-render specific slides after fixes:

```bash
pdftoppm -jpeg -r 150 -f N -l N output.pdf slide-fixed
```

---

## Dependencies

- `pip install "markitdown[pptx]"` - text extraction
- `pip install Pillow` - thumbnail grids
- `pip install Pillow numpy` - figure cropping (`scripts/crop_image.py`; numpy only for `--profile`)
- `npm install -g pptxgenjs` - creating from scratch
- LibreOffice (`soffice`) - PDF conversion (auto-configured for sandboxed environments via `scripts/office/soffice.py`)
- Poppler (`pdftoppm`) - PDF to images

