# Codex Ppt Skill

> Generate image-based PowerPoint presentations using gpt-image-2, converting articles, papers, and reports into visual slide decks

- Skill: `aradotso/codex-ppt-skill` (Agent Skill)
- Install (CLI): `npx skillmds@latest add aradotso/codex-ppt-skill`
- Raw SKILL.md: https://api.skillmd.com/api/skills/aradotso/codex-ppt-skill/raw
- Safety review: pending
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Docs & Writing
- Author: aradotso (https://skillmd.com/u/aradotso)
- Updated: 2026-09-17
- Page: https://skillmd.com/skills/aradotso/codex-ppt-skill

---


# codex-ppt-skill

> Skill by [ara.so](https://ara.so) — Codex Skills collection.

A skill for generating image-based PowerPoint presentations where each slide is a complete 16:9 image generated by `gpt-image-2`. Converts articles, papers, reports, and notes into visually cohesive presentation decks with unified styling.

## What This Skill Does

- **Image-based slides**: Each slide is a full 16:9 image, perfect for strong visual storytelling
- **Multi-agent support**: Works in Codex, Claude Code, OpenClaw, Hermes Agent
- **Style library**: Built-in visual styles (clean professional, scientific defense, e-ink magazine, hand-drawn technical, dashboard, etc.)
- **Unified visual language**: Maintains consistent styling while varying layouts per content
- **Custom images**: Insert specific figures, diagrams, or screenshots on designated slides
- **Local assembly**: Python script packages generated images into `.pptx` with speaker notes

## Installation

### For Codex

```bash
npx -y skills@latest add ningzimu/codex-ppt-skill \
  --skill codex-ppt \
  --agent codex \
  --global
```

Restart Codex after installation.

### For Claude Code

```bash
npx -y skills@latest add ningzimu/codex-ppt-skill \
  --skill codex-ppt \
  --agent claude-code \
  --global
```

### For OpenClaw

```bash
openclaw skills install codex-ppt
```

### For Hermes Agent

```bash
npx -y skills@latest add ningzimu/codex-ppt-skill \
  --skill codex-ppt \
  --agent hermes-agent \
  --global
```

### Manual Installation

Clone and symlink to your agent's skills directory:

```bash
git clone https://github.com/ningzimu/codex-ppt-skill.git
mkdir -p ~/.codex/skills
ln -s $(pwd)/codex-ppt-skill/skills/codex-ppt ~/.codex/skills/codex-ppt
```

## Image Generation Configuration

**Important**: Only configure if you need API/CLI fallback. If using Codex with GPT subscription and built-in image generation works, **skip this section**.

Configure only when:
- Using third-party OpenAI-compatible APIs
- Using Claude Code, OpenClaw, or Hermes Agent
- Codex built-in image generation is unavailable

### Configuration Command

```bash
python3 ~/.codex/skills/codex-ppt/scripts/codex_ppt_runtime.py config \
  --api-key "$OPENAI_API_KEY" \
  --model gpt-image-2
```

With custom base URL (for third-party providers):

```bash
python3 ~/.codex/skills/codex-ppt/scripts/codex_ppt_runtime.py config \
  --api-key "$OPENAI_API_KEY" \
  --base-url "https://api.example.com/v1" \
  --model openai/gpt-image-2
```

Configuration is stored in `~/.codex-ppt-skill/.env` and shared across all agents.

### Environment Variables

If configured, the `.env` file contains:

```bash
OPENAI_API_KEY=your-api-key-here
OPENAI_BASE_URL=https://api.example.com/v1  # optional
OPENAI_IMAGE_MODEL=gpt-image-2
```

## Key Commands and API

### Generate Image via CLI

```bash
python3 ~/.codex/skills/codex-ppt/scripts/codex_ppt_runtime.py generate \
  --prompt "Clean professional slide: Introduction to AI, blue gradient background, large title, 3 bullet points" \
  --output /path/to/slide_01.png \
  --size 2048x1152
```

### Assemble PPT from Images

```bash
python3 ~/.codex/skills/codex-ppt/scripts/assemble_ppt.py \
  --project-dir /path/to/ppt-project \
  --title "My Presentation" \
  --speech-file /path/to/ppt-project/speech.md
```

### Check Configuration

```bash
python3 ~/.codex/skills/codex-ppt/scripts/codex_ppt_runtime.py check-config
```

Output:
```
✓ Configuration file exists
✓ API key configured
✓ Model: gpt-image-2
✓ Base URL: https://api.openai.com/v1
```

## Workflow and Usage Patterns

### Basic Usage

```python
# In agent conversation:
# "Use codex-ppt skill to create a 10-slide presentation from article.md"
```

The skill follows this workflow:

1. **Read content** and create outline
2. **Generate `outline.md`** with slide titles and key points
3. **Request confirmation** on page count and structure
4. **Propose visual styles** (2-3 options with recommendation)
5. **Confirm image generation backend** before first generation
6. **Generate sample slide** for style approval
7. **Create project directory** structure
8. **Generate all slides** with unified styling
9. **Quality check** (text clarity, style consistency)
10. **Generate `speech.md`** with speaker notes
11. **Assemble `.pptx`** using local script

### Project Directory Structure

```
{base_dir}/{ppt_name}/
├── origin_image/
│   ├── slide_01.png      # Title slide
│   ├── slide_02.png      # Content slides
│   ├── slide_03.png
│   └── ...
├── outline.md            # Slide structure
├── speech.md            # Speaker notes (## Slide 1: Title format)
└── {ppt_name}.pptx      # Final presentation
```

### Example: Article to Presentation

```python
# Input: technical_article.md containing AI research summary
# Agent command: "Use codex-ppt skill to make this into slides"

# Step 1: Outline generation
outline = """
# AI Research Presentation Outline

## Slide 1: Title
- "Recent Advances in Large Language Models"
- Speaker name, date

## Slide 2: Background
- Evolution of NLP
- Pre-transformer era vs transformer era
- Key milestones timeline

## Slide 3: Architecture
- Transformer architecture diagram
- Attention mechanism
- Scaling laws

# ... (continues)
"""

# Step 2: Style selection
styles = [
    "clean-professional",      # Recommended for tech talks
    "scientific-defense",      # For academic presentations
    "handdrawn-technical"     # For approachable tech explanation
]

# Step 3: Sample slide generation (2K resolution)
sample_prompt = """
16:9 slide, clean professional style, blue gradient background
Title: "Recent Advances in Large Language Models"
Subtitle: "Dr. Jane Smith | May 2024"
Minimalist design, large readable font, subtle geometric accents
"""

# Step 4: Bulk generation (all remaining slides)
# Agent generates each slide with consistent style but varied layouts
```

### Example: Inserting Custom Images

```python
# Outline with custom figure specification
outline_with_figures = """
## Slide 5: Model Architecture
- Insert: /path/to/architecture_diagram.png
- Transformer architecture
- Multi-head attention
- Feed-forward layers

## Slide 8: Experimental Results
- Insert: /path/to/results_chart.png
- Benchmark comparison
- Performance metrics
"""

# The skill will:
# 1. Use provided image as background/focal element
# 2. Add consistent styling (borders, background, text overlays)
# 3. Maintain visual coherence with other slides
```

### Example: Python Assembly Script Usage

```python
#!/usr/bin/env python3
from pptx import Presentation
from pptx.util import Inches
import os

def assemble_presentation(project_dir: str, title: str, speech_file: str):
    """
    Assemble PPT from images in origin_image/ directory.
    
    Args:
        project_dir: Path to PPT project directory
        title: Presentation title
        speech_file: Path to speech.md with speaker notes
    """
    prs = Presentation()
    prs.slide_width = Inches(10)  # 16:9 aspect ratio
    prs.slide_height = Inches(5.625)
    
    image_dir = os.path.join(project_dir, "origin_image")
    images = sorted([f for f in os.listdir(image_dir) if f.startswith("slide_")])
    
    # Parse speaker notes
    notes_map = parse_speaker_notes(speech_file)
    
    for idx, img_file in enumerate(images, 1):
        slide_layout = prs.slide_layouts[6]  # Blank layout
        slide = prs.slides.add_slide(slide_layout)
        
        # Add image filling entire slide
        img_path = os.path.join(image_dir, img_file)
        slide.shapes.add_picture(
            img_path, 
            Inches(0), Inches(0),
            width=Inches(10), height=Inches(5.625)
        )
        
        # Add speaker notes
        if idx in notes_map:
            slide.notes_slide.notes_text_frame.text = notes_map[idx]
    
    output_path = os.path.join(project_dir, f"{title}.pptx")
    prs.save(output_path)
    return output_path

def parse_speaker_notes(speech_file: str) -> dict:
    """Extract speaker notes by slide number from markdown."""
    notes = {}
    current_slide = None
    current_text = []
    
    with open(speech_file, 'r', encoding='utf-8') as f:
        for line in f:
            if line.startswith("## Slide "):
                if current_slide:
                    notes[current_slide] = "\n".join(current_text).strip()
                # Extract slide number
                current_slide = int(line.split("Slide ")[1].split(":")[0])
                current_text = []
            elif current_slide:
                current_text.append(line.rstrip())
    
    if current_slide:
        notes[current_slide] = "\n".join(current_text).strip()
    
    return notes
```

## Visual Styles Reference

Built-in styles in `skills/codex-ppt/references/styles.md`:

### clean-professional
- Blue/gray gradients, sans-serif fonts
- Minimalist design, ample white space
- Subtle geometric accents
- Best for: corporate, tech talks

### scientific-defense
- Academic journal aesthetic
- Structured layouts with clear sections
- Chart-friendly, equation-compatible
- Best for: research presentations, thesis defense

### e-ink-magazine
- Black and white high contrast
- Editorial typography
- Grid-based layouts
- Best for: content-heavy, text-focused presentations

### handdrawn-technical
- Hand-drawn diagrams and annotations
- Whiteboard aesthetic with digital polish
- Friendly, approachable
- Best for: tutorials, educational content

### data-dashboard
- Dark background with bright data visualizations
- KPI-focused layouts
- Chart and metric emphasis
- Best for: business reviews, analytics presentations

### retro-flat-illustration
- Flat design with vintage color palettes
- Illustrative icons and graphics
- Playful yet professional
- Best for: creative presentations, marketing

### warm-handmade
- Textured backgrounds, craft-paper feel
- Handwritten fonts, natural colors
- Organic, human-centered
- Best for: storytelling, personal projects

## Common Patterns

### Pattern 1: Conference Talk from Paper

```python
# Input: research_paper.pdf (converted to markdown)
# Command: "Create a 15-slide conference talk using scientific-defense style"

# Agent workflow:
# 1. Extract key sections (abstract, methodology, results, conclusions)
# 2. Create outline with 1 title + 2-3 intro + 6-8 technical + 2-3 conclusion slides
# 3. Insert paper figures on relevant slides
# 4. Generate slides with consistent academic styling
# 5. Add detailed speaker notes from paper content
```

### Pattern 2: Business Quarterly Review

```python
# Input: Q4_2024_metrics.md with KPIs and charts
# Command: "Make data-dashboard style slides, insert my 3 chart images"

# Outline example:
"""
## Slide 1: Q4 2024 Review
## Slide 2: Revenue Overview
- Insert: revenue_chart.png
## Slide 3: User Growth
- Insert: growth_chart.png
## Slide 4: Regional Performance
- Insert: regional_map.png
## Slide 5: Key Initiatives
## Slide 6: Q1 2025 Goals
"""
```

### Pattern 3: Course Lecture Slides

```python
# Input: lecture_notes.md (markdown with headings and bullets)
# Command: "Create handdrawn-technical style slides, about 20 pages"

# Agent approach:
# - 1 slide per major concept
# - Visual metaphors for abstract ideas
# - Step-by-step diagrams for processes
# - Summary slide every 5-6 slides
# - Consistent whiteboard aesthetic throughout
```

## Resolution and Quality

### Standard Resolution (2K)
```bash
# 2048x1152 (16:9, sufficient for most presentations)
python3 codex_ppt_runtime.py generate \
  --prompt "..." \
  --output slide.png \
  --size 2048x1152
```

### High Resolution (4K)
```bash
# 3840x2160 (16:9, for text-heavy slides or printing)
python3 codex_ppt_runtime.py generate \
  --prompt "..." \
  --output slide.png \
  --size 3840x2160
```

**When to use 4K:**
- Slides with extensive text (>100 words)
- Complex diagrams or code snippets
- Presentations for large screens or printing
- Fine typography matters

## Troubleshooting

### Issue: Blurry text on slides

**Solution:** Increase resolution to 4K
```python
# In agent: "Regenerate slides 3-5 at 4K resolution for better text clarity"
```

### Issue: Style inconsistency between slides

**Cause:** Vague or varying style prompts

**Solution:** 
1. Lock style reference in first sample slide
2. Reuse exact style description for all slides
3. Only vary content/layout, keep color/font/theme identical

```python
# Good: Consistent base prompt
base_style = "Clean professional, blue gradient (#1e3a8a to #3b82f6), Inter font, minimalist"
slide_2_prompt = f"{base_style}\nTitle: Introduction\n3 bullet points..."
slide_3_prompt = f"{base_style}\nTitle: Methods\nDiagram with 4 boxes..."

# Bad: Different styles per slide
slide_2_prompt = "Blue background, title and bullets"
slide_3_prompt = "Professional slide with diagram"  # Too vague
```

### Issue: Missing speaker notes in assembled PPT

**Cause:** `speech.md` format doesn't match parser

**Solution:** Use strict heading format
```markdown
## Slide 1: Title Slide
This is the opening slide. Introduce yourself and the topic.

## Slide 2: Background
Explain the context. Mention key historical developments.

## Slide 3: Problem Statement
...
```

Parser expects `## Slide N: Title` format exactly.

### Issue: Image generation fails with API error

**Check configuration:**
```bash
python3 codex_ppt_runtime.py check-config
```

**Common fixes:**
- Verify API key in `~/.codex-ppt-skill/.env`
- Check base URL (must end with `/v1`)
- Confirm model name matches provider's requirements
- Test with direct curl:

```bash
curl https://api.openai.com/v1/images/generations \
  -H "Authorization: Bearer $OPENAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-image-2",
    "prompt": "Test slide",
    "size": "2048x1152"
  }'
```

### Issue: Agent not using codex-ppt skill

**Solution:** Explicitly mention the skill
```python
# Weak: "Make slides from this article"
# Strong: "Use codex-ppt skill to generate slides from this article"
```

### Issue: Custom image not properly integrated

**Correct outline format:**
```markdown
## Slide 5: Architecture
- Insert: /absolute/path/to/diagram.png
- Model architecture overview
- Key components
```

Agent should:
1. Place custom image as slide background or main element
2. Add styling overlay consistent with other slides
3. Optionally add title/annotations matching theme

## Advanced Usage

### Modify Existing Slide

```python
# Don't regenerate entire deck
# Command: "Regenerate slide 7 with darker background and larger font"

# Agent should:
# 1. Load slide_07 prompt from history or reconstruct from outline
# 2. Modify only specified parameters
# 3. Overwrite origin_image/slide_07.png
# 4. Re-run assembly script
```

### Custom Style Reference

```python
# Upload reference image
# Command: "Use this slide style for the entire presentation"

# Agent approach:
# 1. Analyze uploaded image (colors, fonts, layout, visual elements)
# 2. Create detailed style description
# 3. Apply to all slide prompts
# 4. Generate sample for confirmation
```

### Batch Export Multiple Formats

```python
from pptx import Presentation
import subprocess

def export_formats(pptx_path: str):
    """Export to PDF and images."""
    base = os.path.splitext(pptx_path)[0]
    
    # PDF (requires LibreOffice)
    subprocess.run([
        "libreoffice", "--headless", "--convert-to", "pdf",
        "--outdir", os.path.dirname(pptx_path),
        pptx_path
    ])
    
    # Individual PNGs already in origin_image/
    print(f"Exported: {base}.pdf")
    print(f"Images: {os.path.dirname(pptx_path)}/origin_image/")
```

## Best Practices

1. **Start with outline confirmation**: Always review structure before generating images
2. **Generate sample first**: Confirm style with 1-2 slides before bulk generation
3. **Use 4K for text-heavy slides**: Don't compromise readability
4. **Keep style locked**: Identical base prompt for all slides in a deck
5. **Semantic layout variation**: Change layout to match content type (intro vs data vs summary)
6. **Speaker notes**: Write detailed notes during generation, not after
7. **Custom images**: Pre-prepare figures at high resolution (min 1920px width)
8. **Iterative refinement**: Fix individual slides, don't regenerate entire decks

## Environment Variables Reference

```bash
# Required only for API/CLI fallback
OPENAI_API_KEY=sk-...                           # Your OpenAI API key
OPENAI_BASE_URL=https://api.openai.com/v1      # Optional, for third-party providers
OPENAI_IMAGE_MODEL=gpt-image-2                  # Model name
```

Never commit these to version control. Store in `~/.codex-ppt-skill/.env`.

