[IMPORTANT] Use TaskCreate to break ALL work into small tasks BEFORE starting — including tasks for each file read. This prevents context loss from long files. For simple tasks, AI MUST ask user whether to skip.
Quick Summary
Goal: Process and generate multimedia content (images, audio, video, documents) using Google Gemini API via Python scripts.
Workflow:
- Identify Modality — Match input type to task (analyze, transcribe, extract, generate)
- Check Limits — Inline max 20MB, File API max 2GB; split large audio at 15min chunks
- Execute — Run
gemini_batch_process.py with appropriate task and files
- Post-Process — Format output as markdown with timestamps, save generated content
Key Rules:
- Requires
GEMINI_API_KEY environment variable
- Always request specific nodes/files, avoid full-file downloads
- Use
media_optimizer.py to compress/split files exceeding limits
Be skeptical. Apply critical thinking, sequential thinking. Every claim needs traced proof, confidence percentages (Idea should be more than 80%).
AI Multimodal
Purpose
Process audio, images, videos, and documents or generate images/videos using Google Gemini's multimodal API via bundled Python scripts.
When to Use
- Analyzing images or screenshots (Gemini vision is preferred over Claude's built-in vision for complex tasks)
- Transcribing audio files (meetings, podcasts, interviews)
- Extracting data from PDFs, scanned documents, or charts
- Processing video content (scene detection, temporal Q&A)
- Generating images with Imagen 4 or videos with Veo 3
- Converting documents to markdown with visual understanding
When NOT to Use
- Simple text-only LLM calls -- use Claude directly
- Reading a file Claude can already read (code, markdown, JSON) -- use
Read tool
- Building AI-powered application features -- use
api-design or frontend-design
- Music composition workflows -- load
references/music-generation.md only when specifically requested
- General prompt engineering -- use
ai-artist skill
Prerequisites
export GEMINI_API_KEY="your-key" # From https://aistudio.google.com/apikey
pip install google-genai python-dotenv pillow
python scripts/check_setup.py # Verify setup
Optional: API key rotation for rate limits (set GEMINI_API_KEY_2, GEMINI_API_KEY_3).
Workflow
Step 1: Identify Modality
| Input Type |
Task |
Command |
| Image (PNG/JPG/WEBP) |
Analyze, caption, OCR |
--task analyze |
| Audio (WAV/MP3/AAC) |
Transcribe, summarize |
--task transcribe |
| Video (MP4/MOV) |
Scene detection, Q&A |
--task analyze |
| PDF/Document |
Extract tables, forms |
--task extract |
| Text prompt |
Generate image |
--task generate |
| Text prompt |
Generate video |
--task generate-video |
Step 2: Check Limits
- Inline upload: max 20MB
- File API: max 2GB (auto-used for large files)
- Audio transcription: split at 15-minute chunks for full transcript
- Video transcription: extract audio first, then split and transcribe
- Formats: Audio (WAV/MP3/AAC, up to 9.5h), Images (PNG/JPEG/WEBP, up to 3.6k), Video (MP4/MOV, up to 6h), PDF (up to 1k pages)
IF file exceeds limits, use scripts/media_optimizer.py to compress/split first.
Step 3: Execute
Quick check: If gemini CLI is available, use: "<prompt>" | gemini -y -m gemini-2.5-flash
Standard: Use the batch processing script:
# Analyze media
python scripts/gemini_batch_process.py --files <file> --task <analyze|transcribe|extract>
# Generate content
python scripts/gemini_batch_process.py --task generate --prompt "description"
python scripts/gemini_batch_process.py --task generate-video --prompt "description"
Stdin support: cat image.png | python scripts/gemini_batch_process.py --task analyze --prompt "Describe this"
Step 4: Post-Processing
- For transcripts: output in markdown with
[HH:MM:SS -> HH:MM:SS] timestamps
- For document extraction: save as structured markdown under
docs/assets/
- For generated images/videos: save to working directory with descriptive filename
Step 5: Verification
- Confirm output matches expected format and completeness
- For long transcripts: verify no truncation occurred (check chunk boundaries)
- For generated content: verify quality meets prompt requirements
Models
| Purpose |
Model |
Notes |
| Analysis (fast) |
gemini-2.5-flash |
Recommended default |
| Analysis (advanced) |
gemini-2.5-pro |
Complex reasoning tasks |
| Image generation |
imagen-4.0-generate-001 |
Standard quality |
| Image generation (quality) |
imagen-4.0-ultra-generate-001 |
Best quality |
| Image generation (speed) |
imagen-4.0-fast-generate-001 |
Fastest |
| Video generation |
veo-3.1-generate-preview |
8s clips with audio |
Scripts Reference
gemini_batch_process.py -- CLI orchestrator for all tasks, auto-resolves API keys and models
media_optimizer.py -- Compress/resize/split media to fit Gemini limits
document_converter.py -- Convert PDFs/images/Office docs to markdown
check_setup.py -- Verify environment, dependencies, and API key
Use --help on any script for full options.
Examples
Example 1: Transcribe a Meeting Recording
Input: 45-minute meeting audio file meeting-2025-01-15.mp3
Steps:
- File is >15min, so split first:
python scripts/media_optimizer.py --input meeting-2025-01-15.mp3 --split-duration 900
- Transcribe each chunk:
python scripts/gemini_batch_process.py --files meeting-part-*.mp3 --task transcribe
- Output: Markdown file with timestamps, speaker detection, and metadata (duration, topics covered)
Example 2: Extract Data from a PDF Report
Input: Quarterly HR report PDF with tables, charts, and forms
Steps:
- Convert and extract:
python scripts/document_converter.py --input quarterly-report.pdf --output docs/assets/
- Output: Structured markdown with tables preserved, chart descriptions, and form field values extracted
Detailed References
Load for in-depth guidance:
| Topic |
File |
| Audio processing |
references/audio-processing.md |
| Vision/image analysis |
references/vision-understanding.md |
| Image generation |
references/image-generation.md |
| Video analysis |
references/video-analysis.md |
| Video generation |
references/video-generation.md |
| Music generation |
references/music-generation.md |
Related Skills
ai-artist -- for prompt engineering and optimization (not media processing)
media-processing -- for FFmpeg-based audio/video encoding without AI
pdf-to-markdown -- for simple PDF text extraction without vision AI
IMPORTANT Task Planning Notes (MUST FOLLOW)
- Always plan and break work into many small todo tasks
- Always add a final review todo task to verify work quality and identify fixes/enhancements
1---2name: ai-multimodal-43description: [AI & Tools] Process and generate multimedia content using Google Gemini API -- vision analysis, audio transcription, video processing, document extraction, image/video generation. Triggers on multimodal, vision API, image recognition, audio transcription, video analysis, gemini, imagen, document extraction.4---5
6> **[IMPORTANT]** Use `TaskCreate` to break ALL work into small tasks BEFORE starting — including tasks for each file read. This prevents context loss from long files. For simple tasks, AI MUST ask user whether to skip.
7
8## Quick Summary
9
10**Goal:** Process and generate multimedia content (images, audio, video, documents) using Google Gemini API via Python scripts.
11
12**Workflow:**
13
141. **Identify Modality** — Match input type to task (analyze, transcribe, extract, generate)
152. **Check Limits** — Inline max 20MB, File API max 2GB; split large audio at 15min chunks
163. **Execute** — Run `gemini_batch_process.py` with appropriate task and files
174. **Post-Process** — Format output as markdown with timestamps, save generated content
18
19**Key Rules:**
20
21- Requires `GEMINI_API_KEY` environment variable
22- Always request specific nodes/files, avoid full-file downloads
23- Use `media_optimizer.py` to compress/split files exceeding limits
24
25**Be skeptical. Apply critical thinking, sequential thinking. Every claim needs traced proof, confidence percentages (Idea should be more than 80%).**
26
27# AI Multimodal
28
29## Purpose
30
31Process audio, images, videos, and documents or generate images/videos using Google Gemini's multimodal API via bundled Python scripts.
32
33## When to Use
34
35- Analyzing images or screenshots (Gemini vision is preferred over Claude's built-in vision for complex tasks)
36- Transcribing audio files (meetings, podcasts, interviews)
37- Extracting data from PDFs, scanned documents, or charts
38- Processing video content (scene detection, temporal Q&A)
39- Generating images with Imagen 4 or videos with Veo 3
40- Converting documents to markdown with visual understanding
41
42## When NOT to Use
43
44- Simple text-only LLM calls -- use Claude directly
45- Reading a file Claude can already read (code, markdown, JSON) -- use `Read` tool
46- Building AI-powered application features -- use `api-design` or `frontend-design`
47- Music composition workflows -- load `references/music-generation.md` only when specifically requested
48- General prompt engineering -- use `ai-artist` skill
49
50## Prerequisites
51
52```bash
53export GEMINI_API_KEY="your-key" # From https://aistudio.google.com/apikey
54pip install google-genai python-dotenv pillow
55python scripts/check_setup.py # Verify setup
56```
57
58Optional: API key rotation for rate limits (set `GEMINI_API_KEY_2`, `GEMINI_API_KEY_3`).
59
60## Workflow
61
62### Step 1: Identify Modality
63
64| Input Type | Task | Command |
65| -------------------- | --------------------- | ----------------------- |
66| Image (PNG/JPG/WEBP) | Analyze, caption, OCR | `--task analyze` |
67| Audio (WAV/MP3/AAC) | Transcribe, summarize | `--task transcribe` |
68| Video (MP4/MOV) | Scene detection, Q&A | `--task analyze` |
69| PDF/Document | Extract tables, forms | `--task extract` |
70| Text prompt | Generate image | `--task generate` |
71| Text prompt | Generate video | `--task generate-video` |
72
73### Step 2: Check Limits
74
75- **Inline upload**: max 20MB
76- **File API**: max 2GB (auto-used for large files)
77- **Audio transcription**: split at 15-minute chunks for full transcript
78- **Video transcription**: extract audio first, then split and transcribe
79- **Formats**: Audio (WAV/MP3/AAC, up to 9.5h), Images (PNG/JPEG/WEBP, up to 3.6k), Video (MP4/MOV, up to 6h), PDF (up to 1k pages)
80
81IF file exceeds limits, use `scripts/media_optimizer.py` to compress/split first.
82
83### Step 3: Execute
84
85**Quick check**: If `gemini` CLI is available, use: `"<prompt>" | gemini -y -m gemini-2.5-flash`
86
87**Standard**: Use the batch processing script:
88
89```bash
90# Analyze media
91python scripts/gemini_batch_process.py --files <file> --task <analyze|transcribe|extract>
92
93# Generate content
94python scripts/gemini_batch_process.py --task generate --prompt "description"
95python scripts/gemini_batch_process.py --task generate-video --prompt "description"
96```
97
98**Stdin support**: `cat image.png | python scripts/gemini_batch_process.py --task analyze --prompt "Describe this"`
99
100### Step 4: Post-Processing
101
102- For transcripts: output in markdown with `[HH:MM:SS -> HH:MM:SS]` timestamps
103- For document extraction: save as structured markdown under `docs/assets/`
104- For generated images/videos: save to working directory with descriptive filename
105
106### Step 5: Verification
107
108- Confirm output matches expected format and completeness
109- For long transcripts: verify no truncation occurred (check chunk boundaries)
110- For generated content: verify quality meets prompt requirements
111
112## Models
113
114| Purpose | Model | Notes |
115| -------------------------- | ------------------------------- | ----------------------- |
116| Analysis (fast) | `gemini-2.5-flash` | Recommended default |
117| Analysis (advanced) | `gemini-2.5-pro` | Complex reasoning tasks |
118| Image generation | `imagen-4.0-generate-001` | Standard quality |
119| Image generation (quality) | `imagen-4.0-ultra-generate-001` | Best quality |
120| Image generation (speed) | `imagen-4.0-fast-generate-001` | Fastest |
121| Video generation | `veo-3.1-generate-preview` | 8s clips with audio |
122
123## Scripts Reference
124
125- **`gemini_batch_process.py`** -- CLI orchestrator for all tasks, auto-resolves API keys and models
126- **`media_optimizer.py`** -- Compress/resize/split media to fit Gemini limits
127- **`document_converter.py`** -- Convert PDFs/images/Office docs to markdown
128- **`check_setup.py`** -- Verify environment, dependencies, and API key
129
130Use `--help` on any script for full options.
131
132## Examples
133
134### Example 1: Transcribe a Meeting Recording
135
136**Input**: 45-minute meeting audio file `meeting-2025-01-15.mp3`
137
138**Steps**:
139
1401. File is >15min, so split first:
141 ```bash
142 python scripts/media_optimizer.py --input meeting-2025-01-15.mp3 --split-duration 900
143 ```
1442. Transcribe each chunk:
145 ```bash
146 python scripts/gemini_batch_process.py --files meeting-part-*.mp3 --task transcribe
147 ```
1483. Output: Markdown file with timestamps, speaker detection, and metadata (duration, topics covered)
149
150### Example 2: Extract Data from a PDF Report
151
152**Input**: Quarterly HR report PDF with tables, charts, and forms
153
154**Steps**:
155
1561. Convert and extract:
157 ```bash
158 python scripts/document_converter.py --input quarterly-report.pdf --output docs/assets/
159 ```
1602. Output: Structured markdown with tables preserved, chart descriptions, and form field values extracted
161
162## Detailed References
163
164Load for in-depth guidance:
165
166| Topic | File |
167| --------------------- | ------------------------------------ |
168| Audio processing | `references/audio-processing.md` |
169| Vision/image analysis | `references/vision-understanding.md` |
170| Image generation | `references/image-generation.md` |
171| Video analysis | `references/video-analysis.md` |
172| Video generation | `references/video-generation.md` |
173| Music generation | `references/music-generation.md` |
174
175## Related Skills
176
177- `ai-artist` -- for prompt engineering and optimization (not media processing)
178- `media-processing` -- for FFmpeg-based audio/video encoding without AI
179- `pdf-to-markdown` -- for simple PDF text extraction without vision AI
180
181---
182
183**IMPORTANT Task Planning Notes (MUST FOLLOW)**
184
185- Always plan and break work into many small todo tasks
186- Always add a final review todo task to verify work quality and identify fixes/enhancements