Gemini Image Generation
Generate images via google/gemini-3.1-flash-image-preview on OpenRouter. Cheap ($0.25/M in, $1.5/M out), fast, good quality.
Quick Start
python3 scripts/generate.py "a watercolor illustration of a cozy café" -o output.png
With reference image (style/character guidance):
python3 scripts/generate.py "same character but waving hello" -o wave.png --ref reference.png
Script path: skills/gemini-image-gen/scripts/generate.py
Requirements
OPENROUTER_API_KEY environment variable (or --api-key flag)
- Python 3.10+ (stdlib only, no pip installs needed)
How It Works
- Calls OpenRouter
/chat/completions with modalities: ["text", "image"]
- Optionally encodes a reference image as base64 in the message
- Extracts generated image from
choices[0].message.images[0].image_url.url (data:image/png;base64,...)
- Decodes and saves to output path
Prompt Engineering Tips (from experience)
Aspect Ratio & Composition
- Gemini respects aspect ratio instructions in the prompt
- For vertical (e.g. phone wallpaper, Xiaohongshu cover): add "vertical composition, 3:4 aspect ratio"
- For horizontal (e.g. banner): add "horizontal composition, 16:9 aspect ratio"
- For square: add "square composition, 1:1 aspect ratio"
- Always specify — without it, Gemini defaults to roughly square and may crop awkwardly
Character Consistency
- When using
--ref, describe the character features explicitly in the prompt AND provide the reference image
- Key details to specify: hair color/style, eye color, clothing, accessories, expression
- Example: "same character from reference: silver-to-ice-blue gradient shoulder-length hair, ice-blue eyes, cream cardigan over light blue shirt, snowflake earring"
- Gemini is decent at maintaining consistency but drifts on small details — always re-specify distinguishing features
Style Control
- Name the art style explicitly: "soft watercolor illustration", "anime cel-shading", "photorealistic", "flat vector", "oil painting"
- For warm/cozy tone: "warm color palette, cream and peach gradient background, bokeh light spots"
- For dark/moody: "dark gradient background, deep navy to black, subtle glow effects"
- Mentioning a well-known art style works: "in the style of Studio Ghibli", "Makoto Shinkai lighting"
Text in Images
- Gemini can render short text in images but it's unreliable for CJK characters
- For English text: works reasonably well if you specify font style ("bold sans-serif", "handwritten script")
- For Chinese/Japanese: avoid — it usually garbles characters. Add text overlays with a separate tool (e.g. ImageMagick, Pillow) instead
Common Pitfalls
- Body proportions: Gemini sometimes compresses/distorts figures. Add "natural human body proportions, do not squash or stretch" for character art
- Hands: Still a weak spot. Minimize visible hands or describe hand pose explicitly
- Multiple subjects: More than 2-3 subjects increases inconsistency. Keep scenes focused
- Batch generation: For generating multiple variations, run the script multiple times — each call is independent. Do NOT ask for "4 options" in one prompt
Sending Images on Feishu
⚠️ Critical: Images must be saved to a path within localRoots (typically your OpenClaw workspace dir). /tmp is NOT whitelisted on Feishu.
# Save to workspace, not /tmp
output_path = "my_image.png" # relative to workspace
# Send via message tool:
# media: "file://<workspace_path>/my_image.png"
# (use 'media' parameter, NOT 'filePath')
After sending, clean up temporary images to avoid workspace clutter.
Advanced: Calling from Python (without CLI)
import os, sys
sys.path.insert(0, "skills/gemini-image-gen/scripts")
from generate import generate
generate(
prompt="a cute robot reading a philosophy book",
output="robot.png",
ref_image=None, # or path to reference image
)
Model Alternatives
| Model |
Cost |
Notes |
google/gemini-3.1-flash-image-preview |
$0.25/$1.5 per M tokens |
Default. Best balance of cost and quality |
google/gemini-3.1-pro-preview |
$2/$12 per M tokens |
Higher quality but 8x more expensive |
openai/gpt-image-1 |
varies |
OpenAI's image model, different API format — not supported by this script |
Troubleshooting
- "No image in response": Check
.debug.json file created alongside output. Usually means the prompt triggered safety filters or the model returned text-only.
- Garbled/distorted output: Try rephrasing. Add "high quality, detailed" and be more specific about composition.
- API error 429: Rate limited. Wait 30s and retry.
- API error 402: Insufficient credits on OpenRouter.
1---2name: gemini-image-gen-23description: Generate images using Google Gemini via OpenRouter API. Supports text-to-image and reference-image-guided generation. Use when the user asks to generate, create, draw, or design images/illustrations/covers/avatars.4---5
6# Gemini Image Generation
7
8Generate images via `google/gemini-3.1-flash-image-preview` on OpenRouter. Cheap ($0.25/M in, $1.5/M out), fast, good quality.
9
10## Quick Start
11
12```bash
13python3 scripts/generate.py "a watercolor illustration of a cozy café" -o output.png
14```
15
16With reference image (style/character guidance):
17```bash
18python3 scripts/generate.py "same character but waving hello" -o wave.png --ref reference.png
19```
20
21Script path: `skills/gemini-image-gen/scripts/generate.py`
22
23## Requirements
24
25- `OPENROUTER_API_KEY` environment variable (or `--api-key` flag)
26- Python 3.10+ (stdlib only, no pip installs needed)
27
28## How It Works
29
301. Calls OpenRouter `/chat/completions` with `modalities: ["text", "image"]`
312. Optionally encodes a reference image as base64 in the message
323. Extracts generated image from `choices[0].message.images[0].image_url.url` (data:image/png;base64,...)
334. Decodes and saves to output path
34
35## Prompt Engineering Tips (from experience)
36
37### Aspect Ratio & Composition
38- Gemini respects aspect ratio instructions in the prompt
39- For vertical (e.g. phone wallpaper, Xiaohongshu cover): add "vertical composition, 3:4 aspect ratio"
40- For horizontal (e.g. banner): add "horizontal composition, 16:9 aspect ratio"
41- For square: add "square composition, 1:1 aspect ratio"
42- **Always specify** — without it, Gemini defaults to roughly square and may crop awkwardly
43
44### Character Consistency
45- When using `--ref`, describe the character features explicitly in the prompt AND provide the reference image
46- Key details to specify: hair color/style, eye color, clothing, accessories, expression
47- Example: "same character from reference: silver-to-ice-blue gradient shoulder-length hair, ice-blue eyes, cream cardigan over light blue shirt, snowflake earring"
48- Gemini is decent at maintaining consistency but drifts on small details — always re-specify distinguishing features
49
50### Style Control
51- Name the art style explicitly: "soft watercolor illustration", "anime cel-shading", "photorealistic", "flat vector", "oil painting"
52- For warm/cozy tone: "warm color palette, cream and peach gradient background, bokeh light spots"
53- For dark/moody: "dark gradient background, deep navy to black, subtle glow effects"
54- Mentioning a well-known art style works: "in the style of Studio Ghibli", "Makoto Shinkai lighting"
55
56### Text in Images
57- Gemini can render short text in images but it's unreliable for CJK characters
58- For English text: works reasonably well if you specify font style ("bold sans-serif", "handwritten script")
59- For Chinese/Japanese: **avoid** — it usually garbles characters. Add text overlays with a separate tool (e.g. ImageMagick, Pillow) instead
60
61### Common Pitfalls
62- **Body proportions**: Gemini sometimes compresses/distorts figures. Add "natural human body proportions, do not squash or stretch" for character art
63- **Hands**: Still a weak spot. Minimize visible hands or describe hand pose explicitly
64- **Multiple subjects**: More than 2-3 subjects increases inconsistency. Keep scenes focused
65- **Batch generation**: For generating multiple variations, run the script multiple times — each call is independent. Do NOT ask for "4 options" in one prompt
66
67## Sending Images on Feishu
68
69⚠️ **Critical**: Images must be saved to a path within `localRoots` (typically your OpenClaw workspace dir). `/tmp` is NOT whitelisted on Feishu.
70
71```python
72# Save to workspace, not /tmp
73output_path = "my_image.png" # relative to workspace
74
75# Send via message tool:
76# media: "file://<workspace_path>/my_image.png"
77# (use 'media' parameter, NOT 'filePath')
78```
79
80After sending, clean up temporary images to avoid workspace clutter.
81
82## Advanced: Calling from Python (without CLI)
83
84```python
85import os, sys
86sys.path.insert(0, "skills/gemini-image-gen/scripts")
87from generate import generate
88
89generate(
90 prompt="a cute robot reading a philosophy book",
91 output="robot.png",
92 ref_image=None, # or path to reference image
93)
94```
95
96## Model Alternatives
97
98| Model | Cost | Notes |
99|-------|------|-------|
100| `google/gemini-3.1-flash-image-preview` | $0.25/$1.5 per M tokens | **Default.** Best balance of cost and quality |
101| `google/gemini-3.1-pro-preview` | $2/$12 per M tokens | Higher quality but 8x more expensive |
102| `openai/gpt-image-1` | varies | OpenAI's image model, different API format — not supported by this script |
103
104## Troubleshooting
105
106- **"No image in response"**: Check `.debug.json` file created alongside output. Usually means the prompt triggered safety filters or the model returned text-only.
107- **Garbled/distorted output**: Try rephrasing. Add "high quality, detailed" and be more specific about composition.
108- **API error 429**: Rate limited. Wait 30s and retry.
109- **API error 402**: Insufficient credits on OpenRouter.