Image Generation (AI SDK)
Official API-based image generation. Supports OpenAI, Azure OpenAI, Google, OpenRouter, DashScope (阿里通义万象), MiniMax, Jimeng (即梦), Seedream (豆包) and Replicate providers.
Script Directory
Agent Execution:
{baseDir} = this SKILL.md file's directory
- Script path =
{baseDir}/scripts/main.ts
- Resolve
${BUN_X} runtime: if bun installed → bun; if npx available → npx -y bun; else suggest installing bun
Step 0: Load Preferences ⛔ BLOCKING
CRITICAL: This step MUST complete BEFORE any image generation. Do NOT skip or defer.
Check EXTEND.md existence (priority: project → user):
# macOS, Linux, WSL, Git Bash
test -f .floracat-skills/floracat-image-gen/EXTEND.md && echo "project"
test -f "$HOME/.floracat-skills/floracat-image-gen/EXTEND.md" && echo "user"
# PowerShell (Windows)
if (Test-Path .floracat-skills/floracat-image-gen/EXTEND.md) { "project" }
if (Test-Path "$HOME/.floracat-skills/floracat-image-gen/EXTEND.md") { "user" }
| Result |
Action |
| Found |
Load, parse, apply settings. If default_model.[provider] is null → ask model only (Flow 2) |
| Not found |
⛔ Run first-time setup (references/config/first-time-setup.md) → Save EXTEND.md → Then continue |
CRITICAL: If not found, complete the full setup (provider + model + quality + save location) using AskUserQuestion BEFORE generating any images. Generation is BLOCKED until EXTEND.md is created.
| Path |
Location |
.floracat-skills/floracat-image-gen/EXTEND.md |
Project directory |
$HOME/.floracat-skills/floracat-image-gen/EXTEND.md |
User home |
EXTEND.md Supports: Default provider | Default quality | Default aspect ratio | Default image size | Default models | Batch worker cap | Provider-specific batch limits
Schema: references/config/preferences-schema.md
Usage
# Basic
${BUN_X} {baseDir}/scripts/main.ts --prompt "A cat" --image cat.png
# With aspect ratio
${BUN_X} {baseDir}/scripts/main.ts --prompt "A landscape" --image out.png --ar 16:9
# High quality
${BUN_X} {baseDir}/scripts/main.ts --prompt "A cat" --image out.png --quality 2k
# From prompt files
${BUN_X} {baseDir}/scripts/main.ts --promptfiles system.md content.md --image out.png
# With reference images (Google, OpenAI, Azure OpenAI, OpenRouter, Replicate, MiniMax, or Seedream 4.0/4.5/5.0)
${BUN_X} {baseDir}/scripts/main.ts --prompt "Make blue" --image out.png --ref source.png
# With reference images (explicit provider/model)
${BUN_X} {baseDir}/scripts/main.ts --prompt "Make blue" --image out.png --provider google --model gemini-3-pro-image-preview --ref source.png
# Azure OpenAI (model means deployment name)
${BUN_X} {baseDir}/scripts/main.ts --prompt "A cat" --image out.png --provider azure --model gpt-image-1.5
# OpenRouter (recommended default model)
${BUN_X} {baseDir}/scripts/main.ts --prompt "A cat" --image out.png --provider openrouter
# OpenRouter with reference images
${BUN_X} {baseDir}/scripts/main.ts --prompt "Make blue" --image out.png --provider openrouter --model google/gemini-3.1-flash-image-preview --ref source.png
# Specific provider
${BUN_X} {baseDir}/scripts/main.ts --prompt "A cat" --image out.png --provider openai
# DashScope (阿里通义万象)
${BUN_X} {baseDir}/scripts/main.ts --prompt "一只可爱的猫" --image out.png --provider dashscope
# DashScope Qwen-Image 2.0 Pro (recommended for custom sizes and text rendering)
${BUN_X} {baseDir}/scripts/main.ts --prompt "为咖啡品牌设计一张 21:9 横幅海报,包含清晰中文标题" --image out.png --provider dashscope --model qwen-image-2.0-pro --size 2048x872
# DashScope legacy Qwen fixed-size model
${BUN_X} {baseDir}/scripts/main.ts --prompt "一张电影感海报" --image out.png --provider dashscope --model qwen-image-max --size 1664x928
# MiniMax
${BUN_X} {baseDir}/scripts/main.ts --prompt "A fashion editorial portrait by a bright studio window" --image out.jpg --provider minimax
# MiniMax with subject reference (best for character/portrait consistency)
${BUN_X} {baseDir}/scripts/main.ts --prompt "A girl stands by the library window, cinematic lighting" --image out.jpg --provider minimax --model image-01 --ref portrait.png --ar 16:9
# MiniMax with custom size (documented for image-01)
${BUN_X} {baseDir}/scripts/main.ts --prompt "A cinematic poster" --image out.jpg --provider minimax --model image-01 --size 1536x1024
# Replicate (google/nano-banana-pro)
${BUN_X} {baseDir}/scripts/main.ts --prompt "A cat" --image out.png --provider replicate
# Replicate with specific model
${BUN_X} {baseDir}/scripts/main.ts --prompt "A cat" --image out.png --provider replicate --model google/nano-banana
# Batch mode with saved prompt files
${BUN_X} {baseDir}/scripts/main.ts --batchfile batch.json
# Batch mode with explicit worker count
${BUN_X} {baseDir}/scripts/main.ts --batchfile batch.json --jobs 4 --json
Batch File Format
{
"jobs": 4,
"tasks": [
{
"id": "hero",
"promptFiles": ["prompts/hero.md"],
"image": "out/hero.png",
"provider": "replicate",
"model": "google/nano-banana-pro",
"ar": "16:9",
"quality": "2k"
},
{
"id": "diagram",
"promptFiles": ["prompts/diagram.md"],
"image": "out/diagram.png",
"ref": ["references/original.png"]
}
]
}
Paths in promptFiles, image, and ref are resolved relative to the batch file's directory. jobs is optional (overridden by CLI --jobs). Top-level array format (without jobs wrapper) is also accepted.
Options
| Option |
Description |
--prompt <text>, -p |
Prompt text |
--promptfiles <files...> |
Read prompt from files (concatenated) |
--image <path> |
Output image path (required in single-image mode) |
--batchfile <path> |
JSON batch file for multi-image generation |
--jobs <count> |
Worker count for batch mode (default: auto, max from config, built-in default 10) |
--provider google|openai|azure|openrouter|dashscope|minimax|jimeng|seedream|replicate |
Force provider (default: auto-detect) |
--model <id>, -m |
Model ID (Google: gemini-3-pro-image-preview; OpenAI: gpt-image-1.5; Azure: deployment name such as gpt-image-1.5 or image-prod; OpenRouter: google/gemini-3.1-flash-image-preview; DashScope: qwen-image-2.0-pro; MiniMax: image-01) |
--ar <ratio> |
Aspect ratio (e.g., 16:9, 1:1, 4:3) |
--size <WxH> |
Size (e.g., 1024x1024) |
--quality normal|2k |
Quality preset (default: 2k) |
--imageSize 1K|2K|4K |
Image size for Google/OpenRouter (default: from quality) |
--ref <files...> |
Reference images. Supported by Google multimodal, OpenAI GPT Image edits, Azure OpenAI edits (PNG/JPG only), OpenRouter multimodal models, Replicate, MiniMax subject-reference, and Seedream 5.0/4.5/4.0. Not supported by Jimeng, Seedream 3.0, or removed SeedEdit 3.0 |
--n <count> |
Number of images |
--json |
JSON output |
Environment Variables
| Variable |
Description |
OPENAI_API_KEY |
OpenAI API key |
AZURE_OPENAI_API_KEY |
Azure OpenAI API key |
OPENROUTER_API_KEY |
OpenRouter API key |
GOOGLE_API_KEY |
Google API key |
DASHSCOPE_API_KEY |
DashScope API key (阿里云) |
MINIMAX_API_KEY |
MiniMax API key |
REPLICATE_API_TOKEN |
Replicate API token |
JIMENG_ACCESS_KEY_ID |
Jimeng (即梦) Volcengine access key |
JIMENG_SECRET_ACCESS_KEY |
Jimeng (即梦) Volcengine secret key |
ARK_API_KEY |
Seedream (豆包) Volcengine ARK API key |
OPENAI_IMAGE_MODEL |
OpenAI model override |
AZURE_OPENAI_DEPLOYMENT |
Azure default deployment name |
AZURE_OPENAI_IMAGE_MODEL |
Backward-compatible alias for Azure default deployment/model name |
OPENROUTER_IMAGE_MODEL |
OpenRouter model override (default: google/gemini-3.1-flash-image-preview) |
GOOGLE_IMAGE_MODEL |
Google model override |
DASHSCOPE_IMAGE_MODEL |
DashScope model override (default: qwen-image-2.0-pro) |
MINIMAX_IMAGE_MODEL |
MiniMax model override (default: image-01) |
REPLICATE_IMAGE_MODEL |
Replicate model override (default: google/nano-banana-pro) |
JIMENG_IMAGE_MODEL |
Jimeng model override (default: jimeng_t2i_v40) |
SEEDREAM_IMAGE_MODEL |
Seedream model override (default: doubao-seedream-5-0-260128) |
OPENAI_BASE_URL |
Custom OpenAI endpoint |
AZURE_OPENAI_BASE_URL |
Azure resource endpoint or deployment endpoint |
AZURE_API_VERSION |
Azure image API version (default: 2025-04-01-preview) |
OPENROUTER_BASE_URL |
Custom OpenRouter endpoint (default: https://openrouter.ai/api/v1) |
OPENROUTER_HTTP_REFERER |
Optional app/site URL for OpenRouter attribution |
OPENROUTER_TITLE |
Optional app name for OpenRouter attribution |
GOOGLE_BASE_URL |
Custom Google endpoint |
DASHSCOPE_BASE_URL |
Custom DashScope endpoint |
MINIMAX_BASE_URL |
Custom MiniMax endpoint (default: https://api.minimax.io) |
REPLICATE_BASE_URL |
Custom Replicate endpoint |
JIMENG_BASE_URL |
Custom Jimeng endpoint (default: https://visual.volcengineapi.com) |
JIMENG_REGION |
Jimeng region (default: cn-north-1) |
SEEDREAM_BASE_URL |
Custom Seedream endpoint (default: https://ark.cn-beijing.volces.com/api/v3) |
FLORACAT_IMAGE_GEN_MAX_WORKERS |
Override batch worker cap |
FLORACAT_IMAGE_GEN_<PROVIDER>_CONCURRENCY |
Override provider concurrency, e.g. FLORACAT_IMAGE_GEN_REPLICATE_CONCURRENCY |
FLORACAT_IMAGE_GEN_<PROVIDER>_START_INTERVAL_MS |
Override provider start gap, e.g. FLORACAT_IMAGE_GEN_REPLICATE_START_INTERVAL_MS |
Load Priority: CLI args > EXTEND.md > env vars > <cwd>/.floracat-skills/.env > ~/.floracat-skills/.env
Model Resolution
Model priority (highest → lowest), applies to all providers:
- CLI flag:
--model <id>
- EXTEND.md:
default_model.[provider]
- Env var:
<PROVIDER>_IMAGE_MODEL (e.g., GOOGLE_IMAGE_MODEL)
- Built-in default
For Azure, --model / default_model.azure should be the Azure deployment name. AZURE_OPENAI_DEPLOYMENT is the preferred env var, and AZURE_OPENAI_IMAGE_MODEL remains as a backward-compatible alias.
EXTEND.md overrides env vars. If both EXTEND.md default_model.google: "gemini-3-pro-image-preview" and env var GOOGLE_IMAGE_MODEL=gemini-3.1-flash-image-preview exist, EXTEND.md wins.
Agent MUST display model info before each generation:
- Show:
Using [provider] / [model]
- Show switch hint:
Switch model: --model <id> | EXTEND.md default_model.[provider] | env <PROVIDER>_IMAGE_MODEL
DashScope Models
Use --model qwen-image-2.0-pro or set default_model.dashscope / DASHSCOPE_IMAGE_MODEL when the user wants official Qwen-Image behavior.
Official DashScope model families:
qwen-image-2.0-pro, qwen-image-2.0-pro-2026-03-03, qwen-image-2.0, qwen-image-2.0-2026-03-03
- Free-form
size in 宽*高 format
- Total pixels must stay between
512*512 and 2048*2048
- Default size is approximately
1024*1024
- Best choice for custom ratios such as
21:9 and text-heavy Chinese/English layouts
qwen-image-max, qwen-image-max-2025-12-30, qwen-image-plus, qwen-image-plus-2026-01-09, qwen-image
- Fixed sizes only:
1664*928, 1472*1104, 1328*1328, 1104*1472, 928*1664
- Default size is
1664*928
qwen-image currently has the same capability as qwen-image-plus
- Legacy DashScope models such as
z-image-turbo, z-image-ultra, wanx-v1
- Keep using them only when the user explicitly asks for legacy behavior or compatibility
When translating CLI args into DashScope behavior:
--size wins over --ar
- For
qwen-image-2.0*, prefer explicit --size; otherwise infer from --ar and use the official recommended resolutions below
- For
qwen-image-max/plus/image, only use the five official fixed sizes; if the requested ratio is not covered, switch to qwen-image-2.0-pro
--quality is a floracat-image-gen compatibility preset, not a native DashScope API field. Mapping normal / 2k onto the qwen-image-2.0* table below is an implementation inference, not an official API guarantee
Recommended qwen-image-2.0* sizes for common aspect ratios:
| Ratio |
normal |
2k |
1:1 |
1024*1024 |
1536*1536 |
2:3 |
768*1152 |
1024*1536 |
3:2 |
1152*768 |
1536*1024 |
3:4 |
960*1280 |
1080*1440 |
4:3 |
1280*960 |
1440*1080 |
9:16 |
720*1280 |
1080*1920 |
16:9 |
1280*720 |
1920*1080 |
21:9 |
1344*576 |
2048*872 |
DashScope official APIs also expose negative_prompt, prompt_extend, and watermark, but floracat-image-gen does not expose them as dedicated CLI flags today.
Official references:
MiniMax Models
Use --model image-01 or set default_model.minimax / MINIMAX_IMAGE_MODEL when the user wants MiniMax image generation.
Official MiniMax image model options currently documented in the API reference:
image-01 (recommended default)
- Supports text-to-image and subject-reference image generation
- Supports official
aspect_ratio values: 1:1, 16:9, 4:3, 3:2, 2:3, 3:4, 9:16, 21:9
- Supports documented custom
width / height output sizes when using --size <WxH>
width and height must both be between 512 and 2048, and both must be divisible by 8
image-01-live
- Lower-latency variant
- Use
--ar for sizing; MiniMax documents custom width / height as only effective for image-01
MiniMax subject reference notes:
--ref files are sent as MiniMax subject_reference
- MiniMax docs currently describe
subject_reference[].type as character
- Official docs say
image_file supports public URLs or Base64 Data URLs; floracat-image-gen sends local refs as Data URLs
- Official docs recommend front-facing portrait references in JPG/JPEG/PNG under 10MB
Official references:
OpenRouter Models
Use full OpenRouter model IDs, e.g.:
google/gemini-3.1-flash-image-preview (recommended, supports image output and reference-image workflows)
google/gemini-2.5-flash-image-preview
black-forest-labs/flux.2-pro
- Other OpenRouter image-capable model IDs
Notes:
- OpenRouter image generation uses
/chat/completions, not the OpenAI /images endpoints
- If
--ref is used, choose a multimodal model that supports image input and image output
--imageSize maps to OpenRouter imageGenerationOptions.size; --size <WxH> is converted to the nearest OpenRouter size and inferred aspect ratio when possible
Replicate Models
Supported model formats:
owner/name (recommended for official models), e.g. google/nano-banana-pro
owner/name:version (community models by version), e.g. stability-ai/sdxl:<version>
Examples:
# Use Replicate default model
${BUN_X} {baseDir}/scripts/main.ts --prompt "A cat" --image out.png --provider replicate
# Override model explicitly
${BUN_X} {baseDir}/scripts/main.ts --prompt "A cat" --image out.png --provider replicate --model google/nano-banana
Provider Selection
--ref provided + no --provider → auto-select Google first, then OpenAI, then Azure, then OpenRouter, then Replicate, then Seedream, then MiniMax (MiniMax subject reference is more specialized toward character/portrait consistency)
--provider specified → use it (if --ref, must be google, openai, azure, openrouter, replicate, seedream, or minimax)
- Only one API key available → use that provider
- Multiple available → default to Google
Quality Presets
| Preset |
Google imageSize |
OpenAI Size |
OpenRouter size |
Replicate resolution |
Use Case |
normal |
1K |
1024px |
1K |
1K |
Quick previews |
2k (default) |
2K |
2048px |
2K |
2K |
Covers, illustrations, infographics |
Google/OpenRouter imageSize: Can be overridden with --imageSize 1K|2K|4K
Aspect Ratios
Supported: 1:1, 16:9, 9:16, 4:3, 3:4, 2.35:1
- Google multimodal: uses
imageConfig.aspectRatio
- OpenAI: maps to closest supported size
- OpenRouter: sends
imageGenerationOptions.aspect_ratio; if only --size <WxH> is given, aspect ratio is inferred automatically
- Replicate: passes
aspect_ratio to model; when --ref is provided without --ar, defaults to match_input_image
- MiniMax: sends official
aspect_ratio values directly; if --size <WxH> is given without --ar, width / height are sent for image-01
Generation Mode
Default: Sequential generation.
Batch Parallel Generation: When --batchfile contains 2 or more pending tasks, the script automatically enables parallel generation.
| Mode |
When to Use |
| Sequential (default) |
Normal usage, single images, small batches |
| Parallel batch |
Batch mode with 2+ tasks |
Execution choice:
| Situation |
Preferred approach |
Why |
| One image, or 1-2 simple images |
Sequential |
Lower coordination overhead and easier debugging |
| Multiple images already have saved prompt files |
Batch (--batchfile) |
Reuses finalized prompts, applies shared throttling/retries, and gives predictable throughput |
| Each image still needs separate reasoning, prompt writing, or style exploration |
Subagents |
The work is still exploratory, so each image may need independent analysis before generation |
Output comes from floracat-rednote with outline.md + prompts/ |
Batch (build-batch.ts -> --batchfile) |
That workflow already produces prompt files, so direct batch execution is the intended path |
Rule of thumb:
- Prefer batch over subagents once prompt files are already saved and the task is "generate all of these"
- Use subagents only when generation is coupled with per-image thinking, rewriting, or divergent creative exploration
Parallel behavior:
- Default worker count is automatic, capped by config, built-in default 10
- Provider-specific throttling is applied only in batch mode, and the built-in defaults are tuned for faster throughput while still avoiding obvious RPM bursts
- You can override worker count with
--jobs <count>
- Each image retries automatically up to 3 attempts
- Final output includes success count, failure count, and per-image failure reasons
Error Handling
- Missing API key → error with setup instructions
- Generation failure → auto-retry up to 3 attempts per image
- Invalid aspect ratio → warning, proceed with default
- Reference images with unsupported provider/model → error with fix hint
Extension Support
Custom configurations via EXTEND.md. See Preferences section for paths and supported options.
1---2name: floracat-image-gen3description: Generate final PNG images from prompt text or prompt files using configured image-generation providers. Supports provider and model selection, aspect ratios, image size, quality presets, reference images, and sequential or batch generation. Use when the user asks to generate, draw, or create images, or when another skill needs direct PNG output instead of HTML/SVG previews.4---56# Image Generation (AI SDK)78Official API-based image generation. Supports OpenAI, Azure OpenAI, Google, OpenRouter, DashScope (阿里通义万象), MiniMax, Jimeng (即梦), Seedream (豆包) and Replicate providers.910## Script Directory1112**Agent Execution**:131. `{baseDir}` = this SKILL.md file's directory142. Script path = `{baseDir}/scripts/main.ts`153. Resolve `${BUN_X}` runtime: if `bun` installed → `bun`; if `npx` available → `npx -y bun`; else suggest installing bun1617## Step 0: Load Preferences ⛔ BLOCKING1819**CRITICAL**: This step MUST complete BEFORE any image generation. Do NOT skip or defer.2021Check EXTEND.md existence (priority: project → user):2223```bash24# macOS, Linux, WSL, Git Bash25test -f .floracat-skills/floracat-image-gen/EXTEND.md && echo "project"26test -f "$HOME/.floracat-skills/floracat-image-gen/EXTEND.md" && echo "user"27```2829```powershell30# PowerShell (Windows)31if (Test-Path .floracat-skills/floracat-image-gen/EXTEND.md) { "project" }32if (Test-Path "$HOME/.floracat-skills/floracat-image-gen/EXTEND.md") { "user" }33```3435| Result | Action |36|--------|--------|37| Found | Load, parse, apply settings. If `default_model.[provider]` is null → ask model only (Flow 2) |38| Not found | ⛔ Run first-time setup ([references/config/first-time-setup.md](references/config/first-time-setup.md)) → Save EXTEND.md → Then continue |3940**CRITICAL**: If not found, complete the full setup (provider + model + quality + save location) using AskUserQuestion BEFORE generating any images. Generation is BLOCKED until EXTEND.md is created.4142| Path | Location |43|------|----------|44| `.floracat-skills/floracat-image-gen/EXTEND.md` | Project directory |45| `$HOME/.floracat-skills/floracat-image-gen/EXTEND.md` | User home |4647**EXTEND.md Supports**: Default provider | Default quality | Default aspect ratio | Default image size | Default models | Batch worker cap | Provider-specific batch limits4849Schema: `references/config/preferences-schema.md`5051## Usage5253```bash54# Basic55${BUN_X} {baseDir}/scripts/main.ts --prompt "A cat" --image cat.png5657# With aspect ratio58${BUN_X} {baseDir}/scripts/main.ts --prompt "A landscape" --image out.png --ar 16:95960# High quality61${BUN_X} {baseDir}/scripts/main.ts --prompt "A cat" --image out.png --quality 2k6263# From prompt files64${BUN_X} {baseDir}/scripts/main.ts --promptfiles system.md content.md --image out.png6566# With reference images (Google, OpenAI, Azure OpenAI, OpenRouter, Replicate, MiniMax, or Seedream 4.0/4.5/5.0)67${BUN_X} {baseDir}/scripts/main.ts --prompt "Make blue" --image out.png --ref source.png6869# With reference images (explicit provider/model)70${BUN_X} {baseDir}/scripts/main.ts --prompt "Make blue" --image out.png --provider google --model gemini-3-pro-image-preview --ref source.png7172# Azure OpenAI (model means deployment name)73${BUN_X} {baseDir}/scripts/main.ts --prompt "A cat" --image out.png --provider azure --model gpt-image-1.57475# OpenRouter (recommended default model)76${BUN_X} {baseDir}/scripts/main.ts --prompt "A cat" --image out.png --provider openrouter7778# OpenRouter with reference images79${BUN_X} {baseDir}/scripts/main.ts --prompt "Make blue" --image out.png --provider openrouter --model google/gemini-3.1-flash-image-preview --ref source.png8081# Specific provider82${BUN_X} {baseDir}/scripts/main.ts --prompt "A cat" --image out.png --provider openai8384# DashScope (阿里通义万象)85${BUN_X} {baseDir}/scripts/main.ts --prompt "一只可爱的猫" --image out.png --provider dashscope8687# DashScope Qwen-Image 2.0 Pro (recommended for custom sizes and text rendering)88${BUN_X} {baseDir}/scripts/main.ts --prompt "为咖啡品牌设计一张 21:9 横幅海报,包含清晰中文标题" --image out.png --provider dashscope --model qwen-image-2.0-pro --size 2048x8728990# DashScope legacy Qwen fixed-size model91${BUN_X} {baseDir}/scripts/main.ts --prompt "一张电影感海报" --image out.png --provider dashscope --model qwen-image-max --size 1664x9289293# MiniMax94${BUN_X} {baseDir}/scripts/main.ts --prompt "A fashion editorial portrait by a bright studio window" --image out.jpg --provider minimax9596# MiniMax with subject reference (best for character/portrait consistency)97${BUN_X} {baseDir}/scripts/main.ts --prompt "A girl stands by the library window, cinematic lighting" --image out.jpg --provider minimax --model image-01 --ref portrait.png --ar 16:99899# MiniMax with custom size (documented for image-01)100${BUN_X} {baseDir}/scripts/main.ts --prompt "A cinematic poster" --image out.jpg --provider minimax --model image-01 --size 1536x1024101102# Replicate (google/nano-banana-pro)103${BUN_X} {baseDir}/scripts/main.ts --prompt "A cat" --image out.png --provider replicate104105# Replicate with specific model106${BUN_X} {baseDir}/scripts/main.ts --prompt "A cat" --image out.png --provider replicate --model google/nano-banana107108# Batch mode with saved prompt files109${BUN_X} {baseDir}/scripts/main.ts --batchfile batch.json110111# Batch mode with explicit worker count112${BUN_X} {baseDir}/scripts/main.ts --batchfile batch.json --jobs 4 --json113```114115### Batch File Format116117```json118{119 "jobs": 4,120 "tasks": [121 {122 "id": "hero",123 "promptFiles": ["prompts/hero.md"],124 "image": "out/hero.png",125 "provider": "replicate",126 "model": "google/nano-banana-pro",127 "ar": "16:9",128 "quality": "2k"129 },130 {131 "id": "diagram",132 "promptFiles": ["prompts/diagram.md"],133 "image": "out/diagram.png",134 "ref": ["references/original.png"]135 }136 ]137}138```139140Paths in `promptFiles`, `image`, and `ref` are resolved relative to the batch file's directory. `jobs` is optional (overridden by CLI `--jobs`). Top-level array format (without `jobs` wrapper) is also accepted.141142## Options143144| Option | Description |145|--------|-------------|146| `--prompt <text>`, `-p` | Prompt text |147| `--promptfiles <files...>` | Read prompt from files (concatenated) |148| `--image <path>` | Output image path (required in single-image mode) |149| `--batchfile <path>` | JSON batch file for multi-image generation |150| `--jobs <count>` | Worker count for batch mode (default: auto, max from config, built-in default 10) |151| `--provider google\|openai\|azure\|openrouter\|dashscope\|minimax\|jimeng\|seedream\|replicate` | Force provider (default: auto-detect) |152| `--model <id>`, `-m` | Model ID (Google: `gemini-3-pro-image-preview`; OpenAI: `gpt-image-1.5`; Azure: deployment name such as `gpt-image-1.5` or `image-prod`; OpenRouter: `google/gemini-3.1-flash-image-preview`; DashScope: `qwen-image-2.0-pro`; MiniMax: `image-01`) |153| `--ar <ratio>` | Aspect ratio (e.g., `16:9`, `1:1`, `4:3`) |154| `--size <WxH>` | Size (e.g., `1024x1024`) |155| `--quality normal\|2k` | Quality preset (default: `2k`) |156| `--imageSize 1K\|2K\|4K` | Image size for Google/OpenRouter (default: from quality) |157| `--ref <files...>` | Reference images. Supported by Google multimodal, OpenAI GPT Image edits, Azure OpenAI edits (PNG/JPG only), OpenRouter multimodal models, Replicate, MiniMax subject-reference, and Seedream 5.0/4.5/4.0. Not supported by Jimeng, Seedream 3.0, or removed SeedEdit 3.0 |158| `--n <count>` | Number of images |159| `--json` | JSON output |160161## Environment Variables162163| Variable | Description |164|----------|-------------|165| `OPENAI_API_KEY` | OpenAI API key |166| `AZURE_OPENAI_API_KEY` | Azure OpenAI API key |167| `OPENROUTER_API_KEY` | OpenRouter API key |168| `GOOGLE_API_KEY` | Google API key |169| `DASHSCOPE_API_KEY` | DashScope API key (阿里云) |170| `MINIMAX_API_KEY` | MiniMax API key |171| `REPLICATE_API_TOKEN` | Replicate API token |172| `JIMENG_ACCESS_KEY_ID` | Jimeng (即梦) Volcengine access key |173| `JIMENG_SECRET_ACCESS_KEY` | Jimeng (即梦) Volcengine secret key |174| `ARK_API_KEY` | Seedream (豆包) Volcengine ARK API key |175| `OPENAI_IMAGE_MODEL` | OpenAI model override |176| `AZURE_OPENAI_DEPLOYMENT` | Azure default deployment name |177| `AZURE_OPENAI_IMAGE_MODEL` | Backward-compatible alias for Azure default deployment/model name |178| `OPENROUTER_IMAGE_MODEL` | OpenRouter model override (default: `google/gemini-3.1-flash-image-preview`) |179| `GOOGLE_IMAGE_MODEL` | Google model override |180| `DASHSCOPE_IMAGE_MODEL` | DashScope model override (default: `qwen-image-2.0-pro`) |181| `MINIMAX_IMAGE_MODEL` | MiniMax model override (default: `image-01`) |182| `REPLICATE_IMAGE_MODEL` | Replicate model override (default: google/nano-banana-pro) |183| `JIMENG_IMAGE_MODEL` | Jimeng model override (default: jimeng_t2i_v40) |184| `SEEDREAM_IMAGE_MODEL` | Seedream model override (default: doubao-seedream-5-0-260128) |185| `OPENAI_BASE_URL` | Custom OpenAI endpoint |186| `AZURE_OPENAI_BASE_URL` | Azure resource endpoint or deployment endpoint |187| `AZURE_API_VERSION` | Azure image API version (default: `2025-04-01-preview`) |188| `OPENROUTER_BASE_URL` | Custom OpenRouter endpoint (default: `https://openrouter.ai/api/v1`) |189| `OPENROUTER_HTTP_REFERER` | Optional app/site URL for OpenRouter attribution |190| `OPENROUTER_TITLE` | Optional app name for OpenRouter attribution |191| `GOOGLE_BASE_URL` | Custom Google endpoint |192| `DASHSCOPE_BASE_URL` | Custom DashScope endpoint |193| `MINIMAX_BASE_URL` | Custom MiniMax endpoint (default: `https://api.minimax.io`) |194| `REPLICATE_BASE_URL` | Custom Replicate endpoint |195| `JIMENG_BASE_URL` | Custom Jimeng endpoint (default: `https://visual.volcengineapi.com`) |196| `JIMENG_REGION` | Jimeng region (default: `cn-north-1`) |197| `SEEDREAM_BASE_URL` | Custom Seedream endpoint (default: `https://ark.cn-beijing.volces.com/api/v3`) |198| `FLORACAT_IMAGE_GEN_MAX_WORKERS` | Override batch worker cap |199| `FLORACAT_IMAGE_GEN_<PROVIDER>_CONCURRENCY` | Override provider concurrency, e.g. `FLORACAT_IMAGE_GEN_REPLICATE_CONCURRENCY` |200| `FLORACAT_IMAGE_GEN_<PROVIDER>_START_INTERVAL_MS` | Override provider start gap, e.g. `FLORACAT_IMAGE_GEN_REPLICATE_START_INTERVAL_MS` |201202**Load Priority**: CLI args > EXTEND.md > env vars > `<cwd>/.floracat-skills/.env` > `~/.floracat-skills/.env`203204## Model Resolution205206Model priority (highest → lowest), applies to all providers:2072081. CLI flag: `--model <id>`2092. EXTEND.md: `default_model.[provider]`2103. Env var: `<PROVIDER>_IMAGE_MODEL` (e.g., `GOOGLE_IMAGE_MODEL`)2114. Built-in default212213For Azure, `--model` / `default_model.azure` should be the Azure deployment name. `AZURE_OPENAI_DEPLOYMENT` is the preferred env var, and `AZURE_OPENAI_IMAGE_MODEL` remains as a backward-compatible alias.214215**EXTEND.md overrides env vars**. If both EXTEND.md `default_model.google: "gemini-3-pro-image-preview"` and env var `GOOGLE_IMAGE_MODEL=gemini-3.1-flash-image-preview` exist, EXTEND.md wins.216217**Agent MUST display model info** before each generation:218- Show: `Using [provider] / [model]`219- Show switch hint: `Switch model: --model <id> | EXTEND.md default_model.[provider] | env <PROVIDER>_IMAGE_MODEL`220221### DashScope Models222223Use `--model qwen-image-2.0-pro` or set `default_model.dashscope` / `DASHSCOPE_IMAGE_MODEL` when the user wants official Qwen-Image behavior.224225Official DashScope model families:226227- `qwen-image-2.0-pro`, `qwen-image-2.0-pro-2026-03-03`, `qwen-image-2.0`, `qwen-image-2.0-2026-03-03`228 - Free-form `size` in `宽*高` format229 - Total pixels must stay between `512*512` and `2048*2048`230 - Default size is approximately `1024*1024`231 - Best choice for custom ratios such as `21:9` and text-heavy Chinese/English layouts232- `qwen-image-max`, `qwen-image-max-2025-12-30`, `qwen-image-plus`, `qwen-image-plus-2026-01-09`, `qwen-image`233 - Fixed sizes only: `1664*928`, `1472*1104`, `1328*1328`, `1104*1472`, `928*1664`234 - Default size is `1664*928`235 - `qwen-image` currently has the same capability as `qwen-image-plus`236- Legacy DashScope models such as `z-image-turbo`, `z-image-ultra`, `wanx-v1`237 - Keep using them only when the user explicitly asks for legacy behavior or compatibility238239When translating CLI args into DashScope behavior:240241- `--size` wins over `--ar`242- For `qwen-image-2.0*`, prefer explicit `--size`; otherwise infer from `--ar` and use the official recommended resolutions below243- For `qwen-image-max/plus/image`, only use the five official fixed sizes; if the requested ratio is not covered, switch to `qwen-image-2.0-pro`244- `--quality` is a floracat-image-gen compatibility preset, not a native DashScope API field. Mapping `normal` / `2k` onto the `qwen-image-2.0*` table below is an implementation inference, not an official API guarantee245246Recommended `qwen-image-2.0*` sizes for common aspect ratios:247248| Ratio | `normal` | `2k` |249|-------|----------|------|250| `1:1` | `1024*1024` | `1536*1536` |251| `2:3` | `768*1152` | `1024*1536` |252| `3:2` | `1152*768` | `1536*1024` |253| `3:4` | `960*1280` | `1080*1440` |254| `4:3` | `1280*960` | `1440*1080` |255| `9:16` | `720*1280` | `1080*1920` |256| `16:9` | `1280*720` | `1920*1080` |257| `21:9` | `1344*576` | `2048*872` |258259DashScope official APIs also expose `negative_prompt`, `prompt_extend`, and `watermark`, but `floracat-image-gen` does not expose them as dedicated CLI flags today.260261Official references:262263- [Qwen-Image API](https://help.aliyun.com/zh/model-studio/qwen-image-api)264- [Text-to-image guide](https://help.aliyun.com/zh/model-studio/text-to-image)265- [Qwen-Image Edit API](https://help.aliyun.com/zh/model-studio/qwen-image-edit-api)266267### MiniMax Models268269Use `--model image-01` or set `default_model.minimax` / `MINIMAX_IMAGE_MODEL` when the user wants MiniMax image generation.270271Official MiniMax image model options currently documented in the API reference:272273- `image-01` (recommended default)274 - Supports text-to-image and subject-reference image generation275 - Supports official `aspect_ratio` values: `1:1`, `16:9`, `4:3`, `3:2`, `2:3`, `3:4`, `9:16`, `21:9`276 - Supports documented custom `width` / `height` output sizes when using `--size <WxH>`277 - `width` and `height` must both be between `512` and `2048`, and both must be divisible by `8`278- `image-01-live`279 - Lower-latency variant280 - Use `--ar` for sizing; MiniMax documents custom `width` / `height` as only effective for `image-01`281282MiniMax subject reference notes:283284- `--ref` files are sent as MiniMax `subject_reference`285- MiniMax docs currently describe `subject_reference[].type` as `character`286- Official docs say `image_file` supports public URLs or Base64 Data URLs; `floracat-image-gen` sends local refs as Data URLs287- Official docs recommend front-facing portrait references in JPG/JPEG/PNG under 10MB288289Official references:290291- [MiniMax Image Generation Guide](https://platform.minimax.io/docs/guides/image-generation)292- [MiniMax Text-to-Image API](https://platform.minimax.io/docs/api-reference/image-generation-t2i)293- [MiniMax Image-to-Image API](https://platform.minimax.io/docs/api-reference/image-generation-i2i)294295### OpenRouter Models296297Use full OpenRouter model IDs, e.g.:298299- `google/gemini-3.1-flash-image-preview` (recommended, supports image output and reference-image workflows)300- `google/gemini-2.5-flash-image-preview`301- `black-forest-labs/flux.2-pro`302- Other OpenRouter image-capable model IDs303304Notes:305306- OpenRouter image generation uses `/chat/completions`, not the OpenAI `/images` endpoints307- If `--ref` is used, choose a multimodal model that supports image input and image output308- `--imageSize` maps to OpenRouter `imageGenerationOptions.size`; `--size <WxH>` is converted to the nearest OpenRouter size and inferred aspect ratio when possible309310### Replicate Models311312Supported model formats:313314- `owner/name` (recommended for official models), e.g. `google/nano-banana-pro`315- `owner/name:version` (community models by version), e.g. `stability-ai/sdxl:<version>`316317Examples:318319```bash320# Use Replicate default model321${BUN_X} {baseDir}/scripts/main.ts --prompt "A cat" --image out.png --provider replicate322323# Override model explicitly324${BUN_X} {baseDir}/scripts/main.ts --prompt "A cat" --image out.png --provider replicate --model google/nano-banana325```326327## Provider Selection3283291. `--ref` provided + no `--provider` → auto-select Google first, then OpenAI, then Azure, then OpenRouter, then Replicate, then Seedream, then MiniMax (MiniMax subject reference is more specialized toward character/portrait consistency)3302. `--provider` specified → use it (if `--ref`, must be `google`, `openai`, `azure`, `openrouter`, `replicate`, `seedream`, or `minimax`)3313. Only one API key available → use that provider3324. Multiple available → default to Google333334## Quality Presets335336| Preset | Google imageSize | OpenAI Size | OpenRouter size | Replicate resolution | Use Case |337|--------|------------------|-------------|-----------------|----------------------|----------|338| `normal` | 1K | 1024px | 1K | 1K | Quick previews |339| `2k` (default) | 2K | 2048px | 2K | 2K | Covers, illustrations, infographics |340341**Google/OpenRouter imageSize**: Can be overridden with `--imageSize 1K|2K|4K`342343## Aspect Ratios344345Supported: `1:1`, `16:9`, `9:16`, `4:3`, `3:4`, `2.35:1`346347- Google multimodal: uses `imageConfig.aspectRatio`348- OpenAI: maps to closest supported size349- OpenRouter: sends `imageGenerationOptions.aspect_ratio`; if only `--size <WxH>` is given, aspect ratio is inferred automatically350- Replicate: passes `aspect_ratio` to model; when `--ref` is provided without `--ar`, defaults to `match_input_image`351- MiniMax: sends official `aspect_ratio` values directly; if `--size <WxH>` is given without `--ar`, `width` / `height` are sent for `image-01`352353## Generation Mode354355**Default**: Sequential generation.356357**Batch Parallel Generation**: When `--batchfile` contains 2 or more pending tasks, the script automatically enables parallel generation.358359| Mode | When to Use |360|------|-------------|361| Sequential (default) | Normal usage, single images, small batches |362| Parallel batch | Batch mode with 2+ tasks |363364Execution choice:365366| Situation | Preferred approach | Why |367|-----------|--------------------|-----|368| One image, or 1-2 simple images | Sequential | Lower coordination overhead and easier debugging |369| Multiple images already have saved prompt files | Batch (`--batchfile`) | Reuses finalized prompts, applies shared throttling/retries, and gives predictable throughput |370| Each image still needs separate reasoning, prompt writing, or style exploration | Subagents | The work is still exploratory, so each image may need independent analysis before generation |371| Output comes from `floracat-rednote` with `outline.md` + `prompts/` | Batch (`build-batch.ts` -> `--batchfile`) | That workflow already produces prompt files, so direct batch execution is the intended path |372373Rule of thumb:374375- Prefer batch over subagents once prompt files are already saved and the task is "generate all of these"376- Use subagents only when generation is coupled with per-image thinking, rewriting, or divergent creative exploration377378Parallel behavior:379380- Default worker count is automatic, capped by config, built-in default 10381- Provider-specific throttling is applied only in batch mode, and the built-in defaults are tuned for faster throughput while still avoiding obvious RPM bursts382- You can override worker count with `--jobs <count>`383- Each image retries automatically up to 3 attempts384- Final output includes success count, failure count, and per-image failure reasons385386## Error Handling387388- Missing API key → error with setup instructions389- Generation failure → auto-retry up to 3 attempts per image390- Invalid aspect ratio → warning, proceed with default391- Reference images with unsupported provider/model → error with fix hint392393## Extension Support394395Custom configurations via EXTEND.md. See **Preferences** section for paths and supported options.