openrouter-image2video
Image + motion-prompt → video via OpenRouter. Default model: google/veo-3.1-fast.
Use this when you already have a strong still image and want to bring it to life with subtle motion (camera pan, parallax, gentle animation) while preserving the composition. For motion-from-scratch, use openrouter-text2video instead.
The script accepts either a remote image URL or a local image path; local files are base64-encoded and inlined into the request as a data URL.
Usage
Resolve TOOL_DIR = the directory containing this SKILL.md. Commands below use TOOL_DIR as a symbolic placeholder; replace it with the resolved, quoted path before running Bash.
export OPENROUTER_API_KEY=sk-or-v1-...
# From a local image you already generated with text2image
python3 TOOL_DIR/scripts/generate_video_from_image.py \
--image PROJECT_DIR/assets/teaser.png \
--prompt "slow parallax push-in, soft drift of ambient particles, no camera shake" \
--duration 5 \
--aspect-ratio 16:9 \
--download PROJECT_DIR/assets/teaser.mp4
# Or from a remote URL
python3 TOOL_DIR/scripts/generate_video_from_image.py \
--image-url "https://example.com/still.png" \
--prompt "subtle camera dolly forward, gentle depth-of-field shift" \
--download PROJECT_DIR/assets/scene.mp4
Flags
| Flag |
Default |
Description |
--prompt |
required |
Motion prompt — describe what should move and how |
--download |
required |
Output MP4 path |
--image |
one of --image / --image-url required |
Local image path (PNG/JPG); will be base64-encoded |
--image-url |
one of --image / --image-url required |
Remote image URL |
--model |
google/veo-3.1-fast |
Any OpenRouter image-to-video-capable model |
--duration |
5 |
Seconds |
--aspect-ratio |
16:9 |
16:9, 9:16, 1:1, 4:3, 3:4, 21:9, 9:21 |
--resolution |
720p |
Model-dependent (e.g. 480p, 720p, 1080p) |
--frame-role |
first |
first or last — anchor frame role for the input image |
--generate-audio |
off |
Generate audio with video (if model supports) |
--poll-interval |
5 |
Seconds between polls |
--max-wait |
600 |
Max total wait time |
Flow
POST /api/v1/videos with body:{
"model": "google/veo-3.1-fast",
"prompt": "...motion prompt...",
"aspect_ratio": "16:9",
"duration": 5,
"resolution": "720p",
"frame_images": [
{
"type": "image_url",
"frame_type": "first_frame",
"image_url": {"url": "data:image/png;base64,..." }
}
]
}
The --frame-role first|last flag maps to frame_type: "first_frame"|"last_frame".
GET /api/v1/videos/{id} every 5s until status == "completed"
GET /api/v1/videos/{id}/content → raw MP4 bytes
Notes
- Veo 3.1 Fast is optimized for low-latency image-to-video. Typical render ≈ 60-180s for a 5s 720p clip.
- The motion prompt should describe motion only, not the subject (the subject comes from the image).
- For subjects with prominent faces, keep motion subtle to avoid uncanny artifacts.
- If you also want a defined ending state, supply two images via
frame_images with roles first and last. The current script wires only one anchor frame; extend body["frame_images"] to add a second.
1---2name: openrouter-image2video3description: Animate a still image into a short video via OpenRouter. Default model google/veo-3.1-fast.4---56# openrouter-image2video78Image + motion-prompt → video via OpenRouter. Default model: `google/veo-3.1-fast`.910Use this when you already have a strong still image and want to bring it to life with subtle motion (camera pan, parallax, gentle animation) while preserving the composition. For motion-from-scratch, use `openrouter-text2video` instead.1112The script accepts either a remote image URL or a local image path; local files are base64-encoded and inlined into the request as a data URL.1314## Usage1516Resolve `TOOL_DIR` = the directory containing this `SKILL.md`. Commands below use `TOOL_DIR` as a symbolic placeholder; replace it with the resolved, quoted path before running Bash.1718```bash19export OPENROUTER_API_KEY=sk-or-v1-...2021# From a local image you already generated with text2image22python3 TOOL_DIR/scripts/generate_video_from_image.py \23 --image PROJECT_DIR/assets/teaser.png \24 --prompt "slow parallax push-in, soft drift of ambient particles, no camera shake" \25 --duration 5 \26 --aspect-ratio 16:9 \27 --download PROJECT_DIR/assets/teaser.mp42829# Or from a remote URL30python3 TOOL_DIR/scripts/generate_video_from_image.py \31 --image-url "https://example.com/still.png" \32 --prompt "subtle camera dolly forward, gentle depth-of-field shift" \33 --download PROJECT_DIR/assets/scene.mp434```3536## Flags3738| Flag | Default | Description |39|---|---|---|40| `--prompt` | required | Motion prompt — describe what should move and how |41| `--download` | required | Output MP4 path |42| `--image` | one of `--image` / `--image-url` required | Local image path (PNG/JPG); will be base64-encoded |43| `--image-url` | one of `--image` / `--image-url` required | Remote image URL |44| `--model` | `google/veo-3.1-fast` | Any OpenRouter image-to-video-capable model |45| `--duration` | `5` | Seconds |46| `--aspect-ratio` | `16:9` | `16:9`, `9:16`, `1:1`, `4:3`, `3:4`, `21:9`, `9:21` |47| `--resolution` | `720p` | Model-dependent (e.g. `480p`, `720p`, `1080p`) |48| `--frame-role` | `first` | `first` or `last` — anchor frame role for the input image |49| `--generate-audio` | off | Generate audio with video (if model supports) |50| `--poll-interval` | `5` | Seconds between polls |51| `--max-wait` | `600` | Max total wait time |5253## Flow54551. `POST /api/v1/videos` with body:56 ```json57 {58 "model": "google/veo-3.1-fast",59 "prompt": "...motion prompt...",60 "aspect_ratio": "16:9",61 "duration": 5,62 "resolution": "720p",63 "frame_images": [64 {65 "type": "image_url",66 "frame_type": "first_frame",67 "image_url": {"url": "data:image/png;base64,..." }68 }69 ]70 }71 ```72 The `--frame-role first|last` flag maps to `frame_type: "first_frame"|"last_frame"`.732. `GET /api/v1/videos/{id}` every 5s until `status == "completed"`743. `GET /api/v1/videos/{id}/content` → raw MP4 bytes7576## Notes7778- Veo 3.1 Fast is optimized for low-latency image-to-video. Typical render ≈ 60-180s for a 5s 720p clip.79- The motion prompt should describe motion only, not the subject (the subject comes from the image).80- For subjects with prominent faces, keep motion subtle to avoid uncanny artifacts.81- If you also want a defined ending state, supply two images via `frame_images` with roles `first` and `last`. The current script wires only one anchor frame; extend `body["frame_images"]` to add a second.