Creative Claw — Gemini Omni
Read shared execution guidance once per task before using tools. It covers existing authorization, model discovery, optional cost checks, imports, and recovery.
Use video/gemini-omni-flash as Creative Claw's default video model. It is the primary recommendation for fast multimodal video generation and source-video editing with native audio.
Core workflow
- Define one clip: purpose, duration, aspect ratio, subject, action, camera, audio, and protected visual details.
- Search for reusable assets and import any ChatGPT attachments into Creative Claw before passing them to URL fields.
- Prefer a storyboard-first workflow when appearance or continuity matters. Generate and approve a clean full-frame start image with
image/nano-banana-2; generate an end frame when the shot needs a precise destination and the current Omni schema exposes end-frame control. - Call
get_model_params({ model: "video/gemini-omni-flash" })before generation when not already fetched for this task. Treat its current schema as authoritative. - Choose
resolutionfrom the current schema when output size matters. The direct Google route supports360p,720p(default),1080p, and4k; 1080p and 4K are upscaled outputs. Explain the selected duration, ratio, resolution, references, and audio plan when useful. Reuse existing authorization, including an explicitly requested batch; ask only when a material choice is unresolved or the proposed work expands the requested scope. - Call
generate_videowithmodel: "video/gemini-omni-flash". - Let the inline viewer monitor the job. Call
check_jobonly when a later tool needs the completed URL or no viewer is monitoring. - Inspect motion, identity, physics, framing, audio, dialogue, and text artifacts before describing the clip as complete.
Choose the input mode
| Intent | Inputs | Prompt emphasis |
|---|---|---|
| Text-to-video | prompt only |
Describe the complete visible scene and sound. |
| Animate a still | image_url |
Describe what begins moving after the supplied first frame. |
| Reference-guided video | image_urls |
Bind every reference to a role with <IMAGE_REF_N>. |
| Edit a source clip | one item in video_urls |
Give one short change followed by “Keep everything else the same.” |
Do not combine modes casually. Use image_url when an image must be the literal first frame. Use image_urls when images should guide identity, product appearance, wardrobe, environment, or style without becoming the opening frame.
Storyboard-first direction
For ads, branded content, character work, and multi-clip sequences:
- Break the concept into short shots with one main action each.
- Generate each clean start frame separately with Nano Banana 2. Do not pass a labeled grid, contact sheet, panels, captions, or prompt text to the video model.
- Approve identity, wardrobe, product geometry, set design, lighting, composition, and ratio before animation.
- Use the approved image as
image_url. - When the next clip must continue the first, extract the last frame of clip N and use it as the start frame of clip N+1.
This reduces visual drift and makes revisions local to one shot.
First and last frames
Start frame
Pass the approved opening image as image_url. Treat it as frame zero. Describe the motion that follows rather than restating every visible detail.
Good:
The woman turns toward the window as rain begins to trace the glass. Her coat,
face, and the room remain unchanged. Slow dolly-in, one continuous shot.
Weak:
A woman wearing a red coat stands in a room by a window.
The weak version invites the model to reinterpret the already-approved frame.
End frame
Google's underlying Omni model supports first-to-last interpolation, but Creative Claw's exposed fields can change. Pass last_frame_url only when get_model_params returns an end-frame field for the current route. Otherwise use Seedance 2.5 or MiniMax H3 Max for controlled first-to-last generation.
When supported, use two frames with the same ratio, subject identity, and plausible spatial continuity. Describe the transition, not two separate scenes:
Begin exactly from the first frame. In one continuous orbital move, the camera
travels clockwise while the product lid opens and blue light grows from inside.
Arrive exactly at the supplied final frame. No cuts or teleporting objects.
Reference syntax
Pass reference images in image_urls. Cite them with zero-based tokens:
<IMAGE_REF_0> = exact product identity and geometry.
<IMAGE_REF_1> = lighting and material reference only.
<IMAGE_REF_2> = wardrobe reference for the actor.
Then direct the shot:
Preserve the product from <IMAGE_REF_0> exactly. Borrow only the cool rim
lighting and glossy black environment from <IMAGE_REF_1>. The actor wears the
outfit from <IMAGE_REF_2>. She places the product on the pedestal as the camera
makes a slow 30-degree orbit. No cuts, no on-screen text, no geometry changes.
Set agentic_prompting: false whenever the prompt contains these tokens, exact dialogue, timecodes, or carefully authored constraints.
Prompt formula
Order information by what the model must protect:
Purpose: [ad, cinematic insert, product reveal, social clip].
References: [token → exact role; attributes to preserve].
Scene: [subject, environment, composition].
Action: [one filmable action with visible motion].
Camera: [framing + one intentional movement].
Look: [lighting, lens/medium, palette, texture].
Audio: [dialogue, ambience, effects, music, or explicit silence].
Timing: [optional natural beats or time ranges].
Constraints: [identity, product geometry, no cuts/text/subtitles/watermarks].
For reusable B-roll, use one subject, one action, and one camera idea. Say “single continuous shot, no scene cuts” when cuts would make the clip unusable.
Timing and audio
Use the current runtime range discovered by get_model_params; the established Creative Claw route commonly supports 3–10 seconds, 16:9 or 9:16, and 360p, 720p, 1080p, or 4k output resolution.
Natural beats work well:
[0–3s] Slow push toward the unopened bottle.
[3–6s] The cap lifts and cold vapor spills across the table.
[6–8s] Hold on the clean hero angle.
Describe native audio explicitly:
- “No dialogue. Quiet studio ambience and a soft mechanical click.”
- “Dialogue, exact line: ‘Ready when you are.’ Natural room tone, no music.”
- “At five seconds, the percussion enters as the product locks into place.”
- “Generate a silent clip; no music, speech, ambience, or sound effects.”
Keep spoken copy short enough to fit naturally. Quote exact lines and disable prompt rewriting.
Source-video editing
Pass one source clip in video_urls. Use a concise delta:
Replace the overcast sky with a warm sunset. Keep the people, timing, camera
motion, buildings, and every other detail unchanged.
Avoid redescribing the source. A long prompt increases unintended changes. Omni editing is best for one clear transformation per pass.
Prompt examples
Product reveal:
Premium ten-second product film. The matte-black headphones remain identical to
the approved first frame. They rotate slowly above a dark reflective plinth as a
thin ribbon of amber light travels across the ear cups. Macro commercial lens,
slow clockwise orbit, deep black background, crisp highlights. Sound design:
low electronic pulse and a soft magnetic click. Single continuous shot. No
people, text, captions, logo changes, or extra objects.
Character scene with references:
<IMAGE_REF_0> is the exact character identity and face. <IMAGE_REF_1> is the
exact wardrobe. Preserve both. In a rain-soaked train station, she looks over
her shoulder and takes one step toward the arriving train. Medium close-up,
slow handheld push-in, cyan and amber practical lights. Natural rain, distant
train brakes, no dialogue. One shot; no face drift, wardrobe changes, text, or
extra people near camera.
Source edit:
Transform the source video into a premium hand-painted anime look. Preserve the
exact motion, timing, people, composition, and camera path. Keep everything else
the same. No subtitles or added text.
Quality and feedback
Reject static-subject pans when subject motion was requested, identity drift, product deformation, unexpected cuts, lip-sync mismatch, duplicate limbs, embedded text, and audio contradicting the prompt. Revise one failure at a time.
Use submit_feedback when Omni repeatedly violates a concrete instruction, produces a model-specific artifact, exposes an unclear parameter, or lacks a requested control. Include the model ID, mode, attempted task, and the observed quality gap.