Renoise Model Routing
Select by task, not by habit. Then prompt for the selected model instead of sending every model the same generic brief.
Hard Boundary
Live model capabilities are authoritative for availability, defaults, input roles, combinations, limits, duration, ratio, resolution, and audio controls. This Skill contains only researched routing preferences and prompting styles.
- Preserve a model explicitly named by the user.
- Filter candidates by the live capabilities required by the task.
- Choose the best specialist below.
- Use the live
isDefault model only when no specialist clearly fits or candidates tie.
- Inspect the selected model's live guidance before writing the final prompt.
- Let the selected model's profile override generic prompt-craft advice when prompt structure, density, or reference wording conflicts.
- Never pass a role or parameter merely because this guide mentions a public model capability; the Renoise deployment may expose a narrower contract.
Live availability alone is not a routing recommendation. Omit models that are fully dominated or lack a clear task advantage; an explicitly named model still follows rule 1.
Routing Questions
Classify the request before selecting:
- Kind: image, video, or audio.
- Operation: create, edit, continue, interpolate, animate a reference, or upscale.
- Priority: final quality, strict instruction following, aesthetics, identity consistency, speed, or cost.
- Structure: exact text/layout, one cinematic shot, multi-shot narrative, dialogue performance, music, or a complete audio scene.
- References: none, one source image, several identity/style references, source video, endpoint frames, or voice/audio references.
Do not call a model “best” without naming the task it is best for.
Image Routing
Decision Order
| Task |
Prefer |
Why |
| Exact text, UI, diagrams, ads, packaging, compositing, identity-sensitive edits, or peak final quality |
gpt-image-2 |
Strong instruction following, photorealism, typography/layout preservation, localization, and complex editing; favor it when fewer retries matter more than latency or cost. |
| General high-quality photorealistic or commercial image |
live image default, currently seedream-5-0-pro |
Strong realism, dense information design, multilingual typography, multi-image compositing, and controlled local edits at the deployment's normal quality/cost balance. |
| High-throughput variants, prototyping, interactive generation, or cost-sensitive production |
nano-banana-2-lite |
Low-latency scale tier for clear, shallow workflows; use another tier when the task depends on many references or sequential editing. |
| Balanced Google workflow, extreme aspect ratios, several references, multilingual localization, or conversational iteration |
nano-banana-2 |
Google-family workhorse balancing quality, latency, text, reference reasoning, and broad formats. |
| Maximum Google-family world knowledge, brand consistency, localization, or reasoning-heavy composition |
nano-banana-pro |
Specialist for intricate professional assets and precise spatial relationships. |
| Best aesthetic exploration, stylized art direction, editorial mood, or concept art |
mj-v8.2 |
Current Midjourney default and strongest current aesthetic prior. |
| Seedream output above 2K or more than ten image references |
seedream-5-0-lite |
Its 3K–4K output and larger reference allowance are the remaining clear reasons to prefer it over Pro. |
| xAI image request |
grok-image for lower-cost iteration; grok-image-quality when higher quality justifies the added cost |
Both support direct natural-language generation and editing through the live Renoise contract. |
Image Prompting Styles
GPT Image 2 — production brief
For complex production briefs, use labeled, ordered instructions:
USE: [ad / packaging / UI / infographic / edit]
SCENE/BACKGROUND: ...
SUBJECT: ...
COMPOSITION: ...
LIGHTING/MATERIALS: ...
EXACT TEXT: "..." with placement, hierarchy, font character, and color
PRESERVE: identity, geometry, layout, logo, and all unmentioned elements
CHANGE ONLY: ...
AVOID: ...
- Put literal in-image text in quotes and specify hierarchy and placement.
- For edits, say change only X; preserve everything else.
- Identify every reference by its job: source, identity, product, layout, or style.
- Prefer one-change iterative edits over rewriting the whole brief.
Nano Banana family — conversational creative direction
[Subject] + [action] + [location] + [composition] + [style]
Camera/viewpoint: ...
Lighting/materials: ...
Exact text: "..."
Keep unchanged: ...
- Start with a clear operation verb and use positive, concrete natural language.
- For references, state each source's job, the relationship among them, and the new scenario.
- For edits, separate the requested change from what remains unchanged; repeat preservation requirements on later turns.
- Quote exact visible copy and specify its hierarchy, placement, visual character, and localization target.
- Nano Banana 2: relate multiple references explicitly and iterate conversationally.
- Lite: keep one clear composition and few dependencies rather than relying on long multi-turn edits.
- Pro: state brand invariants, localization requirements, and complex spatial relationships precisely.
Seedream 5 — design brief
Put the deliverable format first, then subject, layout, light, exact text, and invariants:
[DELIVERABLE FORMAT and ratio/use case]
Subject and setting: ...
Layout and spatial relationships: ...
Lighting, materials, and palette: ...
Exact visible copy: "..."
Keep exactly: ...
- Pro: specify dense layout, multilingual copy, materials, local edit regions, and everything the edit preserves.
- Lite: natural language works well; make causal/spatial relationships explicit for complex transformations.
- Assign each reference a clear role and avoid vague pronouns in multi-reference edits.
Midjourney — aesthetic visual phrase
Use a concise visual description, not a requirements document:
[subject], [medium], [environment], [lighting], [palette], [mood], [composition]
- Prefer concrete visual nouns and adjectives over keyword spam.
- Name the lighting and composition that matter instead of relying on “cinematic.”
- Use Midjourney for look development rather than text-heavy layouts or surgical edits.
Grok Imagine Image — concise scene direction
[subject doing action] in [setting], [camera/framing], [lighting], [material/style]
- Use standard for lower-cost iteration and Quality when the higher-quality tier is worth the added cost.
- For edits, attach the source image or images and describe the requested change directly.
- Keep the prompt concrete enough that the main subject and action remain dominant.
Video Routing
Decision Order
| Task |
Prefer |
Why |
| Explicit Seedance 2.5 request, continuous 16–30 second sequence, rich mixed references, source-video edit, or forward/backward extension |
seedance-2.5-byteplus |
Long-form and reference-heavy specialist with precise edit/extension workflows. |
| General multimodal video, complex physical motion, recurring references, product/character continuity, or image-to-video with synchronized audio |
seedance-2.0-byteplus |
Live generalist with strong image/video/audio role assignment, motion, continuity, and native sound. |
| Fast Seedance draft |
seedance-2.0-fast-byteplus |
Official speed/cost balance tier. |
| Lowest-cost/high-volume Seedance draft |
seedance-2.0-mini-byteplus |
Cost-performance tier for iteration and selection. |
| Short 720p generation, text rendering, or source-video edit within its narrow live contract |
gemini-omni-flash |
Strong short-form instruction following, multi-shot generation, and conversational editing. |
| Exact first frame, last frame, frame interpolation, focused video references, 2K delivery, or reference audio paired with visual references |
hailuo-h3 |
Strong endpoint control and structured multimodal audiovisual direction. |
| Fast H3-family text-to-video or first/last-frame generation where 768p is enough but quality matters more than minimum cost |
hailuo-h3-max |
Quality/speed middle tier between full H3 and Turbo; current arena results remain strong. |
| Fastest, lowest-cost H3-family text-to-video or first/last-frame generation |
h3-max-turbo |
Rapid low-cost iteration when the live contract needs no generic image, video, or audio references; use full H3 instead when references or 2K are required. |
| xAI one-image animation with native sound |
grok-video-1.5 |
Current Renoise 1.5 contract requires exactly one input image. |
| xAI text-to-video or video needing several image references |
grok-video |
The base contract retains text-only and multi-image modes that 1.5 does not expose. |
| Upscale an existing video without changing its content |
upscale-video-topaz-starlight-2.5 |
Dedicated restoration/upscaling model; use a generation or editing model when objects, motion, timing, or style should change. |
Tier and version notes
- Seedance Full, Fast, and Mini are respectively the quality, speed/cost-balance, and cost-performance tiers; live estimates decide the actual trade-off.
- Within the H3 family, use H3 for 2K or multimodal references, H3 Max for faster official-channel first/last-frame work, and H3 Max Turbo for the fastest and cheapest text/first/last-frame iteration. Never infer reference-to-video support for Max or Turbo when the live roles do not expose it.
- Grok's public upstream capabilities may move faster than Renoise's contract. Route from live roles and limits rather than assuming an upstream feature is connected.
- Model preference leaderboards are task-, resolution-, and audio-filter-specific; use them as volatile evidence, not a single aggregate ranking.
Video Prompting Styles
Seedance family — multimodal director brief
Assign each reference an explicit job, then write shots:
References:
- @material:<ID>: subject identity
- @material:<ID>: location/style
- @material:<ID>: motion/camera/voice
Shot 1: framing, subject action, camera move, environment, sound/dialogue.
Shot 2: ...
Continuity: traits and objects that must remain unchanged.
Constraints: no subtitles/logo/watermark unless requested.
- Use a few stable identity features rather than an exhaustive biography.
- Give each beat a clear camera behavior; avoid simultaneous conflicting moves.
- Describe body part, speed, force, and physical consequence for actions.
- Externalize emotion through visible behavior and direct dialogue/audio explicitly.
- Use the same prompt structure for Fast/Mini while keeping draft requests easy to compare.
- For Seedance 2.5, begin with the intended result, map every reference to its purpose, then use non-overlapping timestamp ranges or numbered shots with continuity notes.
- For a source-video edit, name the source, the intended change, its time range when relevant, and what remains unchanged.
- For extension, state forward or backward direction and describe the visual, motion, and audio continuity across the source boundary.
Gemini Omni Flash — conversational generation/editing
- For creation: write a direct natural-language brief with subject, action, camera, setting, visible text, and desired audio.
- Gemini readily creates multi-shot clips; say “single continuous shot,” “single unbroken scene,” or “no scene cuts” when continuity is required.
- Use natural time phrases or
[0–3s] blocks for important beats.
- For editing: identify the source, say change only X, and list what remains unchanged.
- Make one edit per turn when possible; use follow-up refinement rather than replacing the entire prompt.
MiniMax H3 family — mode-specific audiovisual plan
Choose one live mode and prompt accordingly. H3 Max and H3 Max Turbo follow the same first/last-frame prompting pattern when those modes are live, but generic reference mode belongs only to models whose live roles explicitly expose it:
Text-to-video: state the subject, action over time, camera movement, setting, lighting, and sound intent directly.
First frame: describe only the motion, camera path, action development, and sound after the supplied opening state.
Last frame: describe the plausible path that converges on the supplied ending.
First + last: describe the continuous transition between states; avoid re-describing the two stills. Prefer one coherent shot unless a cut is essential.
Reference mode: assign each image, video, and audio an explicit identity, motion, style, voice, action, or sound job.
For complex reference work use:
References: each @material:<ID> image/video/audio and its job.
Task summary: target result and what each source contributes.
Preserve: identity, motion, voice, style, or content that carries over.
[Shot 1] Visual action, camera motion, dialogue, and diegetic sound.
[Shot 2 at time] Cut and continuation.
Overall soundscape: ambience, Foley, and non-verbal sound.
Non-diegetic music: instrumentation, tempo, dynamics — or none.
Specify camera type + amplitude + speed only when they matter. Keep dialogue exact, label speakers consistently, and separate in-scene sound from background score.
Grok Imagine Video — animate the change
For image-to-video, do not re-describe the static source:
[subject motion/change], [one camera move], [atmosphere/lighting change].
Sound: specific dialogue/SFX/ambience; no music if unwanted.
- Use one primary action or a short causal sequence and front-load the important motion.
- For text-to-video, include the subject and setting; for image-to-video, describe the change from the starting frame.
- Assign every live reference a clear visual or motion job.
- Request dialogue, effects, ambience, or music explicitly when native audio matters.
Topaz Starlight — no creative prompt
Upscaling is not regeneration. Supply only the source video and settings required by the live capability; do not add a creative prompt or unrelated media inputs. If the user wants content changed, route to a video editing model instead.
Audio Routing
| Task |
Prefer |
Why |
| Short song excerpt, instrumental, loop, preview, score cue, or music bed |
lyria-clip |
Dedicated 30-second music model for rapid iteration, social assets, and background cues. |
| Dialogue, expressive narration, podcast, dubbing/re-voicing, SFX/ambience, or a complete sound scene |
seed-audio-1.0 |
Unified speech, dialogue, effects, and ambience generation; live default audio model. |
Neither broadly replaces the other: Lyria composes short music clips; Seed Audio handles speech-led and complete sound scenes.
Lyria 3 Clip — composer brief
Genre/style: ...
Mood: ...
Instrumentation: ...
Tempo/rhythm: ... BPM
Vocals: instrumental OR voice type, delivery, and language
Lyrics/theme: exact lyrics or subject
Structure over ~30 seconds: intro → development → ending
Production: mix character, era, space, dynamics
- Explicitly say instrumental when vocals are unwanted.
- Name instruments, tempo, vocal range/texture, backing vocals, and language.
- Keep custom lyrics concise for 30 seconds;
Lyrics:, [Verse], and [Chorus] can clarify structure.
- Use timestamps only for meaningful structural changes; musical alignment follows bars rather than sample-accurate timing.
- Describe genre, era, instrumentation, and production traits instead of requesting a named artist imitation.
- Use an optional guide image only when the live material roles expose it.
Seed Audio 1.0 — sound-scene script
Setting/acoustics: ...
Continuous sound bed: ...
Speaker A (age, accent, timbre, emotion, pace): "Exact line."
Action-tied SFX: ...
Speaker B (...): "Exact line."
Music cue and ending: ...
- For full-scene work, direct it as a scene; for speech-only work, specify voice identity, performance, pacing, and acoustics.
- Label speakers and exact lines; match prompt language to dialogue language.
- Use clean authorized voice references, or an authorized character image for inferred voice, according to the mutually exclusive live reference modes.
- Use concrete Foley and ambience; onomatopoeia can clarify transient sounds.
- Use precise timing for dialogue; describe SFX, ambience, and music timing approximately unless live guidance says otherwise.
Research Basis
Reviewed 2026-09-10. Rankings are directional and decay quickly; live capabilities and task-specific tests override them.
Primary guidance:
- OpenAI GPT Image: https://developers.openai.com/api/docs/models/gpt-image-2, https://developers.openai.com/api/docs/guides/image-generation, and https://developers.openai.com/cookbook/examples/multimodal/image-gen-models-prompting-guide
- Google Gemini image: https://ai.google.dev/gemini-api/docs/image-generation and https://cloud.google.com/blog/products/ai-machine-learning/ultimate-prompting-guide-for-nano-banana
- Midjourney prompting/version docs: https://docs.midjourney.com/hc/en-us/articles/32023408776205-Prompt-Basics, https://docs.midjourney.com/hc/en-us/articles/32199405667853-Version, and https://updates.midjourney.com/version-8-2/
- ByteDance Seedream: https://seed.bytedance.com/en/blog/deeper-thinking-more-accurate-generation-introducing-seedream-5-0-lite and https://seed.bytedance.com/en/blog/beyond-generation-it-understands-design-introducing-seedream-5-0-pro
- xAI Imagine image/video: https://docs.x.ai/developers/model-capabilities/imagine and https://docs.x.ai/developers/model-capabilities/video/generation
- BytePlus Seedance 2.5 and 2.0 prompt guides: https://docs.byteplus.com/en/docs/ModelArk/2607689 and https://docs.byteplus.com/en/docs/ModelArk/2222480
- MiniMax H3 family: https://www.minimax.io/blog/minimax-h3, https://platform.minimax.io/docs/guides/video-generation, https://huggingface.co/MiniMaxAI/MiniMax-H3/tree/main/docs, and https://blog.fal.ai/introducing-h3-max-by-fal/
- Topaz Starlight Precise 2.5: https://developer.topazlabs.com/video-models/starlight/starlight-precise-2.5
- Gemini Omni: https://ai.google.dev/gemini-api/docs/omni
- Lyria: https://ai.google.dev/gemini-api/docs/music-generation, https://deepmind.google/models/lyria/prompt-guide/, and https://cloud.google.com/blog/products/ai-machine-learning/ultimate-prompting-guide-for-lyria-3-pro
- Seed Audio: https://seed.bytedance.com/en/blog/from-speech-to-audio-creation-introducing-the-seed-audio-1-0-audio-creation-model and https://docs.byteplus.com/en/docs/byteplusvoice/seedaudio-01
- Independent preference evidence: https://artificialanalysis.ai/image/leaderboard/text-to-image, https://artificialanalysis.ai/image/leaderboard/editing, https://artificialanalysis.ai/video/leaderboard/text-to-video, https://artificialanalysis.ai/video/leaderboard/image-to-video, and https://artificialanalysis.ai/video/leaderboard/video-editing
1---2name: model-routing3description: Choose the best Renoise image, video, or audio model and write prompts for it. Always use before any Renoise generation estimate or task, and when the user asks for the best model, model comparison, model routing, which model to use, or “用哪个模型 / 模型选择”. Use alongside the active generation workflow; this Skill does not execute tasks.4---56# Renoise Model Routing78Select by task, not by habit. Then prompt for the selected model instead of sending every model the same generic brief.910## Hard Boundary1112Live model capabilities are authoritative for availability, defaults, input roles, combinations, limits, duration, ratio, resolution, and audio controls. This Skill contains only researched routing preferences and prompting styles.13141. Preserve a model explicitly named by the user.152. Filter candidates by the live capabilities required by the task.163. Choose the best specialist below.174. Use the live `isDefault` model only when no specialist clearly fits or candidates tie.185. Inspect the selected model's live guidance before writing the final prompt.196. Let the selected model's profile override generic prompt-craft advice when prompt structure, density, or reference wording conflicts.207. Never pass a role or parameter merely because this guide mentions a public model capability; the Renoise deployment may expose a narrower contract.2122Live availability alone is not a routing recommendation. Omit models that are fully dominated or lack a clear task advantage; an explicitly named model still follows rule 1.2324## Routing Questions2526Classify the request before selecting:2728- **Kind:** image, video, or audio.29- **Operation:** create, edit, continue, interpolate, animate a reference, or upscale.30- **Priority:** final quality, strict instruction following, aesthetics, identity consistency, speed, or cost.31- **Structure:** exact text/layout, one cinematic shot, multi-shot narrative, dialogue performance, music, or a complete audio scene.32- **References:** none, one source image, several identity/style references, source video, endpoint frames, or voice/audio references.3334Do not call a model “best” without naming the task it is best for.3536# Image Routing3738## Decision Order3940| Task | Prefer | Why |41|---|---|---|42| Exact text, UI, diagrams, ads, packaging, compositing, identity-sensitive edits, or peak final quality | `gpt-image-2` | Strong instruction following, photorealism, typography/layout preservation, localization, and complex editing; favor it when fewer retries matter more than latency or cost. |43| General high-quality photorealistic or commercial image | live image default, currently `seedream-5-0-pro` | Strong realism, dense information design, multilingual typography, multi-image compositing, and controlled local edits at the deployment's normal quality/cost balance. |44| High-throughput variants, prototyping, interactive generation, or cost-sensitive production | `nano-banana-2-lite` | Low-latency scale tier for clear, shallow workflows; use another tier when the task depends on many references or sequential editing. |45| Balanced Google workflow, extreme aspect ratios, several references, multilingual localization, or conversational iteration | `nano-banana-2` | Google-family workhorse balancing quality, latency, text, reference reasoning, and broad formats. |46| Maximum Google-family world knowledge, brand consistency, localization, or reasoning-heavy composition | `nano-banana-pro` | Specialist for intricate professional assets and precise spatial relationships. |47| Best aesthetic exploration, stylized art direction, editorial mood, or concept art | `mj-v8.2` | Current Midjourney default and strongest current aesthetic prior. |48| Seedream output above 2K or more than ten image references | `seedream-5-0-lite` | Its 3K–4K output and larger reference allowance are the remaining clear reasons to prefer it over Pro. |49| xAI image request | `grok-image` for lower-cost iteration; `grok-image-quality` when higher quality justifies the added cost | Both support direct natural-language generation and editing through the live Renoise contract. |5051## Image Prompting Styles5253### GPT Image 2 — production brief5455For complex production briefs, use labeled, ordered instructions:5657```text58USE: [ad / packaging / UI / infographic / edit]59SCENE/BACKGROUND: ...60SUBJECT: ...61COMPOSITION: ...62LIGHTING/MATERIALS: ...63EXACT TEXT: "..." with placement, hierarchy, font character, and color64PRESERVE: identity, geometry, layout, logo, and all unmentioned elements65CHANGE ONLY: ...66AVOID: ...67```6869- Put literal in-image text in quotes and specify hierarchy and placement.70- For edits, say **change only X; preserve everything else**.71- Identify every reference by its job: source, identity, product, layout, or style.72- Prefer one-change iterative edits over rewriting the whole brief.7374### Nano Banana family — conversational creative direction7576```text77[Subject] + [action] + [location] + [composition] + [style]78Camera/viewpoint: ...79Lighting/materials: ...80Exact text: "..."81Keep unchanged: ...82```8384- Start with a clear operation verb and use positive, concrete natural language.85- For references, state each source's job, the relationship among them, and the new scenario.86- For edits, separate the requested change from what remains unchanged; repeat preservation requirements on later turns.87- Quote exact visible copy and specify its hierarchy, placement, visual character, and localization target.88- Nano Banana 2: relate multiple references explicitly and iterate conversationally.89- Lite: keep one clear composition and few dependencies rather than relying on long multi-turn edits.90- Pro: state brand invariants, localization requirements, and complex spatial relationships precisely.9192### Seedream 5 — design brief9394Put the deliverable format first, then subject, layout, light, exact text, and invariants:9596```text97[DELIVERABLE FORMAT and ratio/use case]98Subject and setting: ...99Layout and spatial relationships: ...100Lighting, materials, and palette: ...101Exact visible copy: "..."102Keep exactly: ...103```104105- Pro: specify dense layout, multilingual copy, materials, local edit regions, and everything the edit preserves.106- Lite: natural language works well; make causal/spatial relationships explicit for complex transformations.107- Assign each reference a clear role and avoid vague pronouns in multi-reference edits.108109### Midjourney — aesthetic visual phrase110111Use a concise visual description, not a requirements document:112113```text114[subject], [medium], [environment], [lighting], [palette], [mood], [composition]115```116117- Prefer concrete visual nouns and adjectives over keyword spam.118- Name the lighting and composition that matter instead of relying on “cinematic.”119- Use Midjourney for look development rather than text-heavy layouts or surgical edits.120121### Grok Imagine Image — concise scene direction122123```text124[subject doing action] in [setting], [camera/framing], [lighting], [material/style]125```126127- Use standard for lower-cost iteration and Quality when the higher-quality tier is worth the added cost.128- For edits, attach the source image or images and describe the requested change directly.129- Keep the prompt concrete enough that the main subject and action remain dominant.130131# Video Routing132133## Decision Order134135| Task | Prefer | Why |136|---|---|---|137| Explicit Seedance 2.5 request, continuous 16–30 second sequence, rich mixed references, source-video edit, or forward/backward extension | `seedance-2.5-byteplus` | Long-form and reference-heavy specialist with precise edit/extension workflows. |138| General multimodal video, complex physical motion, recurring references, product/character continuity, or image-to-video with synchronized audio | `seedance-2.0-byteplus` | Live generalist with strong image/video/audio role assignment, motion, continuity, and native sound. |139| Fast Seedance draft | `seedance-2.0-fast-byteplus` | Official speed/cost balance tier. |140| Lowest-cost/high-volume Seedance draft | `seedance-2.0-mini-byteplus` | Cost-performance tier for iteration and selection. |141| Short 720p generation, text rendering, or source-video edit within its narrow live contract | `gemini-omni-flash` | Strong short-form instruction following, multi-shot generation, and conversational editing. |142| Exact first frame, last frame, frame interpolation, focused video references, 2K delivery, or reference audio paired with visual references | `hailuo-h3` | Strong endpoint control and structured multimodal audiovisual direction. |143| Fast H3-family text-to-video or first/last-frame generation where 768p is enough but quality matters more than minimum cost | `hailuo-h3-max` | Quality/speed middle tier between full H3 and Turbo; current arena results remain strong. |144| Fastest, lowest-cost H3-family text-to-video or first/last-frame generation | `h3-max-turbo` | Rapid low-cost iteration when the live contract needs no generic image, video, or audio references; use full H3 instead when references or 2K are required. |145| xAI one-image animation with native sound | `grok-video-1.5` | Current Renoise 1.5 contract requires exactly one input image. |146| xAI text-to-video or video needing several image references | `grok-video` | The base contract retains text-only and multi-image modes that 1.5 does not expose. |147| Upscale an existing video without changing its content | `upscale-video-topaz-starlight-2.5` | Dedicated restoration/upscaling model; use a generation or editing model when objects, motion, timing, or style should change. |148149### Tier and version notes150151- Seedance Full, Fast, and Mini are respectively the quality, speed/cost-balance, and cost-performance tiers; live estimates decide the actual trade-off.152- Within the H3 family, use H3 for 2K or multimodal references, H3 Max for faster official-channel first/last-frame work, and H3 Max Turbo for the fastest and cheapest text/first/last-frame iteration. Never infer reference-to-video support for Max or Turbo when the live roles do not expose it.153- Grok's public upstream capabilities may move faster than Renoise's contract. Route from live roles and limits rather than assuming an upstream feature is connected.154- Model preference leaderboards are task-, resolution-, and audio-filter-specific; use them as volatile evidence, not a single aggregate ranking.155156## Video Prompting Styles157158### Seedance family — multimodal director brief159160Assign each reference an explicit job, then write shots:161162```text163References:164- @material:<ID>: subject identity165- @material:<ID>: location/style166- @material:<ID>: motion/camera/voice167168Shot 1: framing, subject action, camera move, environment, sound/dialogue.169Shot 2: ...170Continuity: traits and objects that must remain unchanged.171Constraints: no subtitles/logo/watermark unless requested.172```173174- Use a few stable identity features rather than an exhaustive biography.175- Give each beat a clear camera behavior; avoid simultaneous conflicting moves.176- Describe body part, speed, force, and physical consequence for actions.177- Externalize emotion through visible behavior and direct dialogue/audio explicitly.178- Use the same prompt structure for Fast/Mini while keeping draft requests easy to compare.179- For Seedance 2.5, begin with the intended result, map every reference to its purpose, then use non-overlapping timestamp ranges or numbered shots with continuity notes.180- For a source-video edit, name the source, the intended change, its time range when relevant, and what remains unchanged.181- For extension, state forward or backward direction and describe the visual, motion, and audio continuity across the source boundary.182183### Gemini Omni Flash — conversational generation/editing184185- For creation: write a direct natural-language brief with subject, action, camera, setting, visible text, and desired audio.186- Gemini readily creates multi-shot clips; say “single continuous shot,” “single unbroken scene,” or “no scene cuts” when continuity is required.187- Use natural time phrases or `[0–3s]` blocks for important beats.188- For editing: identify the source, say **change only X**, and list what remains unchanged.189- Make one edit per turn when possible; use follow-up refinement rather than replacing the entire prompt.190191### MiniMax H3 family — mode-specific audiovisual plan192193Choose one live mode and prompt accordingly. H3 Max and H3 Max Turbo follow the same first/last-frame prompting pattern when those modes are live, but generic reference mode belongs only to models whose live roles explicitly expose it:194195- **Text-to-video:** state the subject, action over time, camera movement, setting, lighting, and sound intent directly.196197- **First frame:** describe only the motion, camera path, action development, and sound after the supplied opening state.198- **Last frame:** describe the plausible path that converges on the supplied ending.199- **First + last:** describe the continuous transition between states; avoid re-describing the two stills. Prefer one coherent shot unless a cut is essential.200- **Reference mode:** assign each image, video, and audio an explicit identity, motion, style, voice, action, or sound job.201202For complex reference work use:203204```text205References: each @material:<ID> image/video/audio and its job.206Task summary: target result and what each source contributes.207Preserve: identity, motion, voice, style, or content that carries over.208[Shot 1] Visual action, camera motion, dialogue, and diegetic sound.209[Shot 2 at time] Cut and continuation.210Overall soundscape: ambience, Foley, and non-verbal sound.211Non-diegetic music: instrumentation, tempo, dynamics — or none.212```213214Specify camera **type + amplitude + speed** only when they matter. Keep dialogue exact, label speakers consistently, and separate in-scene sound from background score.215216### Grok Imagine Video — animate the change217218For image-to-video, do not re-describe the static source:219220```text221[subject motion/change], [one camera move], [atmosphere/lighting change].222Sound: specific dialogue/SFX/ambience; no music if unwanted.223```224225- Use one primary action or a short causal sequence and front-load the important motion.226- For text-to-video, include the subject and setting; for image-to-video, describe the change from the starting frame.227- Assign every live reference a clear visual or motion job.228- Request dialogue, effects, ambience, or music explicitly when native audio matters.229230### Topaz Starlight — no creative prompt231232Upscaling is not regeneration. Supply only the source video and settings required by the live capability; do not add a creative prompt or unrelated media inputs. If the user wants content changed, route to a video editing model instead.233234# Audio Routing235236| Task | Prefer | Why |237|---|---|---|238| Short song excerpt, instrumental, loop, preview, score cue, or music bed | `lyria-clip` | Dedicated 30-second music model for rapid iteration, social assets, and background cues. |239| Dialogue, expressive narration, podcast, dubbing/re-voicing, SFX/ambience, or a complete sound scene | `seed-audio-1.0` | Unified speech, dialogue, effects, and ambience generation; live default audio model. |240241Neither broadly replaces the other: Lyria composes short music clips; Seed Audio handles speech-led and complete sound scenes.242243## Lyria 3 Clip — composer brief244245```text246Genre/style: ...247Mood: ...248Instrumentation: ...249Tempo/rhythm: ... BPM250Vocals: instrumental OR voice type, delivery, and language251Lyrics/theme: exact lyrics or subject252Structure over ~30 seconds: intro → development → ending253Production: mix character, era, space, dynamics254```255256- Explicitly say **instrumental** when vocals are unwanted.257- Name instruments, tempo, vocal range/texture, backing vocals, and language.258- Keep custom lyrics concise for 30 seconds; `Lyrics:`, `[Verse]`, and `[Chorus]` can clarify structure.259- Use timestamps only for meaningful structural changes; musical alignment follows bars rather than sample-accurate timing.260- Describe genre, era, instrumentation, and production traits instead of requesting a named artist imitation.261- Use an optional guide image only when the live material roles expose it.262263## Seed Audio 1.0 — sound-scene script264265```text266Setting/acoustics: ...267Continuous sound bed: ...268Speaker A (age, accent, timbre, emotion, pace): "Exact line."269Action-tied SFX: ...270Speaker B (...): "Exact line."271Music cue and ending: ...272```273274- For full-scene work, direct it as a scene; for speech-only work, specify voice identity, performance, pacing, and acoustics.275- Label speakers and exact lines; match prompt language to dialogue language.276- Use clean authorized voice references, or an authorized character image for inferred voice, according to the mutually exclusive live reference modes.277- Use concrete Foley and ambience; onomatopoeia can clarify transient sounds.278- Use precise timing for dialogue; describe SFX, ambience, and music timing approximately unless live guidance says otherwise.279280# Research Basis281282Reviewed 2026-09-10. Rankings are directional and decay quickly; live capabilities and task-specific tests override them.283284Primary guidance:285286- OpenAI GPT Image: https://developers.openai.com/api/docs/models/gpt-image-2, https://developers.openai.com/api/docs/guides/image-generation, and https://developers.openai.com/cookbook/examples/multimodal/image-gen-models-prompting-guide287- Google Gemini image: https://ai.google.dev/gemini-api/docs/image-generation and https://cloud.google.com/blog/products/ai-machine-learning/ultimate-prompting-guide-for-nano-banana288- Midjourney prompting/version docs: https://docs.midjourney.com/hc/en-us/articles/32023408776205-Prompt-Basics, https://docs.midjourney.com/hc/en-us/articles/32199405667853-Version, and https://updates.midjourney.com/version-8-2/289- ByteDance Seedream: https://seed.bytedance.com/en/blog/deeper-thinking-more-accurate-generation-introducing-seedream-5-0-lite and https://seed.bytedance.com/en/blog/beyond-generation-it-understands-design-introducing-seedream-5-0-pro290- xAI Imagine image/video: https://docs.x.ai/developers/model-capabilities/imagine and https://docs.x.ai/developers/model-capabilities/video/generation291- BytePlus Seedance 2.5 and 2.0 prompt guides: https://docs.byteplus.com/en/docs/ModelArk/2607689 and https://docs.byteplus.com/en/docs/ModelArk/2222480292- MiniMax H3 family: https://www.minimax.io/blog/minimax-h3, https://platform.minimax.io/docs/guides/video-generation, https://huggingface.co/MiniMaxAI/MiniMax-H3/tree/main/docs, and https://blog.fal.ai/introducing-h3-max-by-fal/293- Topaz Starlight Precise 2.5: https://developer.topazlabs.com/video-models/starlight/starlight-precise-2.5294- Gemini Omni: https://ai.google.dev/gemini-api/docs/omni295- Lyria: https://ai.google.dev/gemini-api/docs/music-generation, https://deepmind.google/models/lyria/prompt-guide/, and https://cloud.google.com/blog/products/ai-machine-learning/ultimate-prompting-guide-for-lyria-3-pro296- Seed Audio: https://seed.bytedance.com/en/blog/from-speech-to-audio-creation-introducing-the-seed-audio-1-0-audio-creation-model and https://docs.byteplus.com/en/docs/byteplusvoice/seedaudio-01297- Independent preference evidence: https://artificialanalysis.ai/image/leaderboard/text-to-image, https://artificialanalysis.ai/image/leaderboard/editing, https://artificialanalysis.ai/video/leaderboard/text-to-video, https://artificialanalysis.ai/video/leaderboard/image-to-video, and https://artificialanalysis.ai/video/leaderboard/video-editing