Faceless Video
Create a coherent faceless video from a script, reusable visual assets, ordered
scenes, narration, and a verified final edit. This skill owns narrative and
production structure. Use the sibling krea-generate skill for live model
selection, model-specific prompting, and generation constraints.
Boundary
Use this skill when the requested deliverable is a multi-scene production whose story is carried by an off-screen narrator or supplied narration. Topic alone is not enough: “make a video about Rome” is generic video generation unless the request also asks for a faceless channel, narrated explainer, documentary, storybook, picture story, or equivalent format.
Route elsewhere:
- A single clip or generic animation belongs to
krea-generate. - A product, brand, reveal film, advertisement, or UGC piece belongs to the relevant marketing or motion skill.
- A visible presenter or talking head belongs to a presenter-led workflow.
- Planning or ideas without production should be answered without generating.
Delivery contract
Resolve the real narration and post-production path before paid generation. A finished narrated video requires:
- narration supplied with a matching script and scene timings;
- narration generated by an available model that supports the requested audio; or
- a separate available narration capability.
Separate model generations do not guarantee identical narrator identity, exact wording, or clean audio joins. If the user requires one locked voice and the available capabilities cannot provide it, explain that before spending and offer an ordered visual scene package.
Likewise, produce one assembled video only when the environment can actually combine and verify the clips and audio. Otherwise deliver the ordered clips, visual assets, script, and edit manifest, clearly labeled as a scene package. Never describe separate clips as a finished video.
Burn subtitles only when a real transcription or timestamp source is available. Do not estimate word timings from the script.
1. Lock the brief
Collect only missing decisions:
- format: Explainer, History/Documentary, Kids Story, Fairy Tale/Myth, or Picture Story;
- topic, premise, or supplied script;
- target duration and destination;
- aspect ratio;
- visual style or supplied references;
- narration source and language;
- subtitles yes/no;
- cover yes/no.
For an ordinary production request, ask material questions together before generating. For an explicitly hands-off request, choose sensible defaults but state any narration or delivery limitation before the first paid generation.
For factual history, science, health, finance, or current-events scripts, research authoritative sources, retain their URLs, and cross-check important spoken numbers. Mark unsupported claims as unverified rather than filling gaps from uncertain memory.
2. Write the story before generating
Divide the total runtime into clip windows supported by the selected model. Ten-second blocks are a useful creative default, not a universal constraint.
Shape the story as hook → escalating build → turn → payoff:
- Open on the strongest concrete claim, question, or image. Avoid greetings and “in this video” preambles.
- Give each block one narration idea. Every shot in that block should visualize the same idea.
- Escalate stakes, scale, specificity, or surprise. If the build blocks can be reordered without loss, rewrite them.
- Choose one recurring physical through-line—a fuse, growing stack, filling map, changing clock, or travelling object. Advance it in every block and resolve it in the payoff.
- Narrated characters emote and gesture but do not speak or lip-sync unless the user explicitly requests visible dialogue and the selected model supports it.
For each block record its narration, location, required characters and props, through-line state, and timed shots. For a ten-second clip, three to five shots is a useful range; use proportionally fewer for shorter clips. Adjacent shots must change size or angle. Establish a location once, then return on coverage or detail rather than repeating the same wide shot.
Maintain an edit manifest containing:
- brief, aspect ratio, duration, language, and narration path;
- one locked style formula;
- factual source URLs;
- an asset registry with stable IDs, asset type, result location, and generation provenance;
- ordered scene blocks with narration, shots, asset IDs, result location, generation provenance, status, and retry notes;
- final assembly result when one exists.
Persist the manifest when the host supports durable files; otherwise keep it in working context. Never infer scene order from filenames, completion time, or a directory listing.
3. Lock one visual language
Write one reusable style formula describing medium, line or surface treatment, shading, palette, accent color, background treatment, depth, lighting, and motion character. Keep it non-photoreal unless the user asks otherwise. Translate named references into visible craft traits instead of imitating a living artist, studio, or brand.
Inspect supplied references before using them. Treat them as style donors unless the user owns and explicitly wants the depicted character. Do not copy logos, embedded text, or an identifiable real person's likeness without permission.
Create and inspect one style-sample vignette before producing the reusable asset set. Record its result and provenance before any dependent generation uses it.
4. Build reusable assets
Derive the smallest roster that covers the locked script:
- characters with readable full-body silhouettes;
- dressed locations without people and with one named anchor object;
- isolated props, including the through-line object;
- alternate location angles only where the shot plan needs them.
Use the exact locked style formula throughout. Do not begin a dependent scene until every required asset has completed successfully, been inspected, and been recorded in the manifest. Prompt-only substitutes for missing characters, locations, or props create identity and style drift.
5. Generate ordered scenes
Default final scene generation to Seedance 2.5
(bytedance/seedance-2-5) when it is live and its current schema supports the
brief. Read ../krea-generate/references/models/seedance-2.md before writing
its prompts. Use a faster or cheaper model for explicitly economical drafts,
Seedance 2 when 4K is required, or another capable model when the references or
requested behavior are incompatible. Live model availability and schemas remain
authoritative.
Before an expensive run, present the scene count, seconds per scene, resolution, and retry allowance, then obtain approval unless the user already authorized that spend.
Generate blocks in manifest order. For each block:
- Include only its completed location, visible characters, props, and style sample.
- Preserve reference order across retries.
- Repeat the locked style formula unchanged.
- Give every timed shot one camera behavior and one physical event.
- Start motion in the first frame and settle into a still-moving final frame.
- For external narration, exclude dialogue, lip movement, captions, and competing music. For native audio, use the approved narration line while acknowledging voice-consistency limits.
- Put duration, aspect ratio, resolution, audio, and references in the capability's declared inputs rather than relying on prompt prose.
- Wait for completion and record the final result before continuing.
Inspect each completed clip at the beginning, middle, and end. Reject the wrong aspect ratio, missing subjects, severe style drift, accidental speech, frozen openings, or broken continuity.
6. Retry without losing the film
Retry only a terminal failure or a concrete manifest violation:
- Repeat the same request once with identical ordered references.
- If it fails again, clarify the ambiguous action while preserving the story beat.
- If framing caused the defect, change that block's framing once.
- Keep passed blocks immutable; never drop a failed block or replace it with a neighboring result.
- Stop and ask before further paid attempts.
Record every attempt and defect in the manifest.
7. Assemble and deliver
When assembly is available, normalize scene dimensions, frame rate, codecs, pixel format, and audio layout before concatenation. Preserve manifest order. For supplied narration, follow its approved scene timings without pitch shifting or speed changes; adjust visual timing instead. For native per-scene narration, verify every audio join when playback or audio inspection is available and state the limitation when it is not.
Verify the final media can be decoded and inspect the beginning, middle, end, and every scene boundary. Delivery passes only when:
- the output has the requested runtime and aspect ratio;
- every scene appears exactly once in manifest order;
- style, characters, locations, and through-line remain coherent;
- there are no black frames, broken joins, accidental talking characters, or unresolved blocking retries;
- narration, subtitles, cover, and assembly are claimed only when produced;
- the payoff resolves both the hook and the physical through-line.
Deliver either the verified final video or the explicitly labeled ordered scene package. Do not leave the user with raw job metadata as the deliverable.