Morph Explainer Video
A fact in, a finished vertical short out — in one fixed house style, on Unsora. Or just the script, or one segment's prompts, if that's all they want.
The idea in one line: pick the script format → write to its beat template → voice it first → cut the picture to the voice → chain keyframes so the film reads as one unbroken morph → score it → master to −14 dB.
The pipeline in one line:
elicit (format, subject, length) → write script to template → [GATE 1: script]
→ create_voiceover (ONE take) → ffprobe → derive keyframe times & segment durations
→ keyframe + segment plan → [GATE 2: shot plan + credits]
→ create_image K0..Kn (chain: each Ki is referenceImage for Ki+1)
→ create_video per segment (image=Ki, lastImage=Ki+1) → ffmpeg concat
→ create_music (mureka-7.5) → mix: bed + VO + 2 stings → −14 dB RMS → final.mp4
If they want only the script, run Part 1 and stop at Gate 1. If they want one segment's prompts, run Part 1 + Phase 4 for that segment and stop.
Two required reads before writing a single prompt:
references/house-style.md— the twelve slots, filled once and locked. This style does not vary. Read it before any prompt.references/pipeline.md— the exact Unsora calls, the keyframe chain, the ffmpeg mix, the master target.
And one of references/formats/hypothetical.md or references/formats/true-story.md
before writing a word of script.
PART 0 — THE FRONT DOOR
Three parameters. Infer what you can; ask only what's missing. On a client with tappable inputs, use them.
- Format — which script architecture:
hypothetical— "If you did [absurd thing] to a body, what happens?" Ends on a deadpan punchline. ~117 words / ~43s.true-story— "This actually happened to a person." Ends on "now they plan to…" ~94 words / ~33s. A "what if" or "can you" phrasing answers this — don't re-ask. A real named person or event answers it the other way.
- Subject — the body fact or question. If they gave a script, read it and classify against the templates. If they gave a bare topic, you write the script; that's Phase 1.
- Length — default to the format's native length. Only ask if they want a non-standard runtime, and warn: these templates are tuned to their word counts, and stretching them dilutes the format.
Orientation is not a parameter. It is always 9:16. Neither is the look — there is one house style and it is locked.
Never generate (which spends the user's Unsora credits) before Gate 2.
PART 1 — THE LAWS
The style DNA never varies
Unlike a multi-style producer, this skill has one look, defined in
references/house-style.md and locked across every film. What varies is the
script architecture — that's the swappable profile here. Slots 1–11 of the
house style are constants; only slot 12 (the hook rule) forks by format.
Cut freely; morph only inside a shot
This section was wrong until it was measured. An earlier version claimed "one hard cut in 77 seconds" and treated cutting as a rare concession. The opposite is true, and building to the old rule produces glitchy welded joins.
Measured on reference R1 (37.70s): 6 shots, mean 6.3s, median 5.25s, shortest 2.8s, longest 11.1s — hard cuts at 4.90 / 16.03 / 21.17 / 29.53 / 32.30s, plus 22 softer boundaries. Budget 5–8s per shot. Cuts are the default grammar.
Two independently generated clips will not join invisibly. Do not try. Give each shot its own start frame and cut cleanly; continuity comes from consistent character and setting references, not from welding. Morphing is what can happen inside one clip when a shot happens to run long — it is not the structure of the film. Two levels:
- Inside a segment — never cut. One Seedance clip morphs smoothly from its
start frame to its end frame, both pinned (
image+lastImage). A cut appearing within a clip is a real failure (the model inserted a shot change) and the negative prompt kills it. - Between segments — cut. This is the normal case, roughly every 5–8s. Two independently generated clips will not join into an invisible take; forcing it reads as a glitch. A match cut (segment i's end keyframe reused as segment i+1's start) is available when a beat genuinely continues the same action, but it is one option among several, not the default — R1 uses honest scene cuts for most of its six boundaries.
Never paper over a scene change with a crossfade. If the beat changes, cut.
The chain, mechanically. Segment i is generated with image: K_i and
lastImage: K_i+1; segment i+1 is generated with image: K_i+1. For a match
cut those two are the identical file. The morph prompt still says "single
continuous take, no cuts" — that governs the inside of the one clip.
Voice first, picture second
The narration is the spine. The VO is generated as one unbroken take before any picture exists, then the picture is cut to it. This inverts the usual picture-first order, and it's correct here for three measured reasons: the reference films have zero silence (never below −21 dB), the read is one continuous deadpan with no per-shot gaps, and the segment boundaries are loose (8–12s) rather than tight beats. Timing the voice to the picture would fight all three.
Practical upside: one create_voiceover call for a whole 117-word script is
~650 characters = 6 credits, versus 6 credits × 8 segments if you chop it up.
Nothing is ever static
Zero pixels in the reference films hold still. Measured on R1: mean temporal standard deviation across the whole frame is 63.2 (of 255), per-shot range 40.7–63.6, frame-to-frame mean absolute difference median 23.0. (An earlier version of this file claimed 15.9; that number had no source.)
This is a dense, fast edit — subjects cross frame, camera pushes and drifts, action fills the shot. A shot where only the camera moves over a static subject is below standard. Which leads to:
No text in generated frames — but the finished film IS captioned
Two rules, previously conflated into one wrong rule:
- No text in any keyframe prompt. Not a title, caption, label, or lower third. If a prompt would put a word on screen, delete it.
- Captions go on in post. R1 carries a white caption layer throughout, synced to narration. An earlier version of this file claimed the reference films had no text at all — that was an observation error.
The only exception to rule 1 is a word that exists as a physical object in the world (the "SALINE" printed on a syringe barrel, the "NEW PLAN" stamped on a clipboard) — that's a prop, not a title.
165 words per minute
Measured: 161.5 and 169.5 WPM; 2.69 and 2.82 words/second. Write to 2.75 wps and the runtime falls out of the word count. A beat's word budget is fixed by the format template — respect it, because the templates are the format.
Spend nothing before the gate
Two gates. Script approved at Gate 1 (free). Shot plan + credit estimate approved
at Gate 2 (before the first generate call). Unsora has no cost-preflight
parameter — unlike some connectors, you cannot price a job without submitting
it. So quote an estimate, check get_credits, and get an explicit yes.
The two-prompt structure
Every segment is two artifacts:
- Keyframe prompts (image model) — one per state, not per segment. An n-segment film has n+1 keyframes. Ordered: Subject → State → Setting → Style → Camera → Lighting → Constraints. Each keyframe carries the house medium sentence and the AI-default negative.
- Morph prompt (video model) — names both frames by role ("the shot opens on the start frame and resolves to the end frame"), never re-describes their design, states the deformation in the house motion verbs, and closes with the audio line.
The frame-zero problem, inverted
Most producers fight the hero frame being the end state. Here both ends are
pinned — image and lastImage are both supplied — so the model's job is purely
to interpolate. The failure mode flips: the model arrives early and holds,
producing a morph that finishes at 4s and then sits dead for 7s while inventing
motion. The fix is in the prompt: state the deformation as continuous across the
full duration ("the change is gradual and unbroken across all N seconds; it is
never complete before the final frame"). See the failure table.
PART 2 — THE PRODUCTION PIPELINE
Read references/pipeline.md for exact Unsora parameters before Phase 5.
Phase 1 — Script
Read the format profile, then write to its beat table. Every beat has a word
budget; hit it within ±2 words. If the user supplied a script, classify it
against the two templates and say plainly where it deviates — a script with no
punchline isn't hypothetical, it's a true-story that needs a "now they plan
to" ending.
If the anti-slop skill is available, run the draft through it. This is
public-facing copy and the format dies instantly on AI narration voice — no "in
a world where", no "the results may surprise you", no hype adjectives.
Gate 1: show the script and wait for a yes. Show it as the beat table with word counts and estimated seconds. Free to change here, expensive later.
Phase 2 — Voice it, then measure
One create_voiceover call, whole script, one take. Voice, stability, and the
<#x#> pause-tag placement are in references/pipeline.md §2. Then ffprobe
its real duration. That number is the film's runtime. Everything downstream
is derived from it, not from the plan.
Phase 3 — Derive the keyframe plan
Split the VO's real duration into segments. Seedance 2.0 accepts only these lengths — 4, 5, 6, 8, 10, 12, 15s — with 15s the hard ceiling (not 20; Unsora's schema allows 20 across all its models, but Seedance rejects anything over 15 and anything off the fixed set). Work in 8–12s segments, floor 6.
Two constraints fight, and the legal-duration one wins. A boundary wants to land on a beat end and make a segment whose length is on the allowed set. You cannot always have both exactly, so resolve in this order: (1) mark every beat-end time as a candidate boundary; (2) choose the set of candidates whose gaps each round to an allowed value within ±1.5s; (3) if a gap still can't be made legal (e.g. a 13.7s stretch that would round up to 15 and overshoot), move the boundary earlier onto the previous beat end rather than forcing 15 — never let a segment exceed its real content, and never pick a duration off the set. Only after the durations are legal do you lock the split. Place boundaries on beat boundaries, never mid-sentence. Then:
nsegments →n+1keyframes,K0 … KnK0is the film's first frame; it must already show the premise. The hook sentence says "like this" and points at something that is already on screen at 0:00. There is no establishing beat.Knis the last frame — the punchline state (hypothetical) or the future-facing state (true-story).- Each interior
Kiis a designed state on a beat boundary: the thing the deformation has become by then.
A 43s film → 4 segments (10/10/10/12) → 5 keyframes. A 33s film → 3 segments (12/12/10) → 4 keyframes. Every length is on Seedance's allowed set. When a raw gap lands between two legal values, round to the nearer and absorb the ±1–2s slack across the other segments so the total still matches the VO. When the concat is built, a segment generated at a legal length can be trimmed to its exact beat window in ffmpeg — so prefer generating the nearest legal length at or above the raw gap and trimming down, rather than generating short and leaving a gap.
At each boundary, decide the cut type (this is a Gate-2 decision):
- Match cut (default) — same subject continues deforming. The end keyframe of segment i IS the start keyframe of segment i+1 (one shared image), so the join is nearly seamless.
- Scene cut — the beat changes location (e.g. the field → the operating
room). Segment i+1 gets its own fresh
Kand the boundary is an honest cut. Keep these rare; the reference films use about one.
Phase 4 — Write the prompts, then stop
Write all n+1 keyframe prompts and all n morph prompts. Show them in the
per-segment format below.
Gate 2: the shot plan. A table, then the numbers, then wait:
| Seg | t_start → t_end | Sec | Start K | End K | Cut in | The deformation | VO beats |
Below it: the keyframe list (one line each — Ki at t, one sentence of the
state), total runtime, segment count, keyframe count, and — plainly — that
generating spends Unsora credits, roughly n+1 image jobs + n video jobs + 1
music job (the VO is already paid for), and takes roughly n × a few minutes of
polling. Report the current balance from get_credits. Do not call a generation
tool before an explicit go.
Per-segment output format:
### SEGMENT N — [TITLE]+ one line: the deformation and why it lands.- Breakdown — duration; VO beats covered; start state; end state; the deformation; camera; the accent color, if any; audio.
- Start keyframe prompt — fenced block. (Skip if already written for the previous segment's end — each keyframe is written once.)
- End keyframe prompt — fenced block.
- Morph prompt — fenced block.
Phase 4.5 — Calibrate against a real frame. MANDATORY.
Never generate a full keyframe set from an unvalidated prompt. This phase exists because the style vocabulary is genuinely hard to hit and the failure is invisible from the prompt alone — it only shows up next to a reference frame.
Pick the single hardest keyframe in the plan (most character detail, most palette range) and generate one image from it.
Extract the nearest matching frame from a reference video:
ffmpeg -ss <t> -i <ref> -frames:v 1 -q:v 2 ref.jpgLook at both. Compare on four axes, in this order:
Axis Right Wrong Render medium Game-engine uncanny Photoreal, or Pixar-polished Grade Warm browns vs desaturated teal Neutral, teal lost Depth Background falls off hard Everything sharp Framing Subject large, foreground crowds bottom Centered portrait If any axis is off, fix the prompt and re-run one image. Iterate here, where it costs one image, not after twenty.
Only when the frame matches, carry the corrected style block into every keyframe prompt and proceed.
Do not skip this because the prompt "looks right". In the calibration run that produced the current house style, iteration 1 read as perfectly reasonable and was wrong — it landed in the animated-feature ditch because it contained the words "high-end animated feature". One 2-credit image caught it.
If the generator accepts reference images, use them. Passing the actual
reference frame in as a reference beats describing it in words. Check the
generator's image tool for a referenceImages parameter before falling back to
text-only calibration.
Phase 5 — Generate
references/pipeline.md has the exact calls. In order:
- Keyframes, in chain order,
K0first. EachKi+1passesKi's result URL inreferenceImages— this is what keeps the subject from redrawing itself between states. Poll each withwait_for_imagebefore generating the next; the chain is strictly sequential. - Segments.
create_videowithimage: Ki_url,lastImage: Ki+1_url,model: seedance-2.0,aspectRatio: "9:16",duration: <one of 4/5/6/8/10/12/15>,generateAudio: false. Segments are independent of each other and can be fired in parallel — they're pinned at both ends, so nothing drifts. (A scene-cut boundary just means segment i+1 uses its own freshK, not the shared one.) - Music.
create_music,mureka-7.5, instrumental, no lyrics.
Keep the manifest updated after every call so a failed segment regenerates without touching the rest. Never re-roll a completed segment "for consistency" — the keyframes already guarantee it, and a re-roll returns different.
Phase 6 — Assemble and master
ffmpeg concat the segments → picture.mp4. Verify its duration against the
VO's. Then the mix: bed ducked under VO, VO at unity, 2 stings on the beats the
format marks, master to −14 dB RMS with peaks under 0.98. Exact commands in
references/pipeline.md §6.
Verify the final duration against picture.mp4 — a dropped segment yields a
valid file that's simply the wrong length. Present with present_files. Offer,
never perform, a social post.
Failure table
| Symptom | Cause | Fix |
|---|---|---|
| A match-cut boundary looks like a jump, not a continuation | The two keyframes aren't the same file — segment i+1's image isn't the identical URL as segment i's lastImage |
Re-read the manifest; for a match cut both must be the same string. #1 silent failure. (If the boundary was meant to be a scene cut, this is fine — leave it) |
| Morph finishes early, then the shot sits dead and invents motion | Both ends pinned, model raced to the end | State the change as continuous across the full duration and never complete before the final frame; if it persists, shorten the segment |
duration rejected or clamped |
Passed a value off Seedance's set (e.g. 11, 13, 20) | Use only 4/5/6/8/10/12/15; 15 is the max. Re-split so every segment lands on an allowed value |
| Looks like real live-action / stock medical footage | The AI-default betrayal for this style | Add the house medium sentence + the photoreal negative to every keyframe prompt |
| A shot change appears inside one segment | Video model inserted a cut within the clip | "Single continuous take, one shot, no cuts, no scene changes" in every morph prompt + the cut terms in the negative prompt |
| A forced-seamless join looks like a glitch (ghosting, stutter) | Tried to weld a real scene change into one take | Make it an honest cut instead — the format allows between-segment cuts |
| Subject redraws itself between keyframes | referenceImages chain not carried |
Every Ki+1 must pass Ki's URL in referenceImages; regenerate only the drifted keyframe and everything after it |
| Words appear on screen | Text leaked into a keyframe prompt | No titles, ever. Delete. Only physical props may carry printing |
| Cyan drifts to blue/teal between segments | Backdrop hex not restated | The exact hex goes in every keyframe prompt, not just the first |
| Film is shorter than the VO | A segment failed and got skipped | Check the manifest, regenerate the missing segment |
| VO runs past the picture | Segment durations derived from the plan, not the real VO | Re-derive from ffprobe of the VO; the voice is the clock |
| Narration sounds hyped / performed | Wrong stability, or the copy is over-written |
stability: 0.75+, and cut the adjectives — the format is deadpan |
| Mix is quiet next to real feed videos | No master stage | Normalize to −14 dB RMS; the reference films both measure exactly that |
| Music drowns the narration | No ducking | Sidechain the bed under the VO |
Checklist — per segment, before generating
- The deformation is a real change of state, not a camera move on a static thing.
- The house medium sentence and the photoreal negative are in both keyframe prompts.
- Keyframe prompts ordered Subject → State → Setting → Style → Camera → Lighting → Constraints.
- The exact backdrop hex is restated in every keyframe prompt.
- No text in frame anywhere, unless it's printing on a physical prop.
- Morph prompt: only house motion verbs; "single continuous take, no cuts"; the change is continuous across the full duration; closes with the audio line.
aspectRatio: "9:16"on every image and video call.durationis one of 4/5/6/8/10/12/15, identical in the plan and the call.- At a match-cut boundary, segment i+1's
imageis byte-identical to segment i'slastImage. At a scene-cut boundary, segment i+1 uses its own fresh keyframe (and that's intentional).
Checklist — before delivering
- Script hit its template's word budgets; VO is one take; no intro, no outro, no CTA.
- Picture duration matches VO duration within ~0.3s.
- No cut appears inside any single clip — verify with
ffmpeg select='gt(scene,0.3)'. Detections at segment boundaries are expected for scene cuts and fine; a detection mid-segment is a failure. - Exactly 2 stings, on the beats the format profile marks.
- Bed is continuous, never silent; ducked under VO.
- Final master ≈ −14 dB RMS, peak < 0.98.
- No burned-in text; full-bleed 9:16.
Reference files
references/house-style.md— the twelve slots, locked. The medium sentence, the AI-default negative, the exact palette, the motion verbs, the prompt templates. Read before any prompt.references/pipeline.md— exact Unsora calls (create_image,create_videowith theimage/lastImagechain,create_voiceover,create_music, thewait_for_*pollers), the ffmpeg concat and mix, the −14 dB master, the manifest schema, resume-after-failure.references/formats/hypothetical.md— the 8-beat "what if you did this to a body" template. ~117 words / ~43s.references/formats/true-story.md— the 5-beat "this happened to a person" template. ~94 words / ~33s.references/formats/format-framework.md— how to build a third format profile without breaking the house style.