# Morph Explainer Video

> Turn a fact, a body question, or a script into a finished vertical morph-explainer short: the deadpan anatomical/body-horror format that reads as one unbroken morph, the subject deforming continuously against a flat cyan backdrop while a flat narrator reads 165 words a minute over a wall-to-wall bed. Built end to end on Unsora: keyframe chain via create_image, morph segments via create_video (seedance-2.0 image + lastImage), one-take narration via create_voiceover (ElevenLabs v3), bed via create_music, then ffmpeg stitch and a -14 dB master. Use when the user wants a video in this style, says "make a Zack D. Films style video", "body horror explainer", "cursed science short", "morph video", "what if you did X to a human body", "turn this fact into a short", or hands over a script or a weird body fact and wants the finished MP4. Also use for a single morph segment's prompt pair, or to write only the script to the house templates. Default to this for any deadpan anatomical explainer built on continuous morphs.

- Skill: `sadekxd/morph-explainer-video` (Agent Skill, multi-file: 2 files)
- Install (CLI): `npx skillmds@latest add sadekxd/morph-explainer-video`
- Raw SKILL.md: https://api.skillmd.com/api/skills/sadekxd/morph-explainer-video/raw
- Safety review: pending (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: AI & ML
- Author: sadekxD (https://skillmd.com/u/sadekxd)
- Updated: 2026-09-17
- Page: https://skillmd.com/skills/sadekxd/morph-explainer-video

---


# Morph Explainer Video

A fact in, a finished vertical short out — in one fixed house style, on Unsora.
Or just the script, or one segment's prompts, if that's all they want.

**The idea in one line:** pick the script format → write to its beat template →
**voice it first** → cut the picture to the voice → chain keyframes so the film
reads as one unbroken morph → score it → master to −14 dB.

**The pipeline in one line:**

```
elicit (format, subject, length) → write script to template → [GATE 1: script]
  → create_voiceover (ONE take) → ffprobe → derive keyframe times & segment durations
  → keyframe + segment plan → [GATE 2: shot plan + credits]
  → create_image K0..Kn (chain: each Ki is referenceImage for Ki+1)
  → create_video per segment (image=Ki, lastImage=Ki+1) → ffmpeg concat
  → create_music (mureka-7.5) → mix: bed + VO + 2 stings → −14 dB RMS → final.mp4
```

If they want only the script, run Part 1 and stop at Gate 1. If they want one
segment's prompts, run Part 1 + Phase 4 for that segment and stop.

**Two required reads before writing a single prompt:**
1. `references/house-style.md` — the twelve slots, filled once and **locked**.
   This style does not vary. Read it before any prompt.
2. `references/pipeline.md` — the exact Unsora calls, the keyframe chain, the
   ffmpeg mix, the master target.

And one of `references/formats/hypothetical.md` or `references/formats/true-story.md`
before writing a word of script.

---

# PART 0 — THE FRONT DOOR

Three parameters. Infer what you can; ask only what's missing. On a client with
tappable inputs, use them.

1. **Format** — which script architecture:
   - **`hypothetical`** — "If you did [absurd thing] to a body, what happens?"
     Ends on a deadpan punchline. ~117 words / ~43s.
   - **`true-story`** — "This actually happened to a person." Ends on "now they
     plan to…" ~94 words / ~33s.
   A "what if" or "can you" phrasing answers this — don't re-ask. A real named
   person or event answers it the other way.
2. **Subject** — the body fact or question. If they gave a script, read it and
   classify against the templates. If they gave a bare topic, you write the
   script; that's Phase 1.
3. **Length** — default to the format's native length. Only ask if they want a
   non-standard runtime, and warn: these templates are tuned to their word
   counts, and stretching them dilutes the format.

Orientation is **not** a parameter. It is always 9:16. Neither is the look —
there is one house style and it is locked.

Never generate (which spends the user's Unsora credits) before Gate 2.

---

# PART 1 — THE LAWS

## The style DNA never varies

Unlike a multi-style producer, this skill has **one look**, defined in
`references/house-style.md` and locked across every film. What varies is the
**script architecture** — that's the swappable profile here. Slots 1–11 of the
house style are constants; only slot 12 (the hook rule) forks by format.

## Cut freely; morph only inside a shot

**This section was wrong until it was measured.** An earlier version claimed
"one hard cut in 77 seconds" and treated cutting as a rare concession. The
opposite is true, and building to the old rule produces glitchy welded joins.

Measured on reference R1 (37.70s): **6 shots, mean 6.3s, median 5.25s, shortest
2.8s, longest 11.1s** — hard cuts at 4.90 / 16.03 / 21.17 / 29.53 / 32.30s, plus
22 softer boundaries. **Budget 5–8s per shot. Cuts are the default grammar.**

Two independently generated clips will not join invisibly. Do not try. Give each
shot its own start frame and cut cleanly; continuity comes from consistent
character and setting references, not from welding. Morphing is what can happen
*inside* one clip when a shot happens to run long — it is not the structure of
the film. Two levels:

- **Inside a segment — never cut.** One Seedance clip morphs smoothly from its
  start frame to its end frame, both pinned (`image` + `lastImage`). A cut
  appearing *within* a clip is a real failure (the model inserted a shot change)
  and the negative prompt kills it.
- **Between segments — cut.** This is the normal case, roughly every 5–8s. Two
  independently generated clips will not join into an invisible take; forcing it
  reads as a glitch. A **match cut** (segment *i*'s end keyframe reused as
  segment *i+1*'s start) is available when a beat genuinely continues the same
  action, but it is one option among several, not the default — R1 uses honest
  scene cuts for most of its six boundaries.

Never paper over a scene change with a crossfade. If the beat changes, cut.

**The chain, mechanically.** Segment *i* is generated with `image: K_i` and
`lastImage: K_i+1`; segment *i+1* is generated with `image: K_i+1`. For a match
cut those two are the identical file. The morph prompt still says "single
continuous take, no cuts" — that governs the *inside* of the one clip.

## Voice first, picture second

The narration is the spine. The VO is generated **as one unbroken take** before
any picture exists, then the picture is cut to it. This inverts the usual
picture-first order, and it's correct here for three measured reasons: the
reference films have **zero silence** (never below −21 dB), the read is one
continuous deadpan with no per-shot gaps, and the segment boundaries are loose
(8–12s) rather than tight beats. Timing the voice to the picture would fight all
three.

Practical upside: one `create_voiceover` call for a whole 117-word script is
~650 characters = **6 credits**, versus 6 credits × 8 segments if you chop it up.

## Nothing is ever static

Zero pixels in the reference films hold still. Measured on R1: mean temporal
standard deviation across the whole frame is **63.2** (of 255), per-shot range
40.7–63.6, frame-to-frame mean absolute difference median **23.0**. (An earlier
version of this file claimed 15.9; that number had no source.)

This is a dense, fast edit — subjects cross frame, camera pushes and drifts,
action fills the shot. A shot where only the camera moves over a static subject
is below standard. Which leads to:

## No text in generated frames — but the finished film IS captioned

Two rules, previously conflated into one wrong rule:

1. **No text in any keyframe prompt.** Not a title, caption, label, or lower
   third. If a prompt would put a word on screen, delete it.
2. **Captions go on in post.** R1 carries a white caption layer throughout,
   synced to narration. An earlier version of this file claimed the reference
   films had no text at all — that was an observation error.

The only exception to rule 1 is
a word that exists as a **physical object in the world** (the "SALINE" printed on
a syringe barrel, the "NEW PLAN" stamped on a clipboard) — that's a prop, not a
title.

## 165 words per minute

Measured: 161.5 and 169.5 WPM; 2.69 and 2.82 words/second. Write to **2.75 wps**
and the runtime falls out of the word count. A beat's word budget is fixed by the
format template — respect it, because the templates are the format.

## Spend nothing before the gate

Two gates. Script approved at Gate 1 (free). Shot plan + credit estimate approved
at Gate 2 (before the first generate call). Unsora has **no cost-preflight
parameter** — unlike some connectors, you cannot price a job without submitting
it. So quote an estimate, check `get_credits`, and get an explicit yes.

## The two-prompt structure

Every segment is two artifacts:

- **Keyframe prompts** (image model) — one per *state*, not per segment. An
  n-segment film has **n+1 keyframes**. Ordered: Subject → State → Setting →
  Style → Camera → Lighting → Constraints. Each keyframe carries the house
  medium sentence and the AI-default negative.
- **Morph prompt** (video model) — names both frames by role ("the shot opens on
  the start frame and resolves to the end frame"), never re-describes their
  design, states the deformation in the house motion verbs, and closes with the
  audio line.

## The frame-zero problem, inverted

Most producers fight the hero frame being the *end* state. Here both ends are
pinned — `image` and `lastImage` are both supplied — so the model's job is purely
to interpolate. The failure mode flips: the model **arrives early and holds**,
producing a morph that finishes at 4s and then sits dead for 7s while inventing
motion. The fix is in the prompt: state the deformation as *continuous across the
full duration* ("the change is gradual and unbroken across all N seconds; it is
never complete before the final frame"). See the failure table.

---

# PART 2 — THE PRODUCTION PIPELINE

Read `references/pipeline.md` for exact Unsora parameters before Phase 5.

## Phase 1 — Script

Read the format profile, then write to its beat table. Every beat has a word
budget; hit it within ±2 words. If the user supplied a script, classify it
against the two templates and say plainly where it deviates — a script with no
punchline isn't `hypothetical`, it's a `true-story` that needs a "now they plan
to" ending.

If the `anti-slop` skill is available, run the draft through it. This is
public-facing copy and the format dies instantly on AI narration voice — no "in
a world where", no "the results may surprise you", no hype adjectives.

**Gate 1: show the script and wait for a yes.** Show it as the beat table with
word counts and estimated seconds. Free to change here, expensive later.

## Phase 2 — Voice it, then measure

One `create_voiceover` call, whole script, one take. Voice, stability, and the
`<#x#>` pause-tag placement are in `references/pipeline.md` §2. Then `ffprobe`
its real duration. **That number is the film's runtime.** Everything downstream
is derived from it, not from the plan.

## Phase 3 — Derive the keyframe plan

Split the VO's real duration into segments. Seedance 2.0 accepts only these
lengths — **4, 5, 6, 8, 10, 12, 15s** — with **15s the hard ceiling** (not 20;
Unsora's schema allows 20 across all its models, but Seedance rejects anything
over 15 and anything off the fixed set). Work in **8–12s** segments, floor 6.

**Two constraints fight, and the legal-duration one wins.** A boundary wants to
land on a beat end *and* make a segment whose length is on the allowed set. You
cannot always have both exactly, so resolve in this order: (1) mark every beat-end
time as a candidate boundary; (2) choose the set of candidates whose gaps each
**round to an allowed value within ±1.5s**; (3) if a gap still can't be made legal
(e.g. a 13.7s stretch that would round up to 15 and overshoot), **move the
boundary earlier onto the previous beat end** rather than forcing 15 — never let a
segment exceed its real content, and never pick a duration off the set. Only after
the durations are legal do you lock the split. Place boundaries on **beat
boundaries**, never mid-sentence. Then:

- `n` segments → `n+1` keyframes, `K0 … Kn`
- `K0` is the film's first frame; it must already show the premise. The hook
  sentence says "like this" and points at something that is *already on screen at
  0:00*. There is no establishing beat.
- `Kn` is the last frame — the punchline state (`hypothetical`) or the
  future-facing state (`true-story`).
- Each interior `Ki` is a designed *state* on a beat boundary: the thing the
  deformation has become by then.

A 43s film → 4 segments (10/10/10/12) → 5 keyframes. A 33s film → 3 segments
(12/12/10) → 4 keyframes. Every length is on Seedance's allowed set. When a raw
gap lands between two legal values, round to the nearer and absorb the ±1–2s slack
across the other segments so the total still matches the VO. When the concat is
built, a segment generated at a legal length can be **trimmed** to its exact beat
window in ffmpeg — so prefer generating the nearest legal length *at or above* the
raw gap and trimming down, rather than generating short and leaving a gap.

**At each boundary, decide the cut type** (this is a Gate-2 decision):
- **Match cut** (default) — same subject continues deforming. The end keyframe of
  segment *i* IS the start keyframe of segment *i+1* (one shared image), so the
  join is nearly seamless.
- **Scene cut** — the beat changes location (e.g. the field → the operating
  room). Segment *i+1* gets its own fresh `K` and the boundary is an honest cut.
  Keep these rare; the reference films use about one.

## Phase 4 — Write the prompts, then stop

Write all `n+1` keyframe prompts and all `n` morph prompts. Show them in the
per-segment format below.

**Gate 2: the shot plan.** A table, then the numbers, then wait:

| Seg | t_start → t_end | Sec | Start K | End K | Cut in | The deformation | VO beats |

Below it: the keyframe list (one line each — `Ki` at `t`, one sentence of the
state), total runtime, segment count, keyframe count, and — plainly — that
generating spends Unsora credits, roughly `n+1` image jobs + `n` video jobs + 1
music job (the VO is already paid for), and takes roughly `n × a few minutes` of
polling. Report the current balance from `get_credits`. Do not call a generation
tool before an explicit go.

**Per-segment output format:**

1. `### SEGMENT N — [TITLE]` + one line: the deformation and why it lands.
2. **Breakdown** — duration; VO beats covered; start state; end state; the
   deformation; camera; the accent color, if any; audio.
3. **Start keyframe prompt** — fenced block. (Skip if already written for the
   previous segment's end — each keyframe is written **once**.)
4. **End keyframe prompt** — fenced block.
5. **Morph prompt** — fenced block.

## Phase 4.5 — Calibrate against a real frame. MANDATORY.

**Never generate a full keyframe set from an unvalidated prompt.** This phase
exists because the style vocabulary is genuinely hard to hit and the failure is
invisible from the prompt alone — it only shows up next to a reference frame.

1. Pick the single hardest keyframe in the plan (most character detail, most
   palette range) and generate **one** image from it.
2. Extract the nearest matching frame from a reference video:
   `ffmpeg -ss <t> -i <ref> -frames:v 1 -q:v 2 ref.jpg`
3. **Look at both.** Compare on four axes, in this order:

   | Axis | Right | Wrong |
   |---|---|---|
   | Render medium | Game-engine uncanny | Photoreal, or Pixar-polished |
   | Grade | Warm browns vs desaturated teal | Neutral, teal lost |
   | Depth | Background falls off hard | Everything sharp |
   | Framing | Subject large, foreground crowds bottom | Centered portrait |

4. If any axis is off, fix the prompt and re-run **one** image. Iterate here,
   where it costs one image, not after twenty.
5. Only when the frame matches, carry the corrected style block into every
   keyframe prompt and proceed.

**Do not skip this because the prompt "looks right".** In the calibration run
that produced the current house style, iteration 1 read as perfectly reasonable
and was wrong — it landed in the animated-feature ditch because it contained the
words "high-end animated feature". One 2-credit image caught it.

**If the generator accepts reference images, use them.** Passing the actual
reference frame in as a reference beats describing it in words. Check the
generator's image tool for a `referenceImages` parameter before falling back to
text-only calibration.

## Phase 5 — Generate

`references/pipeline.md` has the exact calls. In order:

1. **Keyframes, in chain order, `K0` first.** Each `Ki+1` passes `Ki`'s result
   URL in `referenceImages` — this is what keeps the subject from redrawing
   itself between states. Poll each with `wait_for_image` before generating the
   next; the chain is strictly sequential.
2. **Segments.** `create_video` with `image: Ki_url`, `lastImage: Ki+1_url`,
   `model: seedance-2.0`, `aspectRatio: "9:16"`, `duration: <one of 4/5/6/8/10/12/15>`,
   `generateAudio: false`. Segments are independent of each other and can be
   fired in parallel — they're pinned at both ends, so nothing drifts. (A scene-cut
   boundary just means segment *i+1* uses its own fresh `K`, not the shared one.)
3. **Music.** `create_music`, `mureka-7.5`, instrumental, no lyrics.

Keep the manifest updated after every call so a failed segment regenerates
without touching the rest. Never re-roll a completed segment "for consistency" —
the keyframes already guarantee it, and a re-roll returns different.

## Phase 6 — Assemble and master

`ffmpeg` concat the segments → `picture.mp4`. Verify its duration against the
VO's. Then the mix: bed ducked under VO, VO at unity, 2 stings on the beats the
format marks, **master to −14 dB RMS with peaks under 0.98**. Exact commands in
`references/pipeline.md` §6.

Verify the final duration against `picture.mp4` — a dropped segment yields a
valid file that's simply the wrong length. Present with `present_files`. Offer,
never perform, a social post.

---

## Failure table

| Symptom | Cause | Fix |
|---|---|---|
| A **match-cut** boundary looks like a jump, not a continuation | The two keyframes aren't the same file — segment *i+1*'s `image` isn't the identical URL as segment *i*'s `lastImage` | Re-read the manifest; for a match cut both must be the same string. #1 silent failure. (If the boundary was *meant* to be a scene cut, this is fine — leave it) |
| Morph finishes early, then the shot sits dead and invents motion | Both ends pinned, model raced to the end | State the change as continuous across the full duration and never complete before the final frame; if it persists, shorten the segment |
| `duration` rejected or clamped | Passed a value off Seedance's set (e.g. 11, 13, 20) | Use only 4/5/6/8/10/12/15; 15 is the max. Re-split so every segment lands on an allowed value |
| Looks like real live-action / stock medical footage | The AI-default betrayal for this style | Add the house medium sentence + the photoreal negative to **every** keyframe prompt |
| A shot change appears *inside* one segment | Video model inserted a cut within the clip | "Single continuous take, one shot, no cuts, no scene changes" in every morph prompt + the cut terms in the negative prompt |
| A forced-seamless join looks like a glitch (ghosting, stutter) | Tried to weld a real scene change into one take | Make it an honest cut instead — the format allows between-segment cuts |
| Subject redraws itself between keyframes | `referenceImages` chain not carried | Every `Ki+1` must pass `Ki`'s URL in `referenceImages`; regenerate only the drifted keyframe and everything after it |
| Words appear on screen | Text leaked into a keyframe prompt | No titles, ever. Delete. Only physical props may carry printing |
| Cyan drifts to blue/teal between segments | Backdrop hex not restated | The exact hex goes in every keyframe prompt, not just the first |
| Film is shorter than the VO | A segment failed and got skipped | Check the manifest, regenerate the missing segment |
| VO runs past the picture | Segment durations derived from the plan, not the real VO | Re-derive from `ffprobe` of the VO; the voice is the clock |
| Narration sounds hyped / performed | Wrong `stability`, or the copy is over-written | `stability: 0.75`+, and cut the adjectives — the format is deadpan |
| Mix is quiet next to real feed videos | No master stage | Normalize to −14 dB RMS; the reference films both measure exactly that |
| Music drowns the narration | No ducking | Sidechain the bed under the VO |

## Checklist — per segment, before generating

- The deformation is a real *change of state*, not a camera move on a static thing.
- The house medium sentence **and** the photoreal negative are in both keyframe prompts.
- Keyframe prompts ordered Subject → State → Setting → Style → Camera → Lighting → Constraints.
- The exact backdrop hex is restated in every keyframe prompt.
- No text in frame anywhere, unless it's printing on a physical prop.
- Morph prompt: only house motion verbs; "single continuous take, no cuts"; the
  change is continuous across the full duration; closes with the audio line.
- `aspectRatio: "9:16"` on every image and video call.
- `duration` is one of 4/5/6/8/10/12/15, identical in the plan and the call.
- At a **match-cut** boundary, segment *i+1*'s `image` is byte-identical to
  segment *i*'s `lastImage`. At a **scene-cut** boundary, segment *i+1* uses its
  own fresh keyframe (and that's intentional).

## Checklist — before delivering

- Script hit its template's word budgets; VO is one take; no intro, no outro, no CTA.
- Picture duration matches VO duration within ~0.3s.
- No cut appears *inside* any single clip — verify with
  `ffmpeg select='gt(scene,0.3)'`. Detections at segment *boundaries* are
  expected for scene cuts and fine; a detection *mid-segment* is a failure.
- Exactly 2 stings, on the beats the format profile marks.
- Bed is continuous, never silent; ducked under VO.
- Final master ≈ −14 dB RMS, peak < 0.98.
- No burned-in text; full-bleed 9:16.

---

## Reference files

- `references/house-style.md` — the twelve slots, locked. The medium sentence,
  the AI-default negative, the exact palette, the motion verbs, the prompt
  templates. Read before any prompt.
- `references/pipeline.md` — exact Unsora calls (`create_image`, `create_video`
  with the `image`/`lastImage` chain, `create_voiceover`, `create_music`, the
  `wait_for_*` pollers), the ffmpeg concat and mix, the −14 dB master, the
  manifest schema, resume-after-failure.
- `references/formats/hypothetical.md` — the 8-beat "what if you did this to a
  body" template. ~117 words / ~43s.
- `references/formats/true-story.md` — the 5-beat "this happened to a person"
  template. ~94 words / ~33s.
- `references/formats/format-framework.md` — how to build a third format profile
  without breaking the house style.

