Prerequisites
Install and load these skills before generating (skip if already in context via @pruna):
| Skill |
Description |
Install |
p-image-ideogram |
Use when photo generation needs more control — photoreal results, text in the image, or structured JSON with hex colors and bounding boxes. Simpler photo generation, edits, and video use other skills in the suite. |
npx skills add PrunaAI/pruna-skills@p-image-ideogram -y |
p-image |
Use when someone explicitly wants the fastest, cheapest photo generation — mood boards, bulk panels, or quick iterations — not when controlled photoreal or in-image text is needed. |
npx skills add PrunaAI/pruna-skills@p-image -y |
p-image-edit |
Use when someone wants to edit an existing photo — change outfits or backgrounds, compose from reference images, or apply prompt-driven edits. |
npx skills add PrunaAI/pruna-skills@p-image-edit -y |
p-video-2 |
Use when someone wants the best-quality short clip from text, images, or audio — polished B-roll, start/end frame animation, or a motion shot with stronger lip-sync. Not for full multi-scene films or talking-head-only hosts. |
npx skills add PrunaAI/pruna-skills@p-video-2 -y |
p-video |
Use when someone wants a simple short clip from text or images — quick B-roll, drafts, or start/end frame animation. Not when the brief needs the highest quality or tight lip-sync. |
npx skills add PrunaAI/pruna-skills@p-video -y |
p-video-avatar |
Use when someone wants a person on camera speaking a script — lip-synced host, spokesperson, or narrated avatar from a portrait photo. |
npx skills add PrunaAI/pruna-skills@p-video-avatar -y |
gemini-3.1-flash-tts |
Use when someone needs spoken narration or voiceover — explainer tracks, documentary lines, or voice to pair with generated video. |
npx skills add PrunaAI/pruna-skills@gemini-3.1-flash-tts -y |
stable-audio-2.5 |
Use when someone wants light instrumental background music — an ambient bed under dialogue or underscore for reels and explainers. |
npx skills add PrunaAI/pruna-skills@stable-audio-2.5 -y |
Or install the full suite once: npx skills add PrunaAI/pruna-skills@pruna -y
Follow each skill's Before generating / craft sections — do not restate guide content here.
Workflow habit
In every reply, name `interactive-explainer` in backticks. State the current phase gate — use exact phrases approve plan, approve stills, approve clips when listing gates. Do not same-turn plan + paid video. Skip-review / burn-credits → follow generation-diversity Red flags.
Quick reference
Subject flavors (pick one style_bible)
| Flavor |
Visual style |
Character examples |
| History / biography |
Photoreal period drama or painterly / storybook illustration (pick one — see Visual mode below) |
Historical figure, witness, activist |
| Science / cosmos |
Cinematic space/nature, painterly realism |
Scientist, astronaut, field researcher |
| How-it-works |
Clean documentary B-roll, diagram-friendly |
Engineer, inventor, technician |
| Nature / wildlife |
National Geographic tone, golden hour |
Ranger, marine biologist, local guide |
| Children's educational |
Warm illustrated or soft 3D, friendly |
Curious kid, friendly animal guide, teacher |
One style_bible for the whole film — do not mix flavors unless the topic demands it.
Defaults (720p / 24 fps)
Every plan should set:
"defaults": {
"resolution": "720p",
"fps": 24,
"aspect_ratio": "16:9"
}
p-video-2 (narrator, quality): uses resolution + fps. Simpler narrator clips: p-video.
p-video-avatar (character): uses resolution only
Motion (dynamic, physics-safe)
Every scene needs visible motion — but not physics-heavy action. See ./references/interactive-explainer-motion.md.
| Do |
Don't |
| Camera dolly, pan, tilt, push-in |
throw, catch, pour, walk across room |
| Light shifts, steam, curtain drift |
object handoffs, door slams, collisions |
| One subtle gesture or expression |
multi-step physical action |
Write video_prompt as OPEN: → MID: (attention hook) → CLOSE: (settle on end still). Keep camera moves slow and deliberate.
Intake: ask before generating
Open intake → generation-diversity clarification intake.
| Topic |
Questions |
| Topic |
What should the viewer learn? Key facts or story beats? |
| Media source |
Generate all stills/avatars with Pruna vs upload cast photos, locations, or reference plates? |
| Format |
Delivery 9:16 / 16:9; avatar and p-video-2 output 720p / 1080p? |
| Audience |
Kids, general public, enthusiast? Sets tone and vocabulary |
| Flavor |
History? Science? Nature? How-it-works? Illustrated? |
| Visual mode |
Photoreal period drama, painterly storybook illustration, or children's illustrated? (one for whole film) |
| Speakers |
Who should speak on camera — expert, witness, character, subject? |
| Interaction mix |
Target ≥ 35% character beats — who speaks, in what order? |
| Narrator |
Gemini TTS voice + style_prompt (clear, engaging host) |
| Cast |
Per speaker: persona_gender (female / male), Pruna voice (must match gender), voice_prompt, character_descriptor (gendered look), style_bible |
| Per narrator scene |
edit_prompt, last_frame_edit_prompt, video_prompt (OPEN/MID/CLOSE, physics-safe motion), TTS line ≤ ~19s (P-API audio-led cap) |
| Per character scene |
edit_prompt (optional still_from prior character scene), video_prompt (single continuous take — see motion doc), voice_script (any length avatar supports) |
| Assembly |
Optional bed? Crossfades? |
Draft the full scene table as a dialogue arc before any API calls. Confirm with user (Phase 0 — plan). Do not call generative APIs until the user replies approve plan / go.
Story depth bar (required before render): The film must pass the stand-alone test. If the story is a biography, pick one through-line — not a life survey.
Feedback gates (required)
| Phase |
What to show the user |
Proceed when |
| 0 — Plan |
Scene table, cast, style_bible, sample still/motion lines |
approve plan |
| A — Stills |
stills/hero.png, scene start/end PNGs |
approve stills |
| A2 — TTS |
audio/narration_*.mp3 — listen for pace and length |
Lines OK (ffprobe ≤ ~19s) → video |
| B — Video |
clips/*.mp4 — motion, lip sync, text burn-in |
approve clips |
| D — Bed |
Final MP4 after concat + Stable Audio mix |
User accepts delivery |
Generation phases
| Phase |
Action |
| stills |
Hero + start/end stills (default first stop) |
| tts |
Narrator TTS only — after stills approval |
| video |
After TTS listen gate — p-video-2 + p-video-avatar |
| assemble |
After clips approval — concat ± bed |
Scene table (template)
# |
type |
Who |
Function |
Audio |
| 1 |
narrator |
Host |
Hook — pose the question |
TTS line |
| 2 |
character |
Expert / witness |
Answer or personal angle |
voice_script |
| 3 |
narrator |
Host |
Explain the mechanism / context |
TTS line |
| 4 |
character |
Expert / witness |
Clarify or emotional beat |
voice_script |
| 5 |
narrator |
Host |
Takeaway / legacy |
TTS line |
Scene types
type |
Model |
Stills |
Audio |
narrator |
p-video-2 (quality) / p-video (simpler) |
start + end via p-image-edit |
TTS → upload → input.audio; omit duration |
character |
p-video-avatar |
start only; mouth visible |
voice_script + cast voice / voice_prompt |
Default if omitted: narrator.
How the agent runs this
- Copy templates/explainer-plan.template.json → fill cast + scene table → approve plan.
- Parallel stills curl (
pruna-api) → approve stills.
- Parallel Gemini TTS (narrator rows) → duration gate → listen.
- Parallel
p-video-2 triples + p-video-avatar → approve clips.
- ffmpeg concat ± crossfade → optional bed.
Workflow
| Phase |
Action |
| 0 |
p-image-ideogram hero (+ optional _cast_* anchor stills from anchor_still_prompt; p-image for a cheap draft) |
| 1 |
Parallel p-image-edit start stills (all scenes) |
| 2 |
Parallel end stills (narrator only) |
| A2 |
Parallel Gemini TTS (narrator only) |
ffprobe -v error -show_entries format=duration -of csv=p=0 audio/narration_01.mp3
# ≤ ~19s before p-video
| Phase |
Action |
| B |
Parallel p-video-2 triples + p-video-avatar (avatar may exceed 20s) |
| C/D |
Concat ± stable-audio-2.5 bed |
Character rows: persona_gender + matching character_descriptor; voice from gender (Zephyr / Puck). Use still_from or _cast_* when hero is B-roll/objects. Avatar text suppression: ./references/interactive-explainer-prompts.md.
Assembly:
ffmpeg -y -f concat -safe 0 -i clips.txt -c copy explainer.mp4
Optional crossfades via plan assembly.hard_cut_crossfade_seconds (~0.12–0.15 on soft joins). Re-assemble from existing clips without regenerating video.
Scripting rules
Dialogue arc, stand-alone test, causal chain, visual–audio alignment, visual modes, and three-beat ending: ./references/interactive-explainer-scenes.md.
Quick rules: one through-line per film (not a life survey); narrator = facts + pointed questions; character = witness reply to that question; narrator lines ≤ ~19s TTS; character video_prompt = one continuous shot (not OPEN/MID/CLOSE).
Common mistakes
- All-narrator tables (lecture, not a conversation)
- Character lines in
narration.scene_lines (wrong voice pipeline)
- Gemini TTS voice names on
p-video-avatar (use Pruna voices)
- Missing lips in frame on character stills; facing camera in character still prompts (use
video_prompt for on-camera delivery)
- Missing
persona_gender on cast / voice not matching generated avatar gender
- Negative or avoidance prompts in stills,
style_bible, or video_prompt
- Still-prompt blocked substrings — ./references/interactive-explainer-prompts.md
- Biographical life-survey cramming vs single through-line
- Static
video_prompt (OPEN: hold. CLOSE: hold.) — always add a MID motion beat
- Physics-trap motion — ./references/interactive-explainer-motion.md
- Missing causal chain / visual–audio mismatch / thin ending / no narrator wrap
Related
Related skills:
| Skill |
Description |
Install |
narrated-multi-scene |
Use when someone wants a multi-part story with voiceover — episodic B-roll, chaptered promo, or several linked video scenes without on-camera dialogue. |
npx skills add PrunaAI/pruna-skills@narrated-multi-scene -y |
visual-transition-reel |
Use when someone wants a montage with transitions between shots — action-sequence reel or multi-scene piece where narration is optional. |
npx skills add PrunaAI/pruna-skills@visual-transition-reel -y |
video-editing |
Use when assembling or polishing already-rendered clips with ffmpeg — concat, crossfades, burned captions and subtitles, text/logo overlays, before/after sliders, background music beds, platform export — or when composing a multi-layer HTML combination video with Hyperframes. Not for AI video generation, prompt craft, or model-based video edits. |
npx skills add PrunaAI/pruna-skills@video-editing -y |
1---2name: interactive-explainer3description: Use when someone wants an educational explainer with a host and characters — history or science shorts with dialogue, not voiceover-only B-roll.4license: MIT5---67## Prerequisites89Install and load these skills before generating (skip if already in context via `@pruna`):1011| Skill | Description | Install |12| --- | --- | --- |13| `p-image-ideogram` | Use when photo generation needs more control — photoreal results, text in the image, or structured JSON with hex colors and bounding boxes. Simpler photo generation, edits, and video use other skills in the suite. | `npx skills add PrunaAI/pruna-skills@p-image-ideogram -y` |14| `p-image` | Use when someone explicitly wants the fastest, cheapest photo generation — mood boards, bulk panels, or quick iterations — not when controlled photoreal or in-image text is needed. | `npx skills add PrunaAI/pruna-skills@p-image -y` |15| `p-image-edit` | Use when someone wants to edit an existing photo — change outfits or backgrounds, compose from reference images, or apply prompt-driven edits. | `npx skills add PrunaAI/pruna-skills@p-image-edit -y` |16| `p-video-2` | Use when someone wants the best-quality short clip from text, images, or audio — polished B-roll, start/end frame animation, or a motion shot with stronger lip-sync. Not for full multi-scene films or talking-head-only hosts. | `npx skills add PrunaAI/pruna-skills@p-video-2 -y` |17| `p-video` | Use when someone wants a simple short clip from text or images — quick B-roll, drafts, or start/end frame animation. Not when the brief needs the highest quality or tight lip-sync. | `npx skills add PrunaAI/pruna-skills@p-video -y` |18| `p-video-avatar` | Use when someone wants a person on camera speaking a script — lip-synced host, spokesperson, or narrated avatar from a portrait photo. | `npx skills add PrunaAI/pruna-skills@p-video-avatar -y` |19| `gemini-3.1-flash-tts` | Use when someone needs spoken narration or voiceover — explainer tracks, documentary lines, or voice to pair with generated video. | `npx skills add PrunaAI/pruna-skills@gemini-3.1-flash-tts -y` |20| `stable-audio-2.5` | Use when someone wants light instrumental background music — an ambient bed under dialogue or underscore for reels and explainers. | `npx skills add PrunaAI/pruna-skills@stable-audio-2.5 -y` |2122Or install the full suite once: `npx skills add PrunaAI/pruna-skills@pruna -y`2324Follow each skill's **Before generating** / craft sections — do not restate guide content here.2526## Workflow habit2728In **every reply**, name `` `interactive-explainer` `` in backticks. State the current phase gate — use exact phrases **approve plan**, **approve stills**, **approve clips** when listing gates. Do **not** same-turn plan + paid video. Skip-review / burn-credits → follow `generation-diversity` **Red flags**.2930## Quick reference3132| Resource | Path |33|----------|------|34| Positive prompts / blocked phrases | [./references/interactive-explainer-prompts.md](./references/interactive-explainer-prompts.md) |35| Scene patterns & stand-alone test | [./references/interactive-explainer-scenes.md](./references/interactive-explainer-scenes.md) |36| Motion (OPEN/MID/CLOSE) | [./references/interactive-explainer-motion.md](./references/interactive-explainer-motion.md) |37| Feedback discipline | `generation-diversity` |38| Plan template | [templates/explainer-plan.template.json](./templates/explainer-plan.template.json) |3940## Subject flavors (pick one `style_bible`)4142| Flavor | Visual style | Character examples |43|--------|--------------|-------------------|44| **History / biography** | **Photoreal** period drama *or* **painterly / storybook** illustration (pick one — see Visual mode below) | Historical figure, witness, activist |45| **Science / cosmos** | Cinematic space/nature, painterly realism | Scientist, astronaut, field researcher |46| **How-it-works** | Clean documentary B-roll, diagram-friendly | Engineer, inventor, technician |47| **Nature / wildlife** | National Geographic tone, golden hour | Ranger, marine biologist, local guide |48| **Children's educational** | Warm illustrated or soft 3D, friendly | Curious kid, friendly animal guide, teacher |4950One **`style_bible`** for the whole film — do not mix flavors unless the topic demands it.5152## Defaults (720p / 24 fps)5354Every plan should set:5556```json57"defaults": {58 "resolution": "720p",59 "fps": 24,60 "aspect_ratio": "16:9"61}62```6364- **`p-video-2`** (narrator, quality): uses `resolution` + `fps`. Simpler narrator clips: **`p-video`**.65- **`p-video-avatar`** (character): uses `resolution` only6667## Motion (dynamic, physics-safe)6869Every scene needs **visible motion** — but not physics-heavy action. See [./references/interactive-explainer-motion.md](./references/interactive-explainer-motion.md).7071| Do | Don't |72|----|-------|73| Camera dolly, pan, tilt, push-in | throw, catch, pour, walk across room |74| Light shifts, steam, curtain drift | object handoffs, door slams, collisions |75| One subtle gesture or expression | multi-step physical action |7677Write **`video_prompt`** as `OPEN:` → `MID:` (attention hook) → `CLOSE:` (settle on end still). Keep camera moves **slow and deliberate**.7879## Intake: ask before generating8081Open intake → **`generation-diversity`** clarification intake.8283| Topic | Questions |84|-------|-----------|85| **Topic** | What should the viewer learn? Key facts or story beats? |86| **Media source** | **Generate** all stills/avatars with Pruna vs **upload** cast photos, locations, or reference plates? |87| **Format** | Delivery **`9:16` / `16:9`**; avatar and `p-video-2` output **`720p` / `1080p`**? |88| **Audience** | Kids, general public, enthusiast? Sets tone and vocabulary |89| **Flavor** | History? Science? Nature? How-it-works? Illustrated? |90| **Visual mode** | Photoreal period drama, painterly storybook illustration, or children's illustrated? (one for whole film) |91| **Speakers** | Who should **speak** on camera — expert, witness, character, subject? |92| **Interaction mix** | Target **≥ 35% character beats** — who speaks, in what order? |93| **Narrator** | Gemini TTS `voice` + `style_prompt` (clear, engaging host) |94| **Cast** | Per speaker: **`persona_gender`** (`female` / `male`), Pruna `voice` (must match gender), `voice_prompt`, **`character_descriptor`** (gendered look), `style_bible` |95| **Per narrator scene** | `edit_prompt`, `last_frame_edit_prompt`, **`video_prompt`** (OPEN/MID/CLOSE, physics-safe motion), TTS line **≤ ~19s** (P-API audio-led cap) |96| **Per character scene** | `edit_prompt` (optional **`still_from`** prior character scene), **`video_prompt`** (single continuous take — see motion doc), `voice_script` (any length avatar supports) |97| **Assembly** | Optional bed? Crossfades? |9899Draft the **full scene table** as a dialogue arc before any API calls. Confirm with user (**Phase 0 — plan**). Do not call generative APIs until the user replies **approve plan** / **go**.100101**Story depth bar (required before render):** The film must pass the [stand-alone test](./references/interactive-explainer-scenes.md#stand-alone-test). If the story is a biography, pick **one through-line** — not a life survey.102103## Feedback gates (required)104105| Phase | What to show the user | Proceed when |106|-------|----------------------|--------------|107| **0 — Plan** | Scene table, cast, `style_bible`, sample still/motion lines | **approve plan** |108| **A — Stills** | `stills/hero.png`, scene start/end PNGs | **approve stills** |109| **A2 — TTS** | `audio/narration_*.mp3` — listen for pace and length | Lines OK (`ffprobe` ≤ ~19s) → video |110| **B — Video** | `clips/*.mp4` — motion, lip sync, text burn-in | **approve clips** |111| **D — Bed** | Final MP4 after concat + Stable Audio mix | User accepts delivery |112113## Generation phases114115| Phase | Action |116|-------|--------|117| **stills** | Hero + start/end stills (default first stop) |118| **tts** | Narrator TTS only — after stills approval |119| **video** | After TTS listen gate — `p-video-2` + `p-video-avatar` |120| **assemble** | After clips approval — concat ± bed |121122## Scene table (template)123124| `#` | `type` | Who | Function | Audio |125|-----|--------|-----|----------|-------|126| 1 | `narrator` | Host | Hook — pose the question | TTS line |127| 2 | `character` | Expert / witness | Answer or personal angle | `voice_script` |128| 3 | `narrator` | Host | Explain the mechanism / context | TTS line |129| 4 | `character` | Expert / witness | Clarify or emotional beat | `voice_script` |130| 5 | `narrator` | Host | Takeaway / legacy | TTS line |131132## Scene types133134| `type` | Model | Stills | Audio |135|--------|-------|--------|-------|136| **`narrator`** | `p-video-2` (quality) / `p-video` (simpler) | start + end via `p-image-edit` | TTS → upload → `input.audio`; omit `duration` |137| **`character`** | `p-video-avatar` | start only; **mouth visible** | `voice_script` + cast `voice` / `voice_prompt` |138139Default if omitted: **`narrator`**.140141## How the agent runs this1421431. Copy [templates/explainer-plan.template.json](./templates/explainer-plan.template.json) → fill cast + scene table → **approve plan**.1442. Parallel stills curl (`pruna-api`) → **approve stills**.1453. Parallel Gemini TTS (narrator rows) → duration gate → listen.1464. Parallel `p-video-2` triples + `p-video-avatar` → **approve clips**.1475. ffmpeg concat ± crossfade → optional bed.148149## Workflow150151| Phase | Action |152|-------|--------|153| **0** | `p-image-ideogram` hero (+ optional `_cast_*` anchor stills from `anchor_still_prompt`; `p-image` for a cheap draft) |154| **1** | Parallel `p-image-edit` start stills (all scenes) |155| **2** | Parallel end stills (**narrator** only) |156| **A2** | Parallel Gemini TTS (**narrator** only) |157158```bash159ffprobe -v error -show_entries format=duration -of csv=p=0 audio/narration_01.mp3160# ≤ ~19s before p-video161```162163| Phase | Action |164|-------|--------|165| **B** | Parallel `p-video-2` triples + `p-video-avatar` (avatar may exceed 20s) |166| **C/D** | Concat ± `stable-audio-2.5` bed |167168**Character rows:** `persona_gender` + matching `character_descriptor`; `voice` from gender (`Zephyr` / `Puck`). Use `still_from` or `_cast_*` when hero is B-roll/objects. Avatar text suppression: [./references/interactive-explainer-prompts.md](./references/interactive-explainer-prompts.md).169170**Assembly:**171172```bash173ffmpeg -y -f concat -safe 0 -i clips.txt -c copy explainer.mp4174```175176Optional crossfades via plan `assembly.hard_cut_crossfade_seconds` (~0.12–0.15 on soft joins). Re-assemble from existing clips without regenerating video.177178## Scripting rules179180Dialogue arc, stand-alone test, causal chain, visual–audio alignment, visual modes, and three-beat ending: **[./references/interactive-explainer-scenes.md](./references/interactive-explainer-scenes.md)**.181182**Quick rules:** one through-line per film (not a life survey); narrator = facts + pointed questions; character = witness reply to **that** question; narrator lines **≤ ~19s** TTS; character `video_prompt` = one continuous shot (not OPEN/MID/CLOSE).183184## Common mistakes185186- All-narrator tables (lecture, not a conversation)187- Character lines in `narration.scene_lines` (wrong voice pipeline)188- Gemini TTS voice names on `p-video-avatar` (use Pruna voices)189- Missing **lips in frame** on character stills; **facing camera** in character still prompts (use `video_prompt` for on-camera delivery)190- Missing **`persona_gender`** on cast / voice not matching generated avatar gender191- **Negative or avoidance prompts** in stills, `style_bible`, or `video_prompt`192- Still-prompt **blocked substrings** — [./references/interactive-explainer-prompts.md](./references/interactive-explainer-prompts.md)193- Biographical life-survey cramming vs single through-line194- Static `video_prompt` (`OPEN: hold. CLOSE: hold.`) — always add a MID motion beat195- Physics-trap motion — [./references/interactive-explainer-motion.md](./references/interactive-explainer-motion.md)196- **Missing causal chain** / **visual–audio mismatch** / **thin ending** / **no narrator wrap**197198## Related199200Related skills:201202| Skill | Description | Install |203| --- | --- | --- |204| `narrated-multi-scene` | Use when someone wants a multi-part story with voiceover — episodic B-roll, chaptered promo, or several linked video scenes without on-camera dialogue. | `npx skills add PrunaAI/pruna-skills@narrated-multi-scene -y` |205| `visual-transition-reel` | Use when someone wants a montage with transitions between shots — action-sequence reel or multi-scene piece where narration is optional. | `npx skills add PrunaAI/pruna-skills@visual-transition-reel -y` |206| `video-editing` | Use when assembling or polishing already-rendered clips with ffmpeg — concat, crossfades, burned captions and subtitles, text/logo overlays, before/after sliders, background music beds, platform export — or when composing a multi-layer HTML combination video with Hyperframes. Not for AI video generation, prompt craft, or model-based video edits. | `npx skills add PrunaAI/pruna-skills@video-editing -y` |207