# Demo Video

> Assemble ordered frames or clips with narration into a narrated MP4 via FFmpeg

- Skill: `gabrielmoreira/demo-video` (Agent Skill)
- Install (CLI): `npx skillmds@latest add gabrielmoreira/demo-video`
- Raw SKILL.md: https://api.skillmd.com/api/skills/gabrielmoreira/demo-video/raw
- Safety review: pending (external: skill-scanner PASS, skillspector CAUTION)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Coding & Dev Tools
- License: MIT
- Author: gabrielmoreira (https://skillmd.com/u/gabrielmoreira)
- Updated: 2026-09-21
- Page: https://skillmd.com/skills/gabrielmoreira/demo-video

---


# Demo Video Assembly Skill

This skill assembles a narrated demo video from ordered visual segments and matching narration audio. It is designed for first-pass walkthrough videos that combine captured prototype frames or clips with per-segment voiceover WAV files.

## Overview

The workflow takes a manifest that describes each segment, resolves the visual source, and uses FFmpeg to render each segment into a normalized video clip before concatenating them into a final MP4. The narration track is muxed from WAV files so the output can be reviewed as a polished walkthrough without requiring a separate video-editing tool.

## Manifest Schema

Use a `segments.yml` manifest with optional top-level output settings and an ordered list of segments. Each entry describes a visual source and the narration audio to combine for that portion of the video. All paths resolve relative to the manifest file.

```yaml
output: ./output/demo.mp4   # optional; destination path for the assembled MP4
resolution: 1280x720        # optional; default 1280x720
fps: 24                     # optional; default 24
segments:
  - type: frame
    visual: ./frames/intro.png
    narration: ./audio/intro.wav
    duration: 4.5
  - type: clip
    clip: ./clips/interaction.mp4
    narration: ./audio/interaction.wav
```

### Top-level fields

* `output` sets the destination path for the assembled MP4, resolved relative to the manifest; the `--output` or `-OutputPath` argument overrides it when supplied
* `resolution` controls the output width and height in `WIDTHxHEIGHT` form (default `1280x720`); the `--resolution` or `-Resolution` argument overrides it
* `fps` sets the frame rate applied when rendering each segment (default `24`); the `--fps` or `-Fps` argument overrides it

### Segment fields

* `type` identifies whether the segment is a still image (`frame`) or a motion clip (`clip`)
* `visual` points to an image file for a frame segment
* `clip` points to a motion clip file for a clip segment
* `narration` points to the WAV file generated from narration text (the script also accepts `narration_wav` as an alias)
* `duration` is optional and overrides the inferred duration when you want a fixed segment length

## Quick Start

Use the bash or PowerShell wrappers to invoke the assembler from the skill directory.

```bash
scripts/assemble-video.sh --manifest examples/segments.yml --output ./output/demo.mp4
```

```powershell
scripts/Invoke-AssembleVideo.ps1 -ManifestPath examples/segments.yml -OutputPath ./output/demo.mp4
```

## Parameters Reference

The assembly step accepts the following high-level controls:

* `--manifest` or `-ManifestPath` selects the YAML manifest to process
* `--output` or `-OutputPath` sets the destination MP4 path
* `--fps` or `-Fps` controls the output frame rate for rendered segments
* `--resolution` or `-Resolution` controls the output width and height in the form `WIDTHxHEIGHT`
* `duration` per segment lets you override the inferred length when narration timing is known in advance

## Narration Quality

Narration quality is the single biggest driver of how polished the final video feels. Prioritize neural voices from **Azure AI Speech (part of Azure AI Foundry)** through the `tts-voiceover` skill for any video you intend to share.

* **Recommended:** Use the `tts-voiceover` skill backed by Azure AI Speech neural voices (for example `en-US-Andrew:DragonHDLatestNeural` or `en-US-Jenny:DragonHDLatestNeural`). These produce natural, presentation-grade narration and are the default for shareable output.
* **Fallback only:** Offline open-source engines such as `espeak-ng` require no credentials but sound noticeably robotic. Treat them as a no-network smoke-test fallback, not a delivery format. Regenerate narration with Azure AI Speech before publishing.

See the `tts-voiceover` skill for the neural voice catalog, `--voice` and `--rate` controls, and Azure authentication (Entra ID or key).

## Reuse Bridge

This skill is intentionally designed to fit into the existing media workflow:

* `tts-voiceover` provides the narration WAV files that this skill muxes into the final output; prefer its Azure AI Speech neural voices for production-quality narration
* `vscode-playwright` provides the frame-capture source for prototype walkthroughs and screen-based demos

## Prerequisites

FFmpeg and ffprobe must be available on your PATH.

### Linux

```bash
sudo apt update && sudo apt install ffmpeg
```

### macOS

```bash
brew install ffmpeg
```

### Windows

```powershell
winget install FFmpeg.FFmpeg
```

