# Clayline Ads

> Produce a claymation-style parallel-storyline video ad end-to-end (brief -> plan -> style lock -> keyframe stills -> transition video clips -> VO/music -> ffmpeg assembly), phase-gated for user approval at each stage. Generates all visuals via the Higgsfield MCP server.

- Skill: `likeitdigital/clayline-ads` (Agent Skill, multi-file: 3 files)
- Install (CLI): `npx skillmds@latest add likeitdigital/clayline-ads`
- Raw SKILL.md: https://api.skillmd.com/api/skills/likeitdigital/clayline-ads/raw
- Safety review: pending
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: AI & ML
- Author: likeitdigital (https://skillmd.com/u/likeitdigital)
- Updated: 2026-09-17
- Page: https://skillmd.com/skills/likeitdigital/clayline-ads

---


# Clayline Ads

A phase-gated pipeline for turning a rough story idea into a finished, vertical (9:16), claymation-style video ad — entirely through the Higgsfield MCP server's image/video/audio generation tools, with a checkpoint after every stage.

## Requirements

- `mcp__Higgsfield__*` MCP tools (image, video, audio generation; media upload/confirm)
- `ffmpeg` and `imagemagick` on the local shell (for the contact sheet and final assembly)

## Phase 1 — Brief

Capture: product/thing being advertised, the one-sentence angle, target audience, and a raw story summary in the user's own words. Structure the story beats under the heading **"Parallel Storyline"** — this is the user's own framework, not an external methodology; do not attribute it to any third-party name. Use whatever phase labels fit the specific story naturally (setup / turning point / resolution, or the user's own terms) rather than a fixed borrowed vocabulary.

Present the brief and wait for explicit approval before continuing.

## Phase 2 — Plan

Lock, in writing, before any generation:
- **Image model** (confirm which model + reference-media role naming with the user if unsure; different models use different `medias[].role` values for reference images).
- **Cast**: every recurring character, each with a *verbatim* description clause that will be repeated word-for-word in every prompt where that character appears. Note absolute or relative scale for any non-human-sized character (animal, object, etc.) explicitly, so the model doesn't invent scale.
- **Product/prop(s)**: same verbatim-clause treatment for any recurring product or important prop (not just the "hero" product — a bag, a costume, anything that must stay visually consistent scene to scene).
- **Visual world / style lock**: material (e.g. claymation/plasticine), palette, lighting, a mandatory "richly populated set" clause so backgrounds never read as empty, character-grammar notes (eyes, hands, mouths), and negative constraints (no photoreal, no real logos, etc).
- **Title atoms**: any on-screen caption text for Phase 6.
- **Script + Parallel Storyline table**: one row per planned still, mapping a VO/caption line (if any) to the story beat and the visual action + camera direction for that still.

Present the plan and wait for explicit approval before continuing.

## Phase 3 — Style lock + character masters

Generate the style-lock plate and one master image per recurring character/prop, directly from the locked text descriptions in the plan. Every visual reference the pipeline uses from here on must be a generation this pipeline itself produced — do not source or reuse images from outside the project.

Present all masters together and wait for explicit approval before continuing.

## Phase 4 — Keyframe stills, in order, then self-check + contact sheet

Generate stills **in story order**, each one referencing the immediately preceding still plus the locked cast/prop masters (so drift doesn't compound).

**Camera-perspective guidance for multi-location or multi-beat scenes:** when a scene's action can't be told from one flat wide shot without looking illogical (e.g. two characters ending up at different tables in the same scene), cut it into two shots, use a POV shot, or use an over-the-shoulder / shallow-depth-of-field composition — and add explicit negative clauses ("no other people, no extra background tables/characters") so strangers don't leak in from style-reference drift.

**Mandatory consistency self-check + contact sheet, after all stills are generated and before presenting them:** re-open and compare every still against the locked cast/prop/product descriptions — check color, shape and scale of recurring props, and character scale relative to each other. Fix any drift found by regenerating that still with an explicit corrective clause before moving on. Then build a contact sheet:

```bash
montage scene-01.png scene-02.png ... scene-NN.png \
  -tile 4x2 -geometry 400x711+5+5 -label '%f' \
  -background '#222' -fill white -pointsize 20 \
  contact-sheet.png
```

Send the contact sheet and wait for explicit approval before continuing.

## Phase 5a — Transition video clips

**Generate ONE clip at a time and wait for explicit user approval before generating the next one.** Do not submit every transition clip as one parallel batch by default — a person reviewing a multi-beat ad wants to catch drift or errors per-clip, early, not after the whole sequence has rendered. Only batch multiple clips in parallel if the user explicitly asks for that.

**Each clip must be a true keyframe-to-keyframe transition between two directly adjacent stills** (still N → still N+1) — never a clip that spans multiple stills or multiple story beats at once. A clip that has to bridge several unseen intermediate actions forces the video model to invent large chunks of motion, which is where visible errors come from. Prefer several short clips (roughly 3–6s each), each with exactly one clearly scoped action, over one long multi-beat clip.

Per-clip prompt discipline:
- State the single action only — nothing else happens in this clip.
- Repeat the locked character/prop verbatim clauses that are relevant to this clip.
- Add explicit negative/continuity constraints wherever they matter (e.g. "never sitting, always standing on the ground"; "static shot, no camera movement, no POV/perspective change" — a POV framing that exists in a still is a still-only device, video clips stay on one consistent camera angle unless the story explicitly calls for a cut).
- **Anatomical/physical logic check before submitting any prompt:** if a character's mouth/beak/hand is occupied holding an object, it cannot simultaneously do something else that needs the same body part (e.g. a bird carrying a bag in its beak cannot also grin with an open beak — express emotion through eyes, posture, or another limb/wing instead).
- If a costume or prop is introduced for comedic/narrative effect, explicitly track across the clip prompts when it appears, persists, and disappears/dissolves — don't let it silently vanish between clips with no in-story reason.

If a generation call returns an unrelated preset recommendation instead of submitting the job, decline it (pass the returned preset id back as `declined_preset_id`) and resubmit immediately — no need to check with the user first, just note that it happened.

## Phase 5b — Audio (runs alongside 5a — video and audio are prepared in parallel within the same approval stage)

Higgsfield's connected audio tools generate speech (TTS) only — there is no standalone instrumental/music or sound-effects model available for general use. Tell the user this plainly the first time music or SFX comes up, and offer these options:

1. **Voice-over narration via TTS** — pick a locale-appropriate preset voice for the script's language.
2. **User-sourced music/SFX** — the user gets their own track (free royalty-free libraries, or their own ElevenLabs/other account that does support music generation) and sends the finished audio file back for muxing into the Phase 6 assembly (see Phase 5a for the video-side workflow). Never fetch or reuse audio or visual assets from unnamed third-party sources on the user's behalf — only assets the user explicitly supplies or has licensed themselves.
3. **Silent / VO-only delivery** — the user adds music themselves in their own editor afterward.

When recommending a music style, match the mood *curve* the story needs (e.g. quiet/melancholic opening → playful middle → warm/triumphant ending), not just a genre — a track that's already "carefree" throughout undersells a quiet opening beat.

When the user's own TTS tool needs a script, hand them a punctuation-clean, copy-paste-ready block — proper commas/em-dashes/ellipses control pacing and emphasis in most TTS engines.

## Phase 6 — Assembly

Concatenate the approved clips with `ffmpeg`, mux in the final audio (VO and/or user-supplied music, loudness-normalized), burn in any caption text from the locked title atoms, and deliver the final video file.

## Failure handling

- **Prop/character drift** caught at the Phase 4 self-check or later: fix with a targeted regeneration carrying an explicit corrective clause, not a full restart.
- **A camera-angle problem is a cinematography fix, not a story rewrite** — before changing the plot to solve a continuity issue, try re-shooting the beat as two shots, a POV, or an over-the-shoulder composition first.
- **Rate-limited or failed generation calls**: resubmit the identical request once; if it fails again, surface the error to the user rather than silently changing the prompt.

