# AI Video Generation

> AI video generation: prompt engineering, model selection, and platform integration (Higgsfield, Flux 3, Runway, etc.). Six-part prompt structure, camera/lighting vocabulary, model routing, MCP automation, Flux 3 reasoning-harness techniques.

- Skill: `theheavenlyd3mon/ai-video-generation` (Agent Skill, multi-file: 6 files)
- Install (CLI): `npx skillmds add theheavenlyd3mon/ai-video-generation`
- Raw SKILL.md: https://api.skillmd.com/api/skills/theheavenlyd3mon/ai-video-generation/raw
- Safety review: pending
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: AI & ML
- Author: theheavenlyd3mon (https://skillmd.com/u/theheavenlyd3mon)
- Updated: 2026-09-09
- Page: https://skillmd.com/skills/theheavenlyd3mon/ai-video-generation

---


# AI Video Generation

## When to load
Any task involving AI-generated video: prompting, model selection, scene construction, platform setup, or batch video workflows.

## Critical rule: Don't assume branding
Unless the user explicitly names a brand/project (e.g. "Cheonma shoot"), treat video generation as **general creative work**. Do NOT auto-inject brand triggers, palette DNA, or style tokens. Ask if unclear.

## The Six-Part Prompt Structure

Production-grade AI video prompts have six components **in order**. 60–120 words total. Structure beats length — a structured 60-word prompt outperforms a 200-word stream of consciousness.

```
[Subject] + [Action] + [Camera] + [Lighting] + [Environment] + [Style]
```

### 1. Subject (1–4 concrete nouns, no filler adjectives)
### 2. Action (one verb phrase with direction/speed)
### 3. Camera (cinematic terms only — see vocabulary below)
### 4. Lighting (name source + quality)
### 5. Environment (location, time, weather, atmosphere)
### 6. Style (aesthetic anchors, quality tags)

## Camera Vocabulary (use these, not generic verbs)

| Shot types | Movements |
|---|---|
| Wide establishing shot | Dolly in / out |
| Medium two-shot | Handheld tracking |
| Over the shoulder | Crane up / down |
| Dutch tilt | Whip pan |
| Close-up / extreme close-up | Slow push-in |
| Extreme low angle | Arc shot / orbital |
| POV | Rack focus |

**Never say** "zoom" or "pan" alone — they produce generic output.

## Lighting Vocabulary

Name the **source** and **quality**:
- `golden hour backlight with rim light on subject`
- `single hard source from above, deep contrast`
- `overcast soft key from camera left`
- `neon practicals with blue shadow fill`
- `volumetric god rays through [object]`

**Never say** "good lighting" or "cinematic lighting" alone.

## Timing & Pacing (for models that accept duration cues)
- `slow sustained movement over the full duration`
- `quick three-beat action, hold on final frame`
- `continuous left-to-right drift across four seconds`
- One motion rhythm per shot. Never stack contradictory timing.

## Negative Prompting Rule
Most video models **ignore or backfire** on negative prompts. Instead of "no blur," write `tack sharp focus throughout`. Describe the positive state you want.

## Multi-Shot Sequences
For 3–6 connected shots, structure as:
```
Multi-shot cinematic sequence.

Shot 1: [shot type]. [style]. [subject + action]. [lighting]. [camera move].
Shot 2: ...
Shot N: ... [Freeze frame / hold on final beat.]
```
Tag character references with @image handles when the platform supports it.

## Model Routing (see references/higgsfield-platform.md for full details)

| Need | Model | Why |
|---|---|---|
| Character consistency + 4K | Kling 3.0 | Cheapest, multi-shot storyboard, Voice Binding |
| Multi-shot + native audio | Seedance 2.0 | Audio-video in one pass, 12 ref inputs |
| Atmospheric/outdoor/wide | Veo 3.1 | Global illumination, weather, depth |
| Restyle existing footage | WAN 2.7 | Video-reference style transfer |
| Fast drafts / iteration | MiniMax Hailuo 2.3 / Kling 2.5 Turbo | Low cost, minimal prompting needed |
| Multimodal ref + native audio | MiniMax H3 | 9 imgs + 3 video + 3 audio refs, 2K, five-block prompt, synced stereo audio — see references/minimax-h3.md |
| Physics / destruction | Sora 2 | Object permanence (API sunsetting Sept 2026) |
| In-Hermes native (no external API) | FLUX 3 (bfl_flux3_*) | Reasoning-harness prompting, native audio, multi-shot, continuation chaining |

**Workflow pattern**: Draft cheap → winnow → re-render winners on premium model.

## FLUX 3 (Hermes-native) — Prompting & Techniques

Flux 3 uses a reasoning harness, not a tag encoder. Plain prose beats keyword soup. The harness expands your prompt, so don't restyle it yourself — that stacks a second rewrite and drifts intent. See `references/flux3-techniques.md` for the full technique bank.

### Key differences from tag-based models
- **No keyword tricks, no word order games.** Write what you want to see in plain language.
- **Audio is generated by default.** Name ambient sound, music, and speech as separate layers. Say "no music" when unwanted.
- **Multi-shot works in one generation** but consecutive shots MUST contrast in scale, location, or color or the cut won't read — near-identical coverage blends into a continuous take.
- **Quoted text becomes speech only if a speaker is visible.** Without one, it renders as burned-in text. Always add "no on-screen text, no subtitles" when you want clean frames.
- **720p output.** Mood and composition carry more weight than fine detail. Atmospheric pieces shine; don't rely on legible text or tiny details.

### Tool selection
| Scenario | Tool |
|---|---|
| No input media | `bfl_flux3_text_to_video` |
| Animate one image as frame 0 | `bfl_flux3_image_to_video` |
| Pin 1-10 images at frame positions | `bfl_flux3_keyframes_to_video` |
| Continue from a clip's final frames | `bfl_flux3_video_continuation` |

### Character consistency across a multi-clip sequence (read before >4 clips)
Text-prompting identity drifts past ~4 clips — a locked style prefix won't hold a
character's look. The fix is a PIXEL anchor, not better words:
- **Generate ONE clean "character sheet" reference still** (image_generate):
  front-facing, neutral static pose, even soft light, full kit visible, plain
  background. Feed it as the frame-0 / reference image for every hero shot so the
  character's *actual pixels* seed each clip, not a text description. A turnaround
  sheet built from a generated reference instead of pulled frames.
- **No visible face is an ADVANTAGE for AI consistency.** A helmeted/goggled rider
  has no face to get wrong — identity is carried by the KIT (jacket color, pants,
  helmet, board graphic). Lock the kit and it's drift-proof. Don't pull anchors
  from existing action footage: it's small, motion-blurred, and wardrobe reads
  differently clip-to-clip (a crimson jacket reading coral in another shot IS the
  drift already happening).
- **B-roll is the consistency loophole.** Gear macros, hands, snow, environment,
  and POV/GoPro shots need no identity at all. A ~25–40% B-roll ratio rests the
  eye and hides drift in hero shots. POV is a consistency gimme — high energy,
  zero identity required.
- **One saturated wardrobe accent** (red jacket vs blue-white snow) reads as "same
  rider" even across scale/angle changes — the cheapest identity anchor there is.
- **Anthology casting turns drift into a feature.** For a commercial/ad that reads
  as "a production with a bunch of people" (sports brand film, crew montage), cast
  4–5 DISTINCT riders with distinct kits (one saturated accent each: red deck, blue
  deck, yellow deck…) instead of fighting to hold ONE face across 20 clips. Now any
  inter-clip variation reads as intentional casting, not drift — the hardest problem
  in long AI-video sequences dissolves. Still lock each rider's kit + a recurring
  product hero (deck graphic / shoe) so the world feels unified. Best paired with
  ~30% B-roll + POV (zero-identity shots) to rest the eye between hero riders.

### Action-sports trick shots are the hardest category (learned 2026-07, skate film)
On a golden-hour skate ad, the score split was stark: B-roll / establishing /
environment shots (no identity, no board-rider interaction) consistently scored
6–8/10, but **hero trick shots — a believable human ON a board doing a specific
trick — scored 2–4/10.** The model scrambles hardware + anatomy + physics at
once: the deck vanishes or shows the wrong face (underside instead of grip
tape), trucks/wheels melt or invert, the rider "floats" in a pose matching no
real trick, an arm goes missing. Even image-to-video from a clean starting-pose
still (the usual identity fix) did NOT save it — the board-rider contact is the
failure, not identity.
**Implications for planning an action-sports film:**
- Don't promise clean hero tricks. Budget the edit to *imply* tricks (fast 1–2s
  cuts where morphing reads as motion blur, B-roll of the spot, POV, environment,
  reaction) rather than showing them held and clean.
- Yellow decks drift to lime/olive-green on EVERY text-to-video gen, and
  backlighting pushes them to dark olive. Lock a saturated deck color ONLY via
  image-to-video from a clean reference frame; never trust a text prompt for it.
- POV shots fail as a category (board geometry + speed cues + obstacle legibility
  all break together). Use sparingly as flashes, not heroes.
- Golden-hour lighting and unified grade are the EASY, reliable win — lean on
  mood/cut rhythm over trick fidelity.

### Wardrobe/color drift (happens on EVERY independent generation)
Text-to-video re-rolls wardrobe on every separate generation — even with an
identical style prefix, a "orange-red jacket" prompt produced olive-green sleeves
(POV shot) and a dark cuff (hand macro) in 2 of 8 clips. The style prefix helps
with world consistency but does NOT lock garment color across independent gens.
**Fixes, in order of reliability:**
1. **Image-to-video from a clean reference frame** — the jacket's actual pixels
   seed the generation. Most reliable. Pull the frame from footage where the
   garment is clearly visible in neutral light (not backlit/silhouetted).
2. **Reinforce with a vivid analogy + prohibitions** — "vivid, saturated
   orange-red — the same bright orange-red as a traffic cone" plus "The jacket
   MUST be vivid orange-red throughout, never green, never dark, never black,
   never any other color." Worked on regeneration for both drifted clips.
3. **Make the color a foreground element** — for shots where the garment is
   barely visible (POV sleeves at frame edge, a hand macro cuff), explicitly
   name it as prominent/visible in the framing description so the model can't
   ignore it.

### Saturated prop/product color drift (worse than wardrobe)
Small saturated objects — skateboard decks, shoes, product hero items — drift
MORE than wardrobe in text-to-video, because they occupy fewer pixels and the
model re-rolls their color on every independent generation. **Yellow is the
worst offender**: a "vivid bright yellow" deck consistently rendered as
lime/chartreuse green-yellow or dark olive across multiple text-to-video gens
(4/10 and 3/10 QA scores). Backlighting compounds it catastrophically — a
yellow deck silhouetted against a sunset sky rendered as dark olive-green.
Even explicit prohibitions ("never green, never lime") failed in text-to-video.
**Fixes, in order of reliability:**
1. **Image-to-video from a clean reference frame** where the prop's actual
   pixels seed the generation. Generate the starting frame with image_generate
   (which handles yellow correctly), then animate it. This is the ONLY
   reliable fix for product-hero color accuracy.
2. **Front-fill the product** — never backlight a color-critical product shot.
   Silhouetting suppresses the surface color into shadow. Add "front fill
   light on the [product], not backlit" to the prompt.
3. **Reinforce with analogy + prohibitions** (same as wardrobe) — works for
   wardrobe but is NOT sufficient for small saturated props.

### POV shots (hardest category — 3-5/10 across all attempts)
First-person POV consistently fails on THREE axes simultaneously:
- **Board geometry**: the model shows the deck's underside graphic instead of
  grip tape (impossible while riding), or renders a floating board nub with no
  trucks, wheels, or feet.
- **Speed cues**: no motion blur, no foreground streak, no camera vibration —
  reads as static/parked, not rolling at speed.
- **Obstacle legibility**: the target (rail, ledge, bowl) is either missing,
  ambiguous (a lone vertical pole instead of a handrail), or blown out by the
  sun bloom sitting exactly where the obstacle should be.
**Practical guidance:** Accept POV as sub-second flashes (0.5–1s) in a fast
montage where the eye can't audit details. Do NOT use as held hero shots.
Budget for 2–3 re-rolls per POV clip. If the edit can survive with fewer POV
shots, cut them first — they're the lowest-yield category.

### Mechanical/hardware geometry (trucks, wheels, bearings)
AI video scrambles skateboard hardware anatomy:
- Trucks render inverted (baseplate grinding the rail instead of hanger/axle).
- Wheels go missing from axles, or warp oval, or multiply.
- Bearings render as "metallic soup" — no readable balls, seals, or races.
- Axle ends render as smooth domes instead of threaded hex nuts.
**Practical guidance:** Hardware macros (bearing spins, truck grinds, wheel
close-ups) are gorgeous in mood but mechanically incoherent. Use them as
sub-second texture flashes in a montage — never held shots. At 1s with grain
and motion, the eye registers "warm, shiny, mechanical" and doesn't audit
anatomy. At 3s+ or in slow-mo, the geometry collapses.

### Starting-pose stills for action shots (the image-to-video anchor technique)
For hero trick shots, DON'T use text-to-video. Instead:
1. **Generate a clean starting-pose still** with image_generate: the rider in
   the wind-up position, correct kit, correct deck color, golden-hour lighting,
   plain or simple background. This is the frame-0 pixel anchor.
2. **Feed it to image-to-video** as the opening frame. The rider's actual
   pixels (kit colors, deck color, body proportions) seed every subsequent
   frame — far more reliable than text description.
3. **Pose guidelines for the starting still:**
   - **Clear wheel-to-ground contact** — all four wheels visibly planted on a
     surface. Ambiguous contact (hovering trucks, floating board) produces
     snap/slide/float artifacts in the animation.
   - **All limbs visible** — no occluded arms. The video model must invent
     hidden limbs as the body turns, producing popping/duplicated arms.
   - **Wind-up pose, not spent pose** — a deep crouch that's already "committed"
     leaves no motion vector to animate. A standing wind-up (weight back, about
     to lean forward) gives the model a clear "lean-and-roll" direction.
   - **Lock contact in the video prompt** — explicitly state "wheels stay in
     contact with the concrete throughout" and "tail planted on the coping" to
     prevent the board from floating or snapping.
4. **QA the starting still** before committing to the video gen (which takes
   minutes). Check: correct kit colors, correct deck color, clear pose, no
   anatomy errors, good lighting. A 6/10 starting pose with explicit
   contact-locking language in the video prompt outperforms a 9/10 text-to-video
   gen for identity consistency.

### Aerial / spin / rotation morphing (and the held-pose fix)
Spinning or rotating tricks morph badly: the board changes length/width/shape
frame-to-frame, the body stretches and compresses, the grab hand melts into the
deck. Even with strong anti-morph constraints, a spinning aerial scored 3–4/10.
**Writing the prompt:** minimize rotation. Write a HELD POSE at the apex with
"nearly motionless," "the same held pose," "perfectly rigid, stable, and
consistent in shape and proportion throughout with absolutely no morphing,
warping, stretching, or distortion," and "very high shutter speed for crisp
frozen motion." This reduces but doesn't eliminate morphing.
**Salvage techniques for morphy footage:**
- **Slow it down** (ffmpeg `setpts=2.4*PTS`): at 2.4× with minimal rotation,
  artifacts become nearly invisible. A 5s clip fills 12s of screen time. QA
  confirmed "clean, no warping" after slowing — the strongest part of the edit.
- **Cut it fast** (1–2s): morphing reads as motion blur at speed. Fine for
  montage sections. Never use morphy footage as a clean slow-mo hero at 1×.

### Hard-won techniques
- **Phone/device screens at 720p**: Don't try to render readable text. Describe the screen as a visual object — "dark terminal interface, amber prompt line, blinking cursor, lines of monospace text scrolling." The audience's brain fills in "terminal" from the shape. Specifying actual text content produces garbled output or burned-in subtitles.
- **Hand-object interaction** (pulling phone from pocket, picking up items): This is the hardest thing for any video model. Workaround: start the shot with the object already in hand. "Hands holding a glowing phone" works; "hand reaches into pocket and retrieves phone" usually fails.
- **Narrative beats in multi-shot**: Give each shot a distinct emotional beat, not just a visual change. Approach → command → payoff. The contrast in intent makes cuts read cleaner than contrast in framing alone.
- **Duration pacing**: 15 seconds for three shots gives ~5s per beat. 10 seconds for three shots is too tight — beats feel rushed. Budget at least 4-5 seconds per shot for the action to breathe.
- **Continuation chaining**: Segments max at 15s. Plan each segment as a self-contained beat, then chain with `bfl_flux3_video_continuation`. Open the prompt with "Continue this video from its final frames:" and re-establish the subject.

### Workflow
1. Submit → get job id immediately (video doesn't exist yet)
2. Poll `bfl_flux3_get_result` with the id — generation takes several minutes, long "Generating" phase is normal
3. The poll call blocks while the job runs; if it returns still-generating, just call again — no sleeping
4. On Ready, the clip is downloaded and `saved_path` is returned. Deliver the file per platform conventions.
5. Job survives client restarts — re-poll the same id, never resubmit (duplicates spend budget)
6. **Rate limit: one submission per minute.** If you get a "limited to one attempt
   per minute" error, `sleep 45` then resubmit. When generating a batch serially,
   the natural poll-wait between clips usually covers this, but back-to-back
   submits (e.g. firing two clips in one turn) will trip it.
7. **Transient 504s happen.** Azure Front Door occasionally returns a 504 HTML
   page instead of a job id. Just wait the rate-limit cooldown and resubmit —
   the prompt wasn't consumed.
8. **`save_to` quirk**: if you pass a path that doesn't exist yet and has no
   extension (e.g. `.../new_clips`), the tool creates a FILE with that name
   instead of a directory. Always either pre-create the directory (`mkdir -p`)
   or pass a full file path with `.mp4` extension.

## Multi-Clip Production Pipeline (15–25 clips → short film / commercial)
For a long-form piece (a skate/snowboard/sports ad, a short film) assembled from
many independent clips, run these phases IN ORDER. Skipping ahead burns budget —
each FLUX 3 gen is minutes + rate-limited, so lock everything upstream first.

1. **Spec & storyboard FIRST.** Write an AD_SPEC.md: setting, mood, lighting,
   pacing, brand feel, target length. Then a SHOT_LIST.md with a shot taxonomy:
   establishing / B-roll-texture / POV / hero tricks / ride-away. Budget a
   ~25–35% B-roll ratio (zero-identity shots rest the eye and hide hero drift).
   Shoot MORE clips than you'll use (~24–26 for a 70s film) so the edit has
   options — cut down to the best ~18–20.
2. **Music is the spine — generate it BEFORE the clips.** Map the track's energy
   (see Music-Driven Assembly below) and lock a beat map (cold open over intro,
   action over sustained energy, slow-mo apex in the breakdown, hardest trick on
   the drop). The picture bends to the song, not vice versa.
3. **Cast the riders / characters.** For an anthology (multiple distinct people),
   generate ONE clean character-sheet reference still per rider (front-facing,
   neutral pose, even light, full kit, plain bg) — each with a distinct saturated
   accent (red deck, blue deck, yellow deck…). Anthology casting turns identity
   drift into intentional casting. See Character consistency section.
4. **Generate in phases** (optimized for the rate limit + QA flow):
   - Phase 1: establishing + B-roll (text-to-video, no identity — warms up the
     pipeline, cheapest to re-roll).
   - Phase 2: POV (text-to-video, zero identity — expect 3–5/10, budget re-rolls).
   - Phase 3: hero tricks (image-to-video from starting-pose stills — the
     payload, best identity consistency).
   - Phase 4: outro / ride-away.
   Run as a managed batch: fire one clip → poll to completion → QA → fire next.
   The natural poll-wait covers the rate limit. Report progress in chunks.
5. **QA every clip** (frame-grab + vision_analyze, see Reviewing below). Keep a
   re-roll queue for anything <6/10. Batch re-rolls at the end, using
   image-to-video where color/geometry matters.
6. **Assemble on the song's spine** (Music-Driven Assembly): trim to downbeats,
   unified grade on every clip, 15% Ken Burns punch-in, hard cuts, audio duck,
   fade. Deliver full film + teaser + individual clips + thumbnails + README in a
   numbered folder (01_raw_clips / 02_final_edits / 03_references / 04_audio /
   05_scripts / _scratch/review_proxies).

**Time reality:** 24–26 clips at the ~5-min account-wide rate limit + iteration
≈ 2–3 hours of generation. Don't promise a clip length the tools can't make in
one pass (>20s needs continuation chaining).

## Music-Driven Assembly (cut the picture to the song's spine)
When scoring a multi-clip sequence to a generated track, the music becomes the
timeline and the picture bends to it — cuts land on beats, the slow-mo hero sits
in a breakdown, the biggest action hits on the drop. Don't cut by feel and bolt
music on after.

**Where the track comes from (routing):**
| Scenario | Path |
|---|---|
| fal.ai balance available | `fal-ai/elevenlabs/music` via FAL_KEY — text-to-music, section-by-section composition_plan, ~$0.80/min (see fal-ai-generation skill) |
| fal.ai balance exhausted (403 "User is locked") | **Suno via web UI** — write the prompt with the `suno-music-creation` skill (DSL style prompt → Style field, structure tags → Lyrics field, Instrumental ON); user generates + returns the MP3. Suno returns a FULL song (~90s–4min): mine the cleanest breakdown→drop section. |
| Offline / local | HeartMuLa (CUDA GPU) or AudioCraft MusicGen (CPU/MPS) |
Probe the fal.ai balance before committing (see fal-ai-generation 403 troubleshooting).
Whichever path produced the file, the steps below are identical:

1. **Map the track's structure FIRST.** Run `scripts/audio_energy_map.py` on the
   audio (decodes via ffmpeg to raw PCM, computes an RMS energy envelope with pure
   stdlib — no numpy/aubio needed). It prints a per-second energy bar graph plus
   the quietest windows (breakdowns/intros), loudest windows (drops/peaks), and
   sharp onsets (candidate drop hits). See `references/music-driven-assembly.md`.
2. **Build the picture on the song's own spine.** Place the cold open over the
   quiet intro, fast action over sustained energy, the slow-mo apex in the
   breakdown (drums gone), and the hardest trick on the drop's slam. A clean
   breakdown→drop transition already in the track is worth more than fighting the
   edit to fit.
3. **Assemble with FFmpeg**: trim each clip so its cut point lands on a downbeat,
   apply one unified grade (eq contrast/saturation + colorbalance cool shadows /
   warm mids + light unsharp) to every clip, duck wind/SFX under the music, fade
   in/out. `concat` filter for the join.
- **Honest caveat**: beat-perfect sync to *generated* audio is never sample-accurate
  — a few ms of feel remains. Get it tight by eye on the waveform. If the track's
  drop isn't clean where you need it, iterate the TRACK, not the cut.
- **Python 3.13 note**: `audioop` was removed from stdlib; the script computes RMS
  manually with `struct.iter_unpack("<h", ...)`.

## Reviewing Generated Clips
Two QA paths — pick per need:

**Path A — Frame-grab + vision_analyze (default for batch QA).** Fastest and
most reliable for reviewing many clips. Extract a representative mid-clip frame
and analyze the still:
```
ffmpeg -y -v error -i clip.mp4 -vf "select=eq(n\,60)" -vframes 1 frame.jpg
```
(Frame 60 ≈ 2.5s into a 24fps clip; pick a frame mid-action, not the first.)
Then `vision_analyze(frame.jpg)` with a SPECIFIC question: "Rate usability 1-10
as a [shot type]. Is the deck vivid yellow (not green)? Golden hour lighting?
Note problems: morphing, weird geometry, text, people." This catches color
drift, broken geometry, and composition issues at a glance and scales to 20+
clips without timeouts. Limitation: a still can't confirm temporal morphing —
flag "must check in motion" for clips where motion integrity is the risk.

**Path B — video_analyze tool (for motion-specific questions).** Use when you
must verify temporal stability (does the spin morph? does the hand warp
frame-to-frame?). Requires prep:
- **Strip audio first.** FLUX 3 clips have audio baked in. The Gemini backend
  rejects files with audio tracks ("Audio input modality is not enabled").
  Fix: `ffmpeg -i clip.mp4 -an -c:v copy review/clip.mp4` before analyzing.
- **Downscale for reliability.** Files >5MB frequently hit "Download multimodal
  file timed out." Fix: `ffmpeg -i clip.mp4 -an -vf "scale=640:-2" -c:v libx264
  -crf 32 -preset fast review/clip.mp4` — gets most clips under 200KB.
- **Batch in groups of 3–4.** More than 4 concurrent analyses increases timeout
  rate. Space retries; if a file keeps timing out, re-encode smaller.

**Both paths:** Ask specific questions. "Rate usability 1-10 as a [shot type]" +
"Any morphing/warping?" + "Jacket/deck color?" gets actionable verdicts. Vague
prompts get essays.

## Platform Integration: Higgsfield MCP

Higgsfield exposes 30+ models via MCP at `https://mcp.higgsfield.ai/mcp`.
Hermes config (`~/.hermes/config.yaml`):
```yaml
mcp_servers:
  higgsfield:
    url: "https://mcp.higgsfield.ai/mcp"
    auth: oauth
    enabled: true
```
First connect opens browser for Higgsfield OAuth. Tokens cache at `~/.hermes/mcp-tokens/higgsfield.json`.

**Critical**: Unlimited mode is **web-app only**. MCP/CLI always consumes credits. Strategy: bulk exploration on higgsfield.ai web, targeted finals via Hermes+MCP.

## CLI Execution (via higgsfield CLI)

Pairs with the vendor `higgsfield-generate` skill for full model/param discovery.
Key commands:

```bash
# Text-to-video (blocks until done, prints result URL)
higgsfield generate create seedance_2_0 --prompt "..." --duration 8 --resolution 720p --aspect_ratio 16:9 --wait

# Image-to-video (animate a still)
higgsfield generate create seedance_2_0 --prompt "camera dollies in" --start-image ./frame.png --duration 12 --wait

# Image generation
higgsfield generate create gpt_image_2 --prompt "..." --aspect_ratio 16:9 --resolution 2k --wait

# Audio / SFX
higgsfield generate create seed_audio --prompt "cinematic rain ambience with distant thunder" --wait

# Check model params before submitting
higgsfield model get <job_set_type> --json

# List all available models
higgsfield model list --json
```

**Always use `--wait`** so the command blocks and prints the result URL.
For long renders: `--wait-timeout 20m`. For machine-readable output: `--json`.

Media flags: `--image` (reference), `--start-image` (first frame),
`--end-image` (last frame), `--video` (reference/analysis), `--audio` (lipsync/soundtrack).
Each accepts a local file path (auto-uploaded) or a UUID.

## Pitfalls
- Don't exceed 120 words per prompt — contradiction risk rises (EXCEPTION: Flux 3 uses a reasoning harness, not a tag encoder — longer structured prose is fine and often better)
- Don't use prompt weighting syntax (most 2026 video models don't support it)
- Don't stack multiple camera moves in one shot description
- Don't re-describe the medium/style after a style trigger token
- Don't assume all creative work is brand-specific — ask first
- Veo 3.1 caps at 8s per generation; longer scenes need stitching
- Seedance 2.0 full model needs Plus plan+; Starter only gets Fast variant
- **Brainstorming workflow**: When a user brings a specific concept, execute on it — refine, add depth, solve the hard problems. Don't generate a menu of 10 alternatives unless they explicitly ask for options. A user who says "I want X" wants X made better, not X replaced with Y.
- **CLI PATH**: `npm install -g` puts binaries in `$(npm prefix -g)/bin/`
  which may not be in `$PATH`. Fix: `export PATH="$(npm prefix -g)/bin:$PATH"`
  and add to `~/.zshrc`. On Hermes-managed npm this is `~/.hermes/node/bin/`.
- **Skills install**: use `npx skills add higgsfield-ai/skills --yes` for
  non-interactive. Skills land in `~/.agents/skills/` and symlink to
  `~/.hermes/skills/` — if running a non-default profile, also symlink
  into `~/.hermes/profiles/<profile>/skills/`.

## References
- `references/higgsfield-platform.md` — full model roster, pricing, credit costs, plan tiers, Unlimited vs credit mechanics, MCP setup details
- `references/flux3-techniques.md` — Flux 3 technique bank: multi-shot structure, device screens at 720p, hand-object workarounds, duration budgeting, continuation chaining, proven prompt templates
- `references/minimax-h3.md` — MiniMax H3: five-block prompt structure, Omni Reference roles, timed beats, sound design, API shape, strengths/weaknesses
- `references/music-driven-assembly.md` — cutting a multi-clip sequence to a generated track: energy-envelope mapping, building the picture on the song's spine, FFmpeg grade/duck/concat recipe
- `scripts/audio_energy_map.py` — pure-stdlib RMS energy envelope + drop/breakdown detector for an audio file (decodes via ffmpeg, no numpy/aubio). Run: `python3 audio_energy_map.py <audio> [window_secs]`
- Vendor skill `higgsfield-generate` (installed via `npx skills add higgsfield-ai/skills`) — CLI mechanics, model IDs, media flags, Marketing Studio workflows, Virality Predictor. Load it for execution details; load THIS skill for prompt craft and model selection strategy.

