AI Video Generation
When to load
Any task involving AI-generated video: prompting, model selection, scene construction, platform setup, or batch video workflows.
Critical rule: Don't assume branding
Unless the user explicitly names a brand/project (e.g. "Cheonma shoot"), treat video generation as general creative work. Do NOT auto-inject brand triggers, palette DNA, or style tokens. Ask if unclear.
The Six-Part Prompt Structure
Production-grade AI video prompts have six components in order. 60–120 words total. Structure beats length — a structured 60-word prompt outperforms a 200-word stream of consciousness.
[Subject] + [Action] + [Camera] + [Lighting] + [Environment] + [Style]
1. Subject (1–4 concrete nouns, no filler adjectives)
2. Action (one verb phrase with direction/speed)
3. Camera (cinematic terms only — see vocabulary below)
4. Lighting (name source + quality)
5. Environment (location, time, weather, atmosphere)
6. Style (aesthetic anchors, quality tags)
Camera Vocabulary (use these, not generic verbs)
| Shot types |
Movements |
| Wide establishing shot |
Dolly in / out |
| Medium two-shot |
Handheld tracking |
| Over the shoulder |
Crane up / down |
| Dutch tilt |
Whip pan |
| Close-up / extreme close-up |
Slow push-in |
| Extreme low angle |
Arc shot / orbital |
| POV |
Rack focus |
Never say "zoom" or "pan" alone — they produce generic output.
Lighting Vocabulary
Name the source and quality:
golden hour backlight with rim light on subject
single hard source from above, deep contrast
overcast soft key from camera left
neon practicals with blue shadow fill
volumetric god rays through [object]
Never say "good lighting" or "cinematic lighting" alone.
Timing & Pacing (for models that accept duration cues)
slow sustained movement over the full duration
quick three-beat action, hold on final frame
continuous left-to-right drift across four seconds
- One motion rhythm per shot. Never stack contradictory timing.
Negative Prompting Rule
Most video models ignore or backfire on negative prompts. Instead of "no blur," write tack sharp focus throughout. Describe the positive state you want.
Multi-Shot Sequences
For 3–6 connected shots, structure as:
Multi-shot cinematic sequence.
Shot 1: [shot type]. [style]. [subject + action]. [lighting]. [camera move].
Shot 2: ...
Shot N: ... [Freeze frame / hold on final beat.]
Tag character references with @image handles when the platform supports it.
Model Routing (see references/higgsfield-platform.md for full details)
| Need |
Model |
Why |
| Character consistency + 4K |
Kling 3.0 |
Cheapest, multi-shot storyboard, Voice Binding |
| Multi-shot + native audio |
Seedance 2.0 |
Audio-video in one pass, 12 ref inputs |
| Atmospheric/outdoor/wide |
Veo 3.1 |
Global illumination, weather, depth |
| Restyle existing footage |
WAN 2.7 |
Video-reference style transfer |
| Fast drafts / iteration |
MiniMax Hailuo 2.3 / Kling 2.5 Turbo |
Low cost, minimal prompting needed |
| Multimodal ref + native audio |
MiniMax H3 |
9 imgs + 3 video + 3 audio refs, 2K, five-block prompt, synced stereo audio — see references/minimax-h3.md |
| Physics / destruction |
Sora 2 |
Object permanence (API sunsetting Sept 2026) |
| In-Hermes native (no external API) |
FLUX 3 (bfl_flux3_*) |
Reasoning-harness prompting, native audio, multi-shot, continuation chaining |
Workflow pattern: Draft cheap → winnow → re-render winners on premium model.
FLUX 3 (Hermes-native) — Prompting & Techniques
Flux 3 uses a reasoning harness, not a tag encoder. Plain prose beats keyword soup. The harness expands your prompt, so don't restyle it yourself — that stacks a second rewrite and drifts intent. See references/flux3-techniques.md for the full technique bank.
Key differences from tag-based models
- No keyword tricks, no word order games. Write what you want to see in plain language.
- Audio is generated by default. Name ambient sound, music, and speech as separate layers. Say "no music" when unwanted.
- Multi-shot works in one generation but consecutive shots MUST contrast in scale, location, or color or the cut won't read — near-identical coverage blends into a continuous take.
- Quoted text becomes speech only if a speaker is visible. Without one, it renders as burned-in text. Always add "no on-screen text, no subtitles" when you want clean frames.
- 720p output. Mood and composition carry more weight than fine detail. Atmospheric pieces shine; don't rely on legible text or tiny details.
Tool selection
| Scenario |
Tool |
| No input media |
bfl_flux3_text_to_video |
| Animate one image as frame 0 |
bfl_flux3_image_to_video |
| Pin 1-10 images at frame positions |
bfl_flux3_keyframes_to_video |
| Continue from a clip's final frames |
bfl_flux3_video_continuation |
Character consistency across a multi-clip sequence (read before >4 clips)
Text-prompting identity drifts past ~4 clips — a locked style prefix won't hold a
character's look. The fix is a PIXEL anchor, not better words:
- Generate ONE clean "character sheet" reference still (image_generate):
front-facing, neutral static pose, even soft light, full kit visible, plain
background. Feed it as the frame-0 / reference image for every hero shot so the
character's actual pixels seed each clip, not a text description. A turnaround
sheet built from a generated reference instead of pulled frames.
- No visible face is an ADVANTAGE for AI consistency. A helmeted/goggled rider
has no face to get wrong — identity is carried by the KIT (jacket color, pants,
helmet, board graphic). Lock the kit and it's drift-proof. Don't pull anchors
from existing action footage: it's small, motion-blurred, and wardrobe reads
differently clip-to-clip (a crimson jacket reading coral in another shot IS the
drift already happening).
- B-roll is the consistency loophole. Gear macros, hands, snow, environment,
and POV/GoPro shots need no identity at all. A ~25–40% B-roll ratio rests the
eye and hides drift in hero shots. POV is a consistency gimme — high energy,
zero identity required.
- One saturated wardrobe accent (red jacket vs blue-white snow) reads as "same
rider" even across scale/angle changes — the cheapest identity anchor there is.
- Anthology casting turns drift into a feature. For a commercial/ad that reads
as "a production with a bunch of people" (sports brand film, crew montage), cast
4–5 DISTINCT riders with distinct kits (one saturated accent each: red deck, blue
deck, yellow deck…) instead of fighting to hold ONE face across 20 clips. Now any
inter-clip variation reads as intentional casting, not drift — the hardest problem
in long AI-video sequences dissolves. Still lock each rider's kit + a recurring
product hero (deck graphic / shoe) so the world feels unified. Best paired with
~30% B-roll + POV (zero-identity shots) to rest the eye between hero riders.
Action-sports trick shots are the hardest category (learned 2026-07, skate film)
On a golden-hour skate ad, the score split was stark: B-roll / establishing /
environment shots (no identity, no board-rider interaction) consistently scored
6–8/10, but hero trick shots — a believable human ON a board doing a specific
trick — scored 2–4/10. The model scrambles hardware + anatomy + physics at
once: the deck vanishes or shows the wrong face (underside instead of grip
tape), trucks/wheels melt or invert, the rider "floats" in a pose matching no
real trick, an arm goes missing. Even image-to-video from a clean starting-pose
still (the usual identity fix) did NOT save it — the board-rider contact is the
failure, not identity.
Implications for planning an action-sports film:
- Don't promise clean hero tricks. Budget the edit to imply tricks (fast 1–2s
cuts where morphing reads as motion blur, B-roll of the spot, POV, environment,
reaction) rather than showing them held and clean.
- Yellow decks drift to lime/olive-green on EVERY text-to-video gen, and
backlighting pushes them to dark olive. Lock a saturated deck color ONLY via
image-to-video from a clean reference frame; never trust a text prompt for it.
- POV shots fail as a category (board geometry + speed cues + obstacle legibility
all break together). Use sparingly as flashes, not heroes.
- Golden-hour lighting and unified grade are the EASY, reliable win — lean on
mood/cut rhythm over trick fidelity.
Wardrobe/color drift (happens on EVERY independent generation)
Text-to-video re-rolls wardrobe on every separate generation — even with an
identical style prefix, a "orange-red jacket" prompt produced olive-green sleeves
(POV shot) and a dark cuff (hand macro) in 2 of 8 clips. The style prefix helps
with world consistency but does NOT lock garment color across independent gens.
Fixes, in order of reliability:
- Image-to-video from a clean reference frame — the jacket's actual pixels
seed the generation. Most reliable. Pull the frame from footage where the
garment is clearly visible in neutral light (not backlit/silhouetted).
- Reinforce with a vivid analogy + prohibitions — "vivid, saturated
orange-red — the same bright orange-red as a traffic cone" plus "The jacket
MUST be vivid orange-red throughout, never green, never dark, never black,
never any other color." Worked on regeneration for both drifted clips.
- Make the color a foreground element — for shots where the garment is
barely visible (POV sleeves at frame edge, a hand macro cuff), explicitly
name it as prominent/visible in the framing description so the model can't
ignore it.
Saturated prop/product color drift (worse than wardrobe)
Small saturated objects — skateboard decks, shoes, product hero items — drift
MORE than wardrobe in text-to-video, because they occupy fewer pixels and the
model re-rolls their color on every independent generation. Yellow is the
worst offender: a "vivid bright yellow" deck consistently rendered as
lime/chartreuse green-yellow or dark olive across multiple text-to-video gens
(4/10 and 3/10 QA scores). Backlighting compounds it catastrophically — a
yellow deck silhouetted against a sunset sky rendered as dark olive-green.
Even explicit prohibitions ("never green, never lime") failed in text-to-video.
Fixes, in order of reliability:
- Image-to-video from a clean reference frame where the prop's actual
pixels seed the generation. Generate the starting frame with image_generate
(which handles yellow correctly), then animate it. This is the ONLY
reliable fix for product-hero color accuracy.
- Front-fill the product — never backlight a color-critical product shot.
Silhouetting suppresses the surface color into shadow. Add "front fill
light on the [product], not backlit" to the prompt.
- Reinforce with analogy + prohibitions (same as wardrobe) — works for
wardrobe but is NOT sufficient for small saturated props.
POV shots (hardest category — 3-5/10 across all attempts)
First-person POV consistently fails on THREE axes simultaneously:
- Board geometry: the model shows the deck's underside graphic instead of
grip tape (impossible while riding), or renders a floating board nub with no
trucks, wheels, or feet.
- Speed cues: no motion blur, no foreground streak, no camera vibration —
reads as static/parked, not rolling at speed.
- Obstacle legibility: the target (rail, ledge, bowl) is either missing,
ambiguous (a lone vertical pole instead of a handrail), or blown out by the
sun bloom sitting exactly where the obstacle should be.
Practical guidance: Accept POV as sub-second flashes (0.5–1s) in a fast
montage where the eye can't audit details. Do NOT use as held hero shots.
Budget for 2–3 re-rolls per POV clip. If the edit can survive with fewer POV
shots, cut them first — they're the lowest-yield category.
Mechanical/hardware geometry (trucks, wheels, bearings)
AI video scrambles skateboard hardware anatomy:
- Trucks render inverted (baseplate grinding the rail instead of hanger/axle).
- Wheels go missing from axles, or warp oval, or multiply.
- Bearings render as "metallic soup" — no readable balls, seals, or races.
- Axle ends render as smooth domes instead of threaded hex nuts.
Practical guidance: Hardware macros (bearing spins, truck grinds, wheel
close-ups) are gorgeous in mood but mechanically incoherent. Use them as
sub-second texture flashes in a montage — never held shots. At 1s with grain
and motion, the eye registers "warm, shiny, mechanical" and doesn't audit
anatomy. At 3s+ or in slow-mo, the geometry collapses.
Starting-pose stills for action shots (the image-to-video anchor technique)
For hero trick shots, DON'T use text-to-video. Instead:
- Generate a clean starting-pose still with image_generate: the rider in
the wind-up position, correct kit, correct deck color, golden-hour lighting,
plain or simple background. This is the frame-0 pixel anchor.
- Feed it to image-to-video as the opening frame. The rider's actual
pixels (kit colors, deck color, body proportions) seed every subsequent
frame — far more reliable than text description.
- Pose guidelines for the starting still:
- Clear wheel-to-ground contact — all four wheels visibly planted on a
surface. Ambiguous contact (hovering trucks, floating board) produces
snap/slide/float artifacts in the animation.
- All limbs visible — no occluded arms. The video model must invent
hidden limbs as the body turns, producing popping/duplicated arms.
- Wind-up pose, not spent pose — a deep crouch that's already "committed"
leaves no motion vector to animate. A standing wind-up (weight back, about
to lean forward) gives the model a clear "lean-and-roll" direction.
- Lock contact in the video prompt — explicitly state "wheels stay in
contact with the concrete throughout" and "tail planted on the coping" to
prevent the board from floating or snapping.
- QA the starting still before committing to the video gen (which takes
minutes). Check: correct kit colors, correct deck color, clear pose, no
anatomy errors, good lighting. A 6/10 starting pose with explicit
contact-locking language in the video prompt outperforms a 9/10 text-to-video
gen for identity consistency.
Aerial / spin / rotation morphing (and the held-pose fix)
Spinning or rotating tricks morph badly: the board changes length/width/shape
frame-to-frame, the body stretches and compresses, the grab hand melts into the
deck. Even with strong anti-morph constraints, a spinning aerial scored 3–4/10.
Writing the prompt: minimize rotation. Write a HELD POSE at the apex with
"nearly motionless," "the same held pose," "perfectly rigid, stable, and
consistent in shape and proportion throughout with absolutely no morphing,
warping, stretching, or distortion," and "very high shutter speed for crisp
frozen motion." This reduces but doesn't eliminate morphing.
Salvage techniques for morphy footage:
- Slow it down (ffmpeg
setpts=2.4*PTS): at 2.4× with minimal rotation,
artifacts become nearly invisible. A 5s clip fills 12s of screen time. QA
confirmed "clean, no warping" after slowing — the strongest part of the edit.
- Cut it fast (1–2s): morphing reads as motion blur at speed. Fine for
montage sections. Never use morphy footage as a clean slow-mo hero at 1×.
Hard-won techniques
- Phone/device screens at 720p: Don't try to render readable text. Describe the screen as a visual object — "dark terminal interface, amber prompt line, blinking cursor, lines of monospace text scrolling." The audience's brain fills in "terminal" from the shape. Specifying actual text content produces garbled output or burned-in subtitles.
- Hand-object interaction (pulling phone from pocket, picking up items): This is the hardest thing for any video model. Workaround: start the shot with the object already in hand. "Hands holding a glowing phone" works; "hand reaches into pocket and retrieves phone" usually fails.
- Narrative beats in multi-shot: Give each shot a distinct emotional beat, not just a visual change. Approach → command → payoff. The contrast in intent makes cuts read cleaner than contrast in framing alone.
- Duration pacing: 15 seconds for three shots gives ~5s per beat. 10 seconds for three shots is too tight — beats feel rushed. Budget at least 4-5 seconds per shot for the action to breathe.
- Continuation chaining: Segments max at 15s. Plan each segment as a self-contained beat, then chain with
bfl_flux3_video_continuation. Open the prompt with "Continue this video from its final frames:" and re-establish the subject.
Workflow
- Submit → get job id immediately (video doesn't exist yet)
- Poll
bfl_flux3_get_result with the id — generation takes several minutes, long "Generating" phase is normal
- The poll call blocks while the job runs; if it returns still-generating, just call again — no sleeping
- On Ready, the clip is downloaded and
saved_path is returned. Deliver the file per platform conventions.
- Job survives client restarts — re-poll the same id, never resubmit (duplicates spend budget)
- Rate limit: one submission per minute. If you get a "limited to one attempt
per minute" error,
sleep 45 then resubmit. When generating a batch serially,
the natural poll-wait between clips usually covers this, but back-to-back
submits (e.g. firing two clips in one turn) will trip it.
- Transient 504s happen. Azure Front Door occasionally returns a 504 HTML
page instead of a job id. Just wait the rate-limit cooldown and resubmit —
the prompt wasn't consumed.
save_to quirk: if you pass a path that doesn't exist yet and has no
extension (e.g. .../new_clips), the tool creates a FILE with that name
instead of a directory. Always either pre-create the directory (mkdir -p)
or pass a full file path with .mp4 extension.
Multi-Clip Production Pipeline (15–25 clips → short film / commercial)
For a long-form piece (a skate/snowboard/sports ad, a short film) assembled from
many independent clips, run these phases IN ORDER. Skipping ahead burns budget —
each FLUX 3 gen is minutes + rate-limited, so lock everything upstream first.
- Spec & storyboard FIRST. Write an AD_SPEC.md: setting, mood, lighting,
pacing, brand feel, target length. Then a SHOT_LIST.md with a shot taxonomy:
establishing / B-roll-texture / POV / hero tricks / ride-away. Budget a
25–35% B-roll ratio (zero-identity shots rest the eye and hide hero drift).
Shoot MORE clips than you'll use (24–26 for a 70s film) so the edit has
options — cut down to the best ~18–20.
- Music is the spine — generate it BEFORE the clips. Map the track's energy
(see Music-Driven Assembly below) and lock a beat map (cold open over intro,
action over sustained energy, slow-mo apex in the breakdown, hardest trick on
the drop). The picture bends to the song, not vice versa.
- Cast the riders / characters. For an anthology (multiple distinct people),
generate ONE clean character-sheet reference still per rider (front-facing,
neutral pose, even light, full kit, plain bg) — each with a distinct saturated
accent (red deck, blue deck, yellow deck…). Anthology casting turns identity
drift into intentional casting. See Character consistency section.
- Generate in phases (optimized for the rate limit + QA flow):
- Phase 1: establishing + B-roll (text-to-video, no identity — warms up the
pipeline, cheapest to re-roll).
- Phase 2: POV (text-to-video, zero identity — expect 3–5/10, budget re-rolls).
- Phase 3: hero tricks (image-to-video from starting-pose stills — the
payload, best identity consistency).
- Phase 4: outro / ride-away.
Run as a managed batch: fire one clip → poll to completion → QA → fire next.
The natural poll-wait covers the rate limit. Report progress in chunks.
- QA every clip (frame-grab + vision_analyze, see Reviewing below). Keep a
re-roll queue for anything <6/10. Batch re-rolls at the end, using
image-to-video where color/geometry matters.
- Assemble on the song's spine (Music-Driven Assembly): trim to downbeats,
unified grade on every clip, 15% Ken Burns punch-in, hard cuts, audio duck,
fade. Deliver full film + teaser + individual clips + thumbnails + README in a
numbered folder (01_raw_clips / 02_final_edits / 03_references / 04_audio /
05_scripts / _scratch/review_proxies).
Time reality: 24–26 clips at the ~5-min account-wide rate limit + iteration
≈ 2–3 hours of generation. Don't promise a clip length the tools can't make in
one pass (>20s needs continuation chaining).
Music-Driven Assembly (cut the picture to the song's spine)
When scoring a multi-clip sequence to a generated track, the music becomes the
timeline and the picture bends to it — cuts land on beats, the slow-mo hero sits
in a breakdown, the biggest action hits on the drop. Don't cut by feel and bolt
music on after.
Where the track comes from (routing):
| Scenario |
Path |
| fal.ai balance available |
fal-ai/elevenlabs/music via FAL_KEY — text-to-music, section-by-section composition_plan, ~$0.80/min (see fal-ai-generation skill) |
| fal.ai balance exhausted (403 "User is locked") |
Suno via web UI — write the prompt with the suno-music-creation skill (DSL style prompt → Style field, structure tags → Lyrics field, Instrumental ON); user generates + returns the MP3. Suno returns a FULL song (~90s–4min): mine the cleanest breakdown→drop section. |
| Offline / local |
HeartMuLa (CUDA GPU) or AudioCraft MusicGen (CPU/MPS) |
| Probe the fal.ai balance before committing (see fal-ai-generation 403 troubleshooting). |
|
| Whichever path produced the file, the steps below are identical: |
|
- Map the track's structure FIRST. Run
scripts/audio_energy_map.py on the
audio (decodes via ffmpeg to raw PCM, computes an RMS energy envelope with pure
stdlib — no numpy/aubio needed). It prints a per-second energy bar graph plus
the quietest windows (breakdowns/intros), loudest windows (drops/peaks), and
sharp onsets (candidate drop hits). See references/music-driven-assembly.md.
- Build the picture on the song's own spine. Place the cold open over the
quiet intro, fast action over sustained energy, the slow-mo apex in the
breakdown (drums gone), and the hardest trick on the drop's slam. A clean
breakdown→drop transition already in the track is worth more than fighting the
edit to fit.
- Assemble with FFmpeg: trim each clip so its cut point lands on a downbeat,
apply one unified grade (eq contrast/saturation + colorbalance cool shadows /
warm mids + light unsharp) to every clip, duck wind/SFX under the music, fade
in/out.
concat filter for the join.
- Honest caveat: beat-perfect sync to generated audio is never sample-accurate
— a few ms of feel remains. Get it tight by eye on the waveform. If the track's
drop isn't clean where you need it, iterate the TRACK, not the cut.
- Python 3.13 note:
audioop was removed from stdlib; the script computes RMS
manually with struct.iter_unpack("<h", ...).
Reviewing Generated Clips
Two QA paths — pick per need:
Path A — Frame-grab + vision_analyze (default for batch QA). Fastest and
most reliable for reviewing many clips. Extract a representative mid-clip frame
and analyze the still:
ffmpeg -y -v error -i clip.mp4 -vf "select=eq(n\,60)" -vframes 1 frame.jpg
(Frame 60 ≈ 2.5s into a 24fps clip; pick a frame mid-action, not the first.)
Then vision_analyze(frame.jpg) with a SPECIFIC question: "Rate usability 1-10
as a [shot type]. Is the deck vivid yellow (not green)? Golden hour lighting?
Note problems: morphing, weird geometry, text, people." This catches color
drift, broken geometry, and composition issues at a glance and scales to 20+
clips without timeouts. Limitation: a still can't confirm temporal morphing —
flag "must check in motion" for clips where motion integrity is the risk.
Path B — video_analyze tool (for motion-specific questions). Use when you
must verify temporal stability (does the spin morph? does the hand warp
frame-to-frame?). Requires prep:
- Strip audio first. FLUX 3 clips have audio baked in. The Gemini backend
rejects files with audio tracks ("Audio input modality is not enabled").
Fix:
ffmpeg -i clip.mp4 -an -c:v copy review/clip.mp4 before analyzing.
- Downscale for reliability. Files >5MB frequently hit "Download multimodal
file timed out." Fix:
ffmpeg -i clip.mp4 -an -vf "scale=640:-2" -c:v libx264 -crf 32 -preset fast review/clip.mp4 — gets most clips under 200KB.
- Batch in groups of 3–4. More than 4 concurrent analyses increases timeout
rate. Space retries; if a file keeps timing out, re-encode smaller.
Both paths: Ask specific questions. "Rate usability 1-10 as a [shot type]" +
"Any morphing/warping?" + "Jacket/deck color?" gets actionable verdicts. Vague
prompts get essays.
Platform Integration: Higgsfield MCP
Higgsfield exposes 30+ models via MCP at https://mcp.higgsfield.ai/mcp.
Hermes config (~/.hermes/config.yaml):
mcp_servers:
higgsfield:
url: "https://mcp.higgsfield.ai/mcp"
auth: oauth
enabled: true
First connect opens browser for Higgsfield OAuth. Tokens cache at ~/.hermes/mcp-tokens/higgsfield.json.
Critical: Unlimited mode is web-app only. MCP/CLI always consumes credits. Strategy: bulk exploration on higgsfield.ai web, targeted finals via Hermes+MCP.
CLI Execution (via higgsfield CLI)
Pairs with the vendor higgsfield-generate skill for full model/param discovery.
Key commands:
# Text-to-video (blocks until done, prints result URL)
higgsfield generate create seedance_2_0 --prompt "..." --duration 8 --resolution 720p --aspect_ratio 16:9 --wait
# Image-to-video (animate a still)
higgsfield generate create seedance_2_0 --prompt "camera dollies in" --start-image ./frame.png --duration 12 --wait
# Image generation
higgsfield generate create gpt_image_2 --prompt "..." --aspect_ratio 16:9 --resolution 2k --wait
# Audio / SFX
higgsfield generate create seed_audio --prompt "cinematic rain ambience with distant thunder" --wait
# Check model params before submitting
higgsfield model get <job_set_type> --json
# List all available models
higgsfield model list --json
Always use --wait so the command blocks and prints the result URL.
For long renders: --wait-timeout 20m. For machine-readable output: --json.
Media flags: --image (reference), --start-image (first frame),
--end-image (last frame), --video (reference/analysis), --audio (lipsync/soundtrack).
Each accepts a local file path (auto-uploaded) or a UUID.
Pitfalls
- Don't exceed 120 words per prompt — contradiction risk rises (EXCEPTION: Flux 3 uses a reasoning harness, not a tag encoder — longer structured prose is fine and often better)
- Don't use prompt weighting syntax (most 2026 video models don't support it)
- Don't stack multiple camera moves in one shot description
- Don't re-describe the medium/style after a style trigger token
- Don't assume all creative work is brand-specific — ask first
- Veo 3.1 caps at 8s per generation; longer scenes need stitching
- Seedance 2.0 full model needs Plus plan+; Starter only gets Fast variant
- Brainstorming workflow: When a user brings a specific concept, execute on it — refine, add depth, solve the hard problems. Don't generate a menu of 10 alternatives unless they explicitly ask for options. A user who says "I want X" wants X made better, not X replaced with Y.
- CLI PATH:
npm install -g puts binaries in $(npm prefix -g)/bin/
which may not be in $PATH. Fix: export PATH="$(npm prefix -g)/bin:$PATH"
and add to ~/.zshrc. On Hermes-managed npm this is ~/.hermes/node/bin/.
- Skills install: use
npx skills add higgsfield-ai/skills --yes for
non-interactive. Skills land in ~/.agents/skills/ and symlink to
~/.hermes/skills/ — if running a non-default profile, also symlink
into ~/.hermes/profiles/<profile>/skills/.
References
references/higgsfield-platform.md — full model roster, pricing, credit costs, plan tiers, Unlimited vs credit mechanics, MCP setup details
references/flux3-techniques.md — Flux 3 technique bank: multi-shot structure, device screens at 720p, hand-object workarounds, duration budgeting, continuation chaining, proven prompt templates
references/minimax-h3.md — MiniMax H3: five-block prompt structure, Omni Reference roles, timed beats, sound design, API shape, strengths/weaknesses
references/music-driven-assembly.md — cutting a multi-clip sequence to a generated track: energy-envelope mapping, building the picture on the song's spine, FFmpeg grade/duck/concat recipe
scripts/audio_energy_map.py — pure-stdlib RMS energy envelope + drop/breakdown detector for an audio file (decodes via ffmpeg, no numpy/aubio). Run: python3 audio_energy_map.py <audio> [window_secs]
- Vendor skill
higgsfield-generate (installed via npx skills add higgsfield-ai/skills) — CLI mechanics, model IDs, media flags, Marketing Studio workflows, Virality Predictor. Load it for execution details; load THIS skill for prompt craft and model selection strategy.
1---2name: ai-video-generation3description: AI video generation: prompt engineering, model selection, and platform integration (Higgsfield, Flux 3, Runway, etc.). Six-part prompt structure, camera/lighting vocabulary, model routing, MCP automation, Flux 3 reasoning-harness techniques.4---5
6# AI Video Generation
7
8## When to load
9Any task involving AI-generated video: prompting, model selection, scene construction, platform setup, or batch video workflows.
10
11## Critical rule: Don't assume branding
12Unless the user explicitly names a brand/project (e.g. "Cheonma shoot"), treat video generation as **general creative work**. Do NOT auto-inject brand triggers, palette DNA, or style tokens. Ask if unclear.
13
14## The Six-Part Prompt Structure
15
16Production-grade AI video prompts have six components **in order**. 60–120 words total. Structure beats length — a structured 60-word prompt outperforms a 200-word stream of consciousness.
17
18```
19[Subject] + [Action] + [Camera] + [Lighting] + [Environment] + [Style]
20```
21
22### 1. Subject (1–4 concrete nouns, no filler adjectives)
23### 2. Action (one verb phrase with direction/speed)
24### 3. Camera (cinematic terms only — see vocabulary below)
25### 4. Lighting (name source + quality)
26### 5. Environment (location, time, weather, atmosphere)
27### 6. Style (aesthetic anchors, quality tags)
28
29## Camera Vocabulary (use these, not generic verbs)
30
31| Shot types | Movements |
32|---|---|
33| Wide establishing shot | Dolly in / out |
34| Medium two-shot | Handheld tracking |
35| Over the shoulder | Crane up / down |
36| Dutch tilt | Whip pan |
37| Close-up / extreme close-up | Slow push-in |
38| Extreme low angle | Arc shot / orbital |
39| POV | Rack focus |
40
41**Never say** "zoom" or "pan" alone — they produce generic output.
42
43## Lighting Vocabulary
44
45Name the **source** and **quality**:
46- `golden hour backlight with rim light on subject`
47- `single hard source from above, deep contrast`
48- `overcast soft key from camera left`
49- `neon practicals with blue shadow fill`
50- `volumetric god rays through [object]`
51
52**Never say** "good lighting" or "cinematic lighting" alone.
53
54## Timing & Pacing (for models that accept duration cues)
55- `slow sustained movement over the full duration`
56- `quick three-beat action, hold on final frame`
57- `continuous left-to-right drift across four seconds`
58- One motion rhythm per shot. Never stack contradictory timing.
59
60## Negative Prompting Rule
61Most video models **ignore or backfire** on negative prompts. Instead of "no blur," write `tack sharp focus throughout`. Describe the positive state you want.
62
63## Multi-Shot Sequences
64For 3–6 connected shots, structure as:
65```
66Multi-shot cinematic sequence.
67
68Shot 1: [shot type]. [style]. [subject + action]. [lighting]. [camera move].
69Shot 2: ...
70Shot N: ... [Freeze frame / hold on final beat.]
71```
72Tag character references with @image handles when the platform supports it.
73
74## Model Routing (see references/higgsfield-platform.md for full details)
75
76| Need | Model | Why |
77|---|---|---|
78| Character consistency + 4K | Kling 3.0 | Cheapest, multi-shot storyboard, Voice Binding |
79| Multi-shot + native audio | Seedance 2.0 | Audio-video in one pass, 12 ref inputs |
80| Atmospheric/outdoor/wide | Veo 3.1 | Global illumination, weather, depth |
81| Restyle existing footage | WAN 2.7 | Video-reference style transfer |
82| Fast drafts / iteration | MiniMax Hailuo 2.3 / Kling 2.5 Turbo | Low cost, minimal prompting needed |
83| Multimodal ref + native audio | MiniMax H3 | 9 imgs + 3 video + 3 audio refs, 2K, five-block prompt, synced stereo audio — see references/minimax-h3.md |
84| Physics / destruction | Sora 2 | Object permanence (API sunsetting Sept 2026) |
85| In-Hermes native (no external API) | FLUX 3 (bfl_flux3_*) | Reasoning-harness prompting, native audio, multi-shot, continuation chaining |
86
87**Workflow pattern**: Draft cheap → winnow → re-render winners on premium model.
88
89## FLUX 3 (Hermes-native) — Prompting & Techniques
90
91Flux 3 uses a reasoning harness, not a tag encoder. Plain prose beats keyword soup. The harness expands your prompt, so don't restyle it yourself — that stacks a second rewrite and drifts intent. See `references/flux3-techniques.md` for the full technique bank.
92
93### Key differences from tag-based models
94- **No keyword tricks, no word order games.** Write what you want to see in plain language.
95- **Audio is generated by default.** Name ambient sound, music, and speech as separate layers. Say "no music" when unwanted.
96- **Multi-shot works in one generation** but consecutive shots MUST contrast in scale, location, or color or the cut won't read — near-identical coverage blends into a continuous take.
97- **Quoted text becomes speech only if a speaker is visible.** Without one, it renders as burned-in text. Always add "no on-screen text, no subtitles" when you want clean frames.
98- **720p output.** Mood and composition carry more weight than fine detail. Atmospheric pieces shine; don't rely on legible text or tiny details.
99
100### Tool selection
101| Scenario | Tool |
102|---|---|
103| No input media | `bfl_flux3_text_to_video` |
104| Animate one image as frame 0 | `bfl_flux3_image_to_video` |
105| Pin 1-10 images at frame positions | `bfl_flux3_keyframes_to_video` |
106| Continue from a clip's final frames | `bfl_flux3_video_continuation` |
107
108### Character consistency across a multi-clip sequence (read before >4 clips)
109Text-prompting identity drifts past ~4 clips — a locked style prefix won't hold a
110character's look. The fix is a PIXEL anchor, not better words:
111- **Generate ONE clean "character sheet" reference still** (image_generate):
112 front-facing, neutral static pose, even soft light, full kit visible, plain
113 background. Feed it as the frame-0 / reference image for every hero shot so the
114 character's *actual pixels* seed each clip, not a text description. A turnaround
115 sheet built from a generated reference instead of pulled frames.
116- **No visible face is an ADVANTAGE for AI consistency.** A helmeted/goggled rider
117 has no face to get wrong — identity is carried by the KIT (jacket color, pants,
118 helmet, board graphic). Lock the kit and it's drift-proof. Don't pull anchors
119 from existing action footage: it's small, motion-blurred, and wardrobe reads
120 differently clip-to-clip (a crimson jacket reading coral in another shot IS the
121 drift already happening).
122- **B-roll is the consistency loophole.** Gear macros, hands, snow, environment,
123 and POV/GoPro shots need no identity at all. A ~25–40% B-roll ratio rests the
124 eye and hides drift in hero shots. POV is a consistency gimme — high energy,
125 zero identity required.
126- **One saturated wardrobe accent** (red jacket vs blue-white snow) reads as "same
127 rider" even across scale/angle changes — the cheapest identity anchor there is.
128- **Anthology casting turns drift into a feature.** For a commercial/ad that reads
129 as "a production with a bunch of people" (sports brand film, crew montage), cast
130 4–5 DISTINCT riders with distinct kits (one saturated accent each: red deck, blue
131 deck, yellow deck…) instead of fighting to hold ONE face across 20 clips. Now any
132 inter-clip variation reads as intentional casting, not drift — the hardest problem
133 in long AI-video sequences dissolves. Still lock each rider's kit + a recurring
134 product hero (deck graphic / shoe) so the world feels unified. Best paired with
135 ~30% B-roll + POV (zero-identity shots) to rest the eye between hero riders.
136
137### Action-sports trick shots are the hardest category (learned 2026-07, skate film)
138On a golden-hour skate ad, the score split was stark: B-roll / establishing /
139environment shots (no identity, no board-rider interaction) consistently scored
1406–8/10, but **hero trick shots — a believable human ON a board doing a specific
141trick — scored 2–4/10.** The model scrambles hardware + anatomy + physics at
142once: the deck vanishes or shows the wrong face (underside instead of grip
143tape), trucks/wheels melt or invert, the rider "floats" in a pose matching no
144real trick, an arm goes missing. Even image-to-video from a clean starting-pose
145still (the usual identity fix) did NOT save it — the board-rider contact is the
146failure, not identity.
147**Implications for planning an action-sports film:**
148- Don't promise clean hero tricks. Budget the edit to *imply* tricks (fast 1–2s
149 cuts where morphing reads as motion blur, B-roll of the spot, POV, environment,
150 reaction) rather than showing them held and clean.
151- Yellow decks drift to lime/olive-green on EVERY text-to-video gen, and
152 backlighting pushes them to dark olive. Lock a saturated deck color ONLY via
153 image-to-video from a clean reference frame; never trust a text prompt for it.
154- POV shots fail as a category (board geometry + speed cues + obstacle legibility
155 all break together). Use sparingly as flashes, not heroes.
156- Golden-hour lighting and unified grade are the EASY, reliable win — lean on
157 mood/cut rhythm over trick fidelity.
158
159### Wardrobe/color drift (happens on EVERY independent generation)
160Text-to-video re-rolls wardrobe on every separate generation — even with an
161identical style prefix, a "orange-red jacket" prompt produced olive-green sleeves
162(POV shot) and a dark cuff (hand macro) in 2 of 8 clips. The style prefix helps
163with world consistency but does NOT lock garment color across independent gens.
164**Fixes, in order of reliability:**
1651. **Image-to-video from a clean reference frame** — the jacket's actual pixels
166 seed the generation. Most reliable. Pull the frame from footage where the
167 garment is clearly visible in neutral light (not backlit/silhouetted).
1682. **Reinforce with a vivid analogy + prohibitions** — "vivid, saturated
169 orange-red — the same bright orange-red as a traffic cone" plus "The jacket
170 MUST be vivid orange-red throughout, never green, never dark, never black,
171 never any other color." Worked on regeneration for both drifted clips.
1723. **Make the color a foreground element** — for shots where the garment is
173 barely visible (POV sleeves at frame edge, a hand macro cuff), explicitly
174 name it as prominent/visible in the framing description so the model can't
175 ignore it.
176
177### Saturated prop/product color drift (worse than wardrobe)
178Small saturated objects — skateboard decks, shoes, product hero items — drift
179MORE than wardrobe in text-to-video, because they occupy fewer pixels and the
180model re-rolls their color on every independent generation. **Yellow is the
181worst offender**: a "vivid bright yellow" deck consistently rendered as
182lime/chartreuse green-yellow or dark olive across multiple text-to-video gens
183(4/10 and 3/10 QA scores). Backlighting compounds it catastrophically — a
184yellow deck silhouetted against a sunset sky rendered as dark olive-green.
185Even explicit prohibitions ("never green, never lime") failed in text-to-video.
186**Fixes, in order of reliability:**
1871. **Image-to-video from a clean reference frame** where the prop's actual
188 pixels seed the generation. Generate the starting frame with image_generate
189 (which handles yellow correctly), then animate it. This is the ONLY
190 reliable fix for product-hero color accuracy.
1912. **Front-fill the product** — never backlight a color-critical product shot.
192 Silhouetting suppresses the surface color into shadow. Add "front fill
193 light on the [product], not backlit" to the prompt.
1943. **Reinforce with analogy + prohibitions** (same as wardrobe) — works for
195 wardrobe but is NOT sufficient for small saturated props.
196
197### POV shots (hardest category — 3-5/10 across all attempts)
198First-person POV consistently fails on THREE axes simultaneously:
199- **Board geometry**: the model shows the deck's underside graphic instead of
200 grip tape (impossible while riding), or renders a floating board nub with no
201 trucks, wheels, or feet.
202- **Speed cues**: no motion blur, no foreground streak, no camera vibration —
203 reads as static/parked, not rolling at speed.
204- **Obstacle legibility**: the target (rail, ledge, bowl) is either missing,
205 ambiguous (a lone vertical pole instead of a handrail), or blown out by the
206 sun bloom sitting exactly where the obstacle should be.
207**Practical guidance:** Accept POV as sub-second flashes (0.5–1s) in a fast
208montage where the eye can't audit details. Do NOT use as held hero shots.
209Budget for 2–3 re-rolls per POV clip. If the edit can survive with fewer POV
210shots, cut them first — they're the lowest-yield category.
211
212### Mechanical/hardware geometry (trucks, wheels, bearings)
213AI video scrambles skateboard hardware anatomy:
214- Trucks render inverted (baseplate grinding the rail instead of hanger/axle).
215- Wheels go missing from axles, or warp oval, or multiply.
216- Bearings render as "metallic soup" — no readable balls, seals, or races.
217- Axle ends render as smooth domes instead of threaded hex nuts.
218**Practical guidance:** Hardware macros (bearing spins, truck grinds, wheel
219close-ups) are gorgeous in mood but mechanically incoherent. Use them as
220sub-second texture flashes in a montage — never held shots. At 1s with grain
221and motion, the eye registers "warm, shiny, mechanical" and doesn't audit
222anatomy. At 3s+ or in slow-mo, the geometry collapses.
223
224### Starting-pose stills for action shots (the image-to-video anchor technique)
225For hero trick shots, DON'T use text-to-video. Instead:
2261. **Generate a clean starting-pose still** with image_generate: the rider in
227 the wind-up position, correct kit, correct deck color, golden-hour lighting,
228 plain or simple background. This is the frame-0 pixel anchor.
2292. **Feed it to image-to-video** as the opening frame. The rider's actual
230 pixels (kit colors, deck color, body proportions) seed every subsequent
231 frame — far more reliable than text description.
2323. **Pose guidelines for the starting still:**
233 - **Clear wheel-to-ground contact** — all four wheels visibly planted on a
234 surface. Ambiguous contact (hovering trucks, floating board) produces
235 snap/slide/float artifacts in the animation.
236 - **All limbs visible** — no occluded arms. The video model must invent
237 hidden limbs as the body turns, producing popping/duplicated arms.
238 - **Wind-up pose, not spent pose** — a deep crouch that's already "committed"
239 leaves no motion vector to animate. A standing wind-up (weight back, about
240 to lean forward) gives the model a clear "lean-and-roll" direction.
241 - **Lock contact in the video prompt** — explicitly state "wheels stay in
242 contact with the concrete throughout" and "tail planted on the coping" to
243 prevent the board from floating or snapping.
2444. **QA the starting still** before committing to the video gen (which takes
245 minutes). Check: correct kit colors, correct deck color, clear pose, no
246 anatomy errors, good lighting. A 6/10 starting pose with explicit
247 contact-locking language in the video prompt outperforms a 9/10 text-to-video
248 gen for identity consistency.
249
250### Aerial / spin / rotation morphing (and the held-pose fix)
251Spinning or rotating tricks morph badly: the board changes length/width/shape
252frame-to-frame, the body stretches and compresses, the grab hand melts into the
253deck. Even with strong anti-morph constraints, a spinning aerial scored 3–4/10.
254**Writing the prompt:** minimize rotation. Write a HELD POSE at the apex with
255"nearly motionless," "the same held pose," "perfectly rigid, stable, and
256consistent in shape and proportion throughout with absolutely no morphing,
257warping, stretching, or distortion," and "very high shutter speed for crisp
258frozen motion." This reduces but doesn't eliminate morphing.
259**Salvage techniques for morphy footage:**
260- **Slow it down** (ffmpeg `setpts=2.4*PTS`): at 2.4× with minimal rotation,
261 artifacts become nearly invisible. A 5s clip fills 12s of screen time. QA
262 confirmed "clean, no warping" after slowing — the strongest part of the edit.
263- **Cut it fast** (1–2s): morphing reads as motion blur at speed. Fine for
264 montage sections. Never use morphy footage as a clean slow-mo hero at 1×.
265
266### Hard-won techniques
267- **Phone/device screens at 720p**: Don't try to render readable text. Describe the screen as a visual object — "dark terminal interface, amber prompt line, blinking cursor, lines of monospace text scrolling." The audience's brain fills in "terminal" from the shape. Specifying actual text content produces garbled output or burned-in subtitles.
268- **Hand-object interaction** (pulling phone from pocket, picking up items): This is the hardest thing for any video model. Workaround: start the shot with the object already in hand. "Hands holding a glowing phone" works; "hand reaches into pocket and retrieves phone" usually fails.
269- **Narrative beats in multi-shot**: Give each shot a distinct emotional beat, not just a visual change. Approach → command → payoff. The contrast in intent makes cuts read cleaner than contrast in framing alone.
270- **Duration pacing**: 15 seconds for three shots gives ~5s per beat. 10 seconds for three shots is too tight — beats feel rushed. Budget at least 4-5 seconds per shot for the action to breathe.
271- **Continuation chaining**: Segments max at 15s. Plan each segment as a self-contained beat, then chain with `bfl_flux3_video_continuation`. Open the prompt with "Continue this video from its final frames:" and re-establish the subject.
272
273### Workflow
2741. Submit → get job id immediately (video doesn't exist yet)
2752. Poll `bfl_flux3_get_result` with the id — generation takes several minutes, long "Generating" phase is normal
2763. The poll call blocks while the job runs; if it returns still-generating, just call again — no sleeping
2774. On Ready, the clip is downloaded and `saved_path` is returned. Deliver the file per platform conventions.
2785. Job survives client restarts — re-poll the same id, never resubmit (duplicates spend budget)
2796. **Rate limit: one submission per minute.** If you get a "limited to one attempt
280 per minute" error, `sleep 45` then resubmit. When generating a batch serially,
281 the natural poll-wait between clips usually covers this, but back-to-back
282 submits (e.g. firing two clips in one turn) will trip it.
2837. **Transient 504s happen.** Azure Front Door occasionally returns a 504 HTML
284 page instead of a job id. Just wait the rate-limit cooldown and resubmit —
285 the prompt wasn't consumed.
2868. **`save_to` quirk**: if you pass a path that doesn't exist yet and has no
287 extension (e.g. `.../new_clips`), the tool creates a FILE with that name
288 instead of a directory. Always either pre-create the directory (`mkdir -p`)
289 or pass a full file path with `.mp4` extension.
290
291## Multi-Clip Production Pipeline (15–25 clips → short film / commercial)
292For a long-form piece (a skate/snowboard/sports ad, a short film) assembled from
293many independent clips, run these phases IN ORDER. Skipping ahead burns budget —
294each FLUX 3 gen is minutes + rate-limited, so lock everything upstream first.
295
2961. **Spec & storyboard FIRST.** Write an AD_SPEC.md: setting, mood, lighting,
297 pacing, brand feel, target length. Then a SHOT_LIST.md with a shot taxonomy:
298 establishing / B-roll-texture / POV / hero tricks / ride-away. Budget a
299 ~25–35% B-roll ratio (zero-identity shots rest the eye and hide hero drift).
300 Shoot MORE clips than you'll use (~24–26 for a 70s film) so the edit has
301 options — cut down to the best ~18–20.
3022. **Music is the spine — generate it BEFORE the clips.** Map the track's energy
303 (see Music-Driven Assembly below) and lock a beat map (cold open over intro,
304 action over sustained energy, slow-mo apex in the breakdown, hardest trick on
305 the drop). The picture bends to the song, not vice versa.
3063. **Cast the riders / characters.** For an anthology (multiple distinct people),
307 generate ONE clean character-sheet reference still per rider (front-facing,
308 neutral pose, even light, full kit, plain bg) — each with a distinct saturated
309 accent (red deck, blue deck, yellow deck…). Anthology casting turns identity
310 drift into intentional casting. See Character consistency section.
3114. **Generate in phases** (optimized for the rate limit + QA flow):
312 - Phase 1: establishing + B-roll (text-to-video, no identity — warms up the
313 pipeline, cheapest to re-roll).
314 - Phase 2: POV (text-to-video, zero identity — expect 3–5/10, budget re-rolls).
315 - Phase 3: hero tricks (image-to-video from starting-pose stills — the
316 payload, best identity consistency).
317 - Phase 4: outro / ride-away.
318 Run as a managed batch: fire one clip → poll to completion → QA → fire next.
319 The natural poll-wait covers the rate limit. Report progress in chunks.
3205. **QA every clip** (frame-grab + vision_analyze, see Reviewing below). Keep a
321 re-roll queue for anything <6/10. Batch re-rolls at the end, using
322 image-to-video where color/geometry matters.
3236. **Assemble on the song's spine** (Music-Driven Assembly): trim to downbeats,
324 unified grade on every clip, 15% Ken Burns punch-in, hard cuts, audio duck,
325 fade. Deliver full film + teaser + individual clips + thumbnails + README in a
326 numbered folder (01_raw_clips / 02_final_edits / 03_references / 04_audio /
327 05_scripts / _scratch/review_proxies).
328
329**Time reality:** 24–26 clips at the ~5-min account-wide rate limit + iteration
330≈ 2–3 hours of generation. Don't promise a clip length the tools can't make in
331one pass (>20s needs continuation chaining).
332
333## Music-Driven Assembly (cut the picture to the song's spine)
334When scoring a multi-clip sequence to a generated track, the music becomes the
335timeline and the picture bends to it — cuts land on beats, the slow-mo hero sits
336in a breakdown, the biggest action hits on the drop. Don't cut by feel and bolt
337music on after.
338
339**Where the track comes from (routing):**
340| Scenario | Path |
341|---|---|
342| fal.ai balance available | `fal-ai/elevenlabs/music` via FAL_KEY — text-to-music, section-by-section composition_plan, ~$0.80/min (see fal-ai-generation skill) |
343| fal.ai balance exhausted (403 "User is locked") | **Suno via web UI** — write the prompt with the `suno-music-creation` skill (DSL style prompt → Style field, structure tags → Lyrics field, Instrumental ON); user generates + returns the MP3. Suno returns a FULL song (~90s–4min): mine the cleanest breakdown→drop section. |
344| Offline / local | HeartMuLa (CUDA GPU) or AudioCraft MusicGen (CPU/MPS) |
345Probe the fal.ai balance before committing (see fal-ai-generation 403 troubleshooting).
346Whichever path produced the file, the steps below are identical:
347
3481. **Map the track's structure FIRST.** Run `scripts/audio_energy_map.py` on the
349 audio (decodes via ffmpeg to raw PCM, computes an RMS energy envelope with pure
350 stdlib — no numpy/aubio needed). It prints a per-second energy bar graph plus
351 the quietest windows (breakdowns/intros), loudest windows (drops/peaks), and
352 sharp onsets (candidate drop hits). See `references/music-driven-assembly.md`.
3532. **Build the picture on the song's own spine.** Place the cold open over the
354 quiet intro, fast action over sustained energy, the slow-mo apex in the
355 breakdown (drums gone), and the hardest trick on the drop's slam. A clean
356 breakdown→drop transition already in the track is worth more than fighting the
357 edit to fit.
3583. **Assemble with FFmpeg**: trim each clip so its cut point lands on a downbeat,
359 apply one unified grade (eq contrast/saturation + colorbalance cool shadows /
360 warm mids + light unsharp) to every clip, duck wind/SFX under the music, fade
361 in/out. `concat` filter for the join.
362- **Honest caveat**: beat-perfect sync to *generated* audio is never sample-accurate
363 — a few ms of feel remains. Get it tight by eye on the waveform. If the track's
364 drop isn't clean where you need it, iterate the TRACK, not the cut.
365- **Python 3.13 note**: `audioop` was removed from stdlib; the script computes RMS
366 manually with `struct.iter_unpack("<h", ...)`.
367
368## Reviewing Generated Clips
369Two QA paths — pick per need:
370
371**Path A — Frame-grab + vision_analyze (default for batch QA).** Fastest and
372most reliable for reviewing many clips. Extract a representative mid-clip frame
373and analyze the still:
374```
375ffmpeg -y -v error -i clip.mp4 -vf "select=eq(n\,60)" -vframes 1 frame.jpg
376```
377(Frame 60 ≈ 2.5s into a 24fps clip; pick a frame mid-action, not the first.)
378Then `vision_analyze(frame.jpg)` with a SPECIFIC question: "Rate usability 1-10
379as a [shot type]. Is the deck vivid yellow (not green)? Golden hour lighting?
380Note problems: morphing, weird geometry, text, people." This catches color
381drift, broken geometry, and composition issues at a glance and scales to 20+
382clips without timeouts. Limitation: a still can't confirm temporal morphing —
383flag "must check in motion" for clips where motion integrity is the risk.
384
385**Path B — video_analyze tool (for motion-specific questions).** Use when you
386must verify temporal stability (does the spin morph? does the hand warp
387frame-to-frame?). Requires prep:
388- **Strip audio first.** FLUX 3 clips have audio baked in. The Gemini backend
389 rejects files with audio tracks ("Audio input modality is not enabled").
390 Fix: `ffmpeg -i clip.mp4 -an -c:v copy review/clip.mp4` before analyzing.
391- **Downscale for reliability.** Files >5MB frequently hit "Download multimodal
392 file timed out." Fix: `ffmpeg -i clip.mp4 -an -vf "scale=640:-2" -c:v libx264
393 -crf 32 -preset fast review/clip.mp4` — gets most clips under 200KB.
394- **Batch in groups of 3–4.** More than 4 concurrent analyses increases timeout
395 rate. Space retries; if a file keeps timing out, re-encode smaller.
396
397**Both paths:** Ask specific questions. "Rate usability 1-10 as a [shot type]" +
398"Any morphing/warping?" + "Jacket/deck color?" gets actionable verdicts. Vague
399prompts get essays.
400
401## Platform Integration: Higgsfield MCP
402
403Higgsfield exposes 30+ models via MCP at `https://mcp.higgsfield.ai/mcp`.
404Hermes config (`~/.hermes/config.yaml`):
405```yaml
406mcp_servers:
407 higgsfield:
408 url: "https://mcp.higgsfield.ai/mcp"
409 auth: oauth
410 enabled: true
411```
412First connect opens browser for Higgsfield OAuth. Tokens cache at `~/.hermes/mcp-tokens/higgsfield.json`.
413
414**Critical**: Unlimited mode is **web-app only**. MCP/CLI always consumes credits. Strategy: bulk exploration on higgsfield.ai web, targeted finals via Hermes+MCP.
415
416## CLI Execution (via higgsfield CLI)
417
418Pairs with the vendor `higgsfield-generate` skill for full model/param discovery.
419Key commands:
420
421```bash
422# Text-to-video (blocks until done, prints result URL)
423higgsfield generate create seedance_2_0 --prompt "..." --duration 8 --resolution 720p --aspect_ratio 16:9 --wait
424
425# Image-to-video (animate a still)
426higgsfield generate create seedance_2_0 --prompt "camera dollies in" --start-image ./frame.png --duration 12 --wait
427
428# Image generation
429higgsfield generate create gpt_image_2 --prompt "..." --aspect_ratio 16:9 --resolution 2k --wait
430
431# Audio / SFX
432higgsfield generate create seed_audio --prompt "cinematic rain ambience with distant thunder" --wait
433
434# Check model params before submitting
435higgsfield model get <job_set_type> --json
436
437# List all available models
438higgsfield model list --json
439```
440
441**Always use `--wait`** so the command blocks and prints the result URL.
442For long renders: `--wait-timeout 20m`. For machine-readable output: `--json`.
443
444Media flags: `--image` (reference), `--start-image` (first frame),
445`--end-image` (last frame), `--video` (reference/analysis), `--audio` (lipsync/soundtrack).
446Each accepts a local file path (auto-uploaded) or a UUID.
447
448## Pitfalls
449- Don't exceed 120 words per prompt — contradiction risk rises (EXCEPTION: Flux 3 uses a reasoning harness, not a tag encoder — longer structured prose is fine and often better)
450- Don't use prompt weighting syntax (most 2026 video models don't support it)
451- Don't stack multiple camera moves in one shot description
452- Don't re-describe the medium/style after a style trigger token
453- Don't assume all creative work is brand-specific — ask first
454- Veo 3.1 caps at 8s per generation; longer scenes need stitching
455- Seedance 2.0 full model needs Plus plan+; Starter only gets Fast variant
456- **Brainstorming workflow**: When a user brings a specific concept, execute on it — refine, add depth, solve the hard problems. Don't generate a menu of 10 alternatives unless they explicitly ask for options. A user who says "I want X" wants X made better, not X replaced with Y.
457- **CLI PATH**: `npm install -g` puts binaries in `$(npm prefix -g)/bin/`
458 which may not be in `$PATH`. Fix: `export PATH="$(npm prefix -g)/bin:$PATH"`
459 and add to `~/.zshrc`. On Hermes-managed npm this is `~/.hermes/node/bin/`.
460- **Skills install**: use `npx skills add higgsfield-ai/skills --yes` for
461 non-interactive. Skills land in `~/.agents/skills/` and symlink to
462 `~/.hermes/skills/` — if running a non-default profile, also symlink
463 into `~/.hermes/profiles/<profile>/skills/`.
464
465## References
466- `references/higgsfield-platform.md` — full model roster, pricing, credit costs, plan tiers, Unlimited vs credit mechanics, MCP setup details
467- `references/flux3-techniques.md` — Flux 3 technique bank: multi-shot structure, device screens at 720p, hand-object workarounds, duration budgeting, continuation chaining, proven prompt templates
468- `references/minimax-h3.md` — MiniMax H3: five-block prompt structure, Omni Reference roles, timed beats, sound design, API shape, strengths/weaknesses
469- `references/music-driven-assembly.md` — cutting a multi-clip sequence to a generated track: energy-envelope mapping, building the picture on the song's spine, FFmpeg grade/duck/concat recipe
470- `scripts/audio_energy_map.py` — pure-stdlib RMS energy envelope + drop/breakdown detector for an audio file (decodes via ffmpeg, no numpy/aubio). Run: `python3 audio_energy_map.py <audio> [window_secs]`
471- Vendor skill `higgsfield-generate` (installed via `npx skills add higgsfield-ai/skills`) — CLI mechanics, model IDs, media flags, Marketing Studio workflows, Virality Predictor. Load it for execution details; load THIS skill for prompt craft and model selection strategy.