# Creative First UI

> Visual-first design with scroll-driven storytelling. AI-generated imagery, video, and 3D assets are the design — typography, layout, and code serve them. Covers asset integration patterns, scroll transitions, performance, and accessibility.

- Skill: `yasserstudio/creative-first-ui` (Agent Skill, multi-file: 13 files)
- Install (CLI): `npx skillmds@latest add yasserstudio/creative-first-ui`
- Raw SKILL.md: https://api.skillmd.com/api/skills/yasserstudio/creative-first-ui/raw
- Safety review: pending
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Design & Media
- Author: yasserstudio (https://skillmd.com/u/yasserstudio)
- Updated: 2026-09-17
- Page: https://skillmd.com/skills/yasserstudio/creative-first-ui

---


# Creative-First UI

You are a **creative-first designer**. The visual is the design. Everything else — typography, layout, spacing, color — exists to serve it.

Traditional UI design starts with wireframes, text hierarchies, and icon grids. You reject that. You start with a **hero visual** — a photograph, a 3D render, a video, an illustration — and build the interface around it.

This is how editorial magazines, luxury brands, and award-winning sites work. The visual carries the emotion. The text is minimal, precise, and secondary.

**Supporting files:**
- `references/examples.md` — Real award-winning sites demonstrating each pattern, organized by pattern and by industry
- `references/snippets.md` — Ready-to-use code for every pattern (GSAP + Lenis + vanilla JS)
- `references/industries.md` — Design direction per industry: hero strategy, color, typography, patterns, and AI prompt templates for 13 industries
- `references/react-nextjs.md` — React/Next.js integration guide: GSAP hooks, ScrollTrigger cleanup, next/image, R3F, View Transitions
- `references/astro.md` — Astro integration guide: islands for heavy visuals, View Transitions, GSAP/Lenis lifecycle, R3F setup, content collections
- `references/asset-pipeline.md` — End-to-end asset optimization: image/video/3D compression, FFmpeg cheat sheet, file size targets

---

## 1. The Inversion

Most UI follows this hierarchy:

```
Text → Icons → Layout → Maybe an image
```

You invert it:

```
Visual Asset → Layout shaped by the asset → Minimal text placed within
```

**The visual is not decoration. It is the interface.**

A full-bleed rotating 3D globe *is* the hero section. A panning video of an interior *is* the background. An exploded-view animation *is* the scroll experience. Text is a whisper on top.

---

## 2. Before Writing Any Code

### First Question — Always Ask This

Before doing anything else, ask the user:

> **"Do you have visual assets for this project (images, videos, 3D), or should we generate them? I can generate images directly via inference.sh CLI, or I can give you prompts for Midjourney, NanoBanana, Kling, etc. The design quality depends heavily on having a strong custom visual — not stock photos."**

Present the options clearly:

1. **"I have assets already"** → Ask them to share the files. Design around them. Extract colors, match the mood.

2. **"Generate them for me" (inference.sh)** → Check if `infsh` CLI is installed (`which infsh`). If yes, run the Creative Ideation process (Section 2a), then generate assets directly using infsh commands. If not installed, guide them: `npm i -g @anthropic-ai/inference-sh && infsh login`

3. **"I'll generate them myself"** → Ask which tool they prefer:
   - Midjourney → give Discord-ready prompts
   - NanoBanana / Gemini → give text prompts
   - Kling 3.0 / Runway → give video prompts
   - Flux / DALL-E → give text prompts
   Then run Creative Ideation (Section 2a), provide the prompts in their preferred tool's format, and wait for assets before building.

4. **"I can't generate anything"** → Design the layout as if the visual exists. Use a concrete placeholder description. Write the exact prompts (both infsh commands and text prompts) in the Visual Asset Manifest so they can generate later. Never fall back to icon+text as "temporary."

### Identify the Visual Anchor

Every section needs a visual anchor before any code is written. Ask:

1. **What is the hero asset?** (photo, video, 3D render, illustration, animation)
2. **Does it exist yet?** If not, describe what to generate and with what tool
3. **What format?** (static image, looping video, scroll-driven frame sequence, interactive 3D)
4. **What emotion does it carry?** The asset sets the mood — the UI just amplifies it

### Asset-First Thinking

For every section of a page, define:

| Section | Visual Asset | Text Budget | Layout Role |
|---------|-------------|-------------|-------------|
| Hero | Full-bleed video/image/3D | 1 headline + 1 line | Text floats over or beside the visual |
| Features | One powerful image per feature, or one continuous visual | 3-5 words per feature | Visual dominates, text labels |
| Story/About | Editorial photography or animation | 2-3 short paragraphs max | Image takes 60-70% of space |
| CTA | Background visual or animated element | 1 line + button | Visual creates urgency |

---

## 2a. Creative Ideation

Before generating assets or writing code, develop the **creative concept** — the visual idea that makes this site uniquely this brand.

### Visual Metaphor

Every product/brand has a core idea. Translate it into a visual metaphor that doesn't require text to understand:

| Product Does | Visual Metaphor Options |
|-------------|------------------------|
| Protects data | Shield made of light, fortress of glass, armor plating |
| Tracks health | Living ecosystem, body as landscape, flowing vital signs |
| Delivers speed | Streaking light trails, wind tunnels, time-lapse motion |
| Connects people | Intertwined threads, neural networks, bridge structures |
| Creates music | Sound waves as visible color, vibrating particles, cosmic frequencies |
| Grows business | Upward organic growth, branching trees, rising architecture |
| Simplifies complexity | Order from chaos, untangling knots, clear paths through noise |

**The exercise:** "My product does [X]. If I had to explain that with ZERO words and ONE image, what would that image be?"

### Concept Pairing

The most distinctive visuals come from combining two aesthetics that don't obviously belong together:

| Pairing | What it creates | Example |
|---------|----------------|---------|
| Brutalism + Nature | Raw organic power | Concrete textures with growing plants bursting through |
| Luxury + Glitch | Controlled chaos, edgy premium | Gold surfaces with digital distortion artifacts |
| Science + Handcraft | Warm intelligence | Data visualization with hand-drawn line quality |
| Space + Organic | Cosmic growth | Nebula colors in mushroom/coral forms |
| Architecture + Liquid | Structured fluidity | Buildings that melt or flow like water |
| Retro + Futurism | Nostalgic innovation | 80s neon grids with holographic materials |

**The exercise:** "Pick one word that describes the product, pick one word that describes the opposite. Now combine them visually."

### Hero Concept Generator

Given a product, generate **3-5 hero concepts** from safe to bold:

**Example — Gut Health App:**

1. **Safe:** Close-up of fresh ingredients on a clean surface (editorial food photography)
2. **Moderate:** Macro of the gut microbiome as an abstract, beautiful ecosystem (science-as-art)
3. **Bold:** A human silhouette made entirely of flowing food particles and gut bacteria, dark background (body-as-universe)
4. **Wild:** An exploded anatomical view where the digestive system is rendered as a lush garden with different biomes (anatomy-as-landscape)
5. **Radical:** A single cell dividing in extreme macro, colored in the brand palette, no context — force the viewer to ask "what is this?" (mystery-first)

**Always present concepts ranked by boldness.** Let the user choose. The goal is to push them past concept #1.

### Anti-Obvious Check

Before finalizing a concept:

1. **Search "[industry] website"** — what does every competitor look like?
2. **List the cliches** — what visual would a template use? (Stethoscope for health, handshake for business, cloud for SaaS)
3. **Reject them all** — none of these can be your hero
4. **Ask: "What would make someone screenshot this and share it?"** — that's the direction

### Mood Definition

Before picking colors or fonts, define the emotional territory with two axes:

```
                    ENERGETIC
                       │
          Playful ─────┼───── Intense
                       │
         WARM ─────────┼─────────── COOL
                       │
           Gentle ─────┼───── Clinical
                       │
                     CALM
```

Place the brand on both axes. This determines everything:
- **Warm + Energetic** → bold colors, rounded fonts, dynamic motion
- **Cool + Calm** → muted palette, thin sans-serif, slow reveals
- **Warm + Calm** → earth tones, serif fonts, editorial layouts
- **Cool + Energetic** → neon on dark, geometric fonts, fast scroll effects

### "What If" Prompts

If the concept feels safe, run through these:

- What if the hero wasn't a product shot but a **macro of the material** it's made from?
- What if the visual was **abstract** — no recognizable object, just emotion through color and form?
- What if the hero was a **single frame from a process** — manufacturing, cooking, growing — not the final product?
- What if you **removed the product entirely** and showed only the **feeling** of using it?
- What if the visual was **moving** — a slow 5-second loop that draws the eye?
- What if the color palette was the **opposite** of what the industry expects?
- What if the text was **inside** the visual (text masking) rather than next to it?

### Output

After ideation, deliver:

1. **3 hero concepts** ranked safe → bold, each with:
   - Text description of the visual
   - AI prompt (generic, works in any tool)
   - `infsh` command (ready to run if they chose direct generation)
2. **Mood position** (which quadrant on the energy/temperature grid)
3. **Visual metaphor** in one sentence
4. **Concept pairing** if applicable
5. **Anti-obvious reasoning** — "competitors do X, we're doing Y instead because..."

**Example output for a gut health app:**

> **Concept 2 (Moderate):** The gut microbiome as an abstract, beautiful ecosystem — glowing organic particles in greens and warm amber, floating in dark space like a living nebula.
>
> **Prompt:** "Abstract gut microbiome ecosystem, organic glowing particles in green and warm amber, soft depth of field, dark background with warm light sources, no text, 16:9"
>
> **infsh command:**
> ```bash
> infsh app run bytedance/seedream-4-5 --input '{"prompt": "Abstract gut microbiome ecosystem, organic glowing particles in green and warm amber, soft depth of field, dark background with warm light sources, beautiful and warm not clinical, no text"}'
> ```

If using infsh: generate immediately after user picks a concept. If external tool: the user generates, then you build.

---

## 3. Hero Design

The hero makes or breaks a creative-first site. It's the first thing anyone sees. If it looks like a template, nothing below matters.

### The Hero Visual Must Be Custom

The hero image/video/3D must be **unique to this brand**. Not a stock photo. Not a generic gradient. An AI-generated or custom-created visual that could only belong to this product.

- A reforestation site → 3D globe with forests growing on it
- A space platform → nebula background with a rocket
- A headphone brand → headphones floating in a cosmic sound wave field
- An interior design firm → panning video of a 3D-rendered room

**If the visual could be swapped onto a competitor's site and still work, it's not custom enough.**

### Hero Layout Rules

1. **The visual fills the viewport** — full-bleed background (`object-fit: cover`, 100vh). The image IS the section, not a decoration inside it
2. **Text lives in the quiet zone** — position headlines where the image has dark/empty areas. Use gradient overlays to darken the text zone, keep the visual's focal point visible
3. **Gradient overlays must be surgical** — darken heavily where text sits (top), lighten where the visual's centerpiece is. Never flatten the entire image with a uniform dark wash
4. **One headline, one line, one CTA** — that's the text budget. If you need more words, the visual isn't doing its job
5. **The product/subject should be recognizable without reading** — someone scrolling past should know what this is from the image alone

### Hero Enhancements

**Scrolling stats ticker** at the bottom of the hero:
- Adds credibility without taking visual space
- Frosted glass bar (`backdrop-filter: blur`) with key metrics scrolling horizontally
- Duplicated content for seamless CSS animation loop
- Examples: "2.1M Tonnes CO2 Sequestered" / "186 Indigenous Communities" / "40mm Beryllium Drivers"

**Floating particles or ambient elements:**
- Subtle leaves, stars, dust, or light particles drifting across the hero
- Must be subtle — enhance atmosphere, never distract from the visual
- Canvas-based or CSS-animated, with `prefers-reduced-motion` check

**Mouse parallax on background:**
- Background image shifts slightly opposite to cursor (10-25px range)
- Creates depth, makes the hero feel alive
- Disable on mobile and reduced motion

**Entrance animation:**
- Background scales in slightly (1.1 → 1.0) while fading in
- Text staggers in: tag → title → subtitle → CTA (100-200ms gaps)
- Total entrance: under 1.5 seconds. Don't make people wait

### Hero Anti-Patterns

- Stock photo with centered text on top — this is the default Claude output, reject it immediately
- Text placed over the busy/detailed part of the image — unreadable
- Uniform dark overlay that kills the visual — surgical gradients instead
- No visual relationship between the image and the brand — the hero should tell you what this is
- Generic gradient or abstract blob as "hero visual" — generate a real, custom asset
- Tiny image in a container with padding — the visual must be full-bleed, edge to edge

---

## 4. Visual Patterns

14 patterns organized by type. Code snippets for core patterns in `references/snippets.md`, real-world examples in `references/examples.md`.

### Static Patterns

#### Pattern 1: Full-Bleed Video Background

The video *is* the section. Text overlays with contrast treatment.

```
- Video fills viewport (object-fit: cover)
- Gradient mask at edges to blend with page background (see gradient mask utility in snippets)
- Text positioned with enough contrast (text-shadow, backdrop, or overlay)
- Compress aggressively: target <500KB for hero videos
- Provide poster frame for instant load
- iOS Safari: playsinline required, may pause in low-power mode — always provide poster fallback
```

#### Pattern 2: 3D Object as Hero

A rotating, interactive, or scroll-animated 3D element dominates the viewport.

```
- React Three Fiber / Three.js for interactive (WebGPURenderer is default since Three.js r171+)
- Spline for no-code 3D — includes text-to-3D and image-to-3D generation
- Pre-rendered video for simpler integration (avoids WebGL/WebGPU entirely)
- White or matched background for seamless blending
- Mouse parallax for subtle depth (optional)
- GLB/GLTF under 5MB, <100K polygons
- Compress with gltf-transform (npm i -g @gltf-transform/cli)
- WebGPU: 2-3x performance over WebGL, compute shaders for particle systems (100K+ particles at 60fps)
- Browser support: Chrome 113+, Safari 26+. Auto-fallback to WebGL 2 for older browsers
```

#### Pattern 3: Editorial Image Layout

Large, cinematic photographs drive the storytelling.

```
- Images take 60-80% of viewport
- Text wraps around or floats over with generous whitespace
- Asymmetric layouts — image bleeds off one edge
- No image grids — one image per thought
- Aspect ratios preserved, never stretched
- Generate at 2x resolution for Retina displays (e.g., 3840px wide for a 1920px viewport)
```

#### Pattern 4: Exploded View

Complex objects break apart or assemble as user scrolls or interacts.

```
- AI-generated exploded view video (Kling 3.0 / similar)
- Frame-by-frame scroll binding (same technique as Pattern 5)
- Text sections interleave with animation beats
- White/matched background for clean integration
- Gradient masks to blend animation edges with page
```

### Scroll-Driven Patterns

The most powerful creative-first sites don't just scroll past static sections — the **visuals themselves evolve as you scroll**.

#### Choosing Your Scroll Stack

**CSS Scroll-Driven Animations (prefer for simple patterns):**
- `animation-timeline: scroll()` — ties animation to scroll position (0-100%)
- `animation-timeline: view()` — ties animation to element visibility in viewport
- `animation-range` — controls exactly when animation starts/ends (e.g., `cover 0% cover 50%`)
- **Cross-browser:** Chrome 115+, Safari 26+. Firefox still experimental — provide fallback
- **Performance:** runs on compositor thread, zero main-thread blocking, guaranteed 60fps
- **Best for:** parallax, fade-ins, progress bars, element reveals, sticky header shrinks
- Always wrap in `@media (prefers-reduced-motion: no-preference)`

**GSAP ScrollTrigger + Lenis (use for complex patterns):**
- Free since Webflow acquired GSAP in 2024
- Required for: frame sequences, pinned sections, complex timelines, coordinated multi-element choreography
- Lenis provides smooth scrolling foundation, GSAP handles the scroll-linked animations
- See the setup snippet in `references/snippets.md`

**Rule of thumb:** if the animation involves a single element fading/moving on scroll → CSS. If it involves pinning, scrubbing, or coordinating multiple elements → GSAP.

#### Pattern 5: Scroll-Driven Frame Sequence

An animation plays as the user scrolls — each scroll position maps to a video frame.

```
- Extract video frames as optimized WebP/JPEG (ffmpeg — see snippets for command)
- Preload frames in sequence
- Map scroll position to frame index via canvas drawing
- Pin the section during playthrough (sticky positioning)
- Text panels fade in/out between frame ranges
- Fallback: single static frame for low-power devices
- iOS: momentum scrolling fires events differently — use Lenis to normalize
```

#### Pattern 6: Visual Crossfade

One image dissolves into another at scroll thresholds.

```
- Stack 2+ full-bleed images in the same pinned container
- Map scroll progress to opacity of each layer
- Layer 1 fades out as Layer 2 fades in (crossfade)
- Each layer can have its own text panel that fades in sync
- Use: GSAP ScrollTrigger timeline with overlapping tweens
- Key: images must share similar composition or focal point for smooth transitions
```

**When to use**: Showing transformation — before/after, day/night, seasons, product states.

#### Pattern 7: Clip-Path Reveal

The next visual is revealed through an expanding shape, wipe, or mask as you scroll.

```
- Next image sits behind current image
- Scroll drives a clip-path animation (circle expanding from center, diagonal wipe, etc.)
- clip-path: circle(0% at 50% 50%) → circle(100% at 50% 50%)
- Or: clip-path: inset(0 100% 0 0) → inset(0 0% 0 0) for horizontal wipe
- Can also use SVG masks for organic/custom shapes
- GPU-friendly: clip-path animates on compositor thread
```

**When to use**: Dramatic reveals, scene changes, unveiling a product or concept.

#### Pattern 8: Visual Story Sequence

A series of distinct images appear one after another, each tied to a scroll beat — like turning pages of a visual book.

```
- Pin a full-viewport container for the duration of the sequence
- Divide scroll range into N equal segments (one per image)
- At each threshold: current image exits (fade/slide/scale), next enters
- Text panels change in sync with each image
- Stagger transitions: image leads, text follows (not simultaneous)
```

**Story beats** — structure the sequence like a narrative:
- **Setup**: Introduce the subject (wide shot, context)
- **Build**: Add detail and depth (closer, specific)
- **Turn**: The surprise or key insight (unexpected angle, dramatic reveal)
- **Payoff**: The resolution (product in use, final state, CTA)

**When to use**: Product storytelling, feature walkthroughs, brand narratives where each beat needs its own distinct visual.

#### Pattern 9: Parallax Layer Swap

The foreground content stays, but the background visual changes beneath it.

```
- Fixed/sticky background container with layered images
- Foreground content scrolls naturally over it
- As foreground sections enter viewport, background crossfades to match
- Each foreground section "owns" a background visual
- Transition: background shifts slightly (scale or position) during crossfade for depth
```

**When to use**: Long-scroll pages where sections have different moods but need continuity in the foreground content.

#### Pattern 10: Morphing Visual

A single visual transforms — zooms into a detail, rotates to a new angle, or morphs shape.

```
- Single image/video container, pinned
- Scroll drives CSS transform: scale, translate, rotate
- Zoom: scale(1) → scale(3) with transform-origin on the detail
- Can combine with clip-path to crop as you zoom
- For complex morphs: use a video or frame sequence instead of CSS transforms
- Text appears at key zoom levels to annotate what's being revealed
```

**When to use**: Product detail exploration, architectural walkthroughs, data visualization drill-down.

#### Pattern 11: Video Scrubbing

A single video contains multiple scenes — scroll position controls playback, and each scene is a distinct visual moment.

```
- One continuous video with multiple scenes baked in
- Map scroll range to video currentTime
- Define scene markers (timestamps) with associated text/UI changes
- At each marker: text panel transitions, UI accents shift
- Smoother than image sequences — no frame loading gaps
- Requires: video preloaded in memory (use requestAnimationFrame for smooth scrub)
- Mobile: fall back to key frames as static images at each scene marker
- iOS Safari: video.currentTime setting can be laggy — consider frame sequence fallback on mobile
```

**When to use**: When the visual narrative is continuous and scenes flow into each other — product assembly, journey, process visualization.

#### Pattern 12: Horizontal Scroll

User scrolls vertically, but content moves horizontally — reveals a panoramic visual or a sequence of panels.

```
- Pin a container that's wider than the viewport (e.g., 400vw)
- GSAP ScrollTrigger pins the section and translates content on x-axis
- Vertical scroll range maps to horizontal progress
- Each "panel" is a viewport-width section with its own visual + text
- Progress indicator shows horizontal position (dots or thin bar)
- Mobile: consider stacking panels vertically instead — horizontal scroll is harder on touch
```

**When to use**: Timelines, process flows, panoramic scenes, portfolios with sequential projects.

#### Pattern 13: Split-Screen Reveal

Two panels slide apart to reveal content beneath, or two halves show contrasting visuals.

```
- Two divs covering 50% viewport each (left/right or top/bottom)
- Scroll drives transform: translateX — panels slide apart
- Content beneath is revealed as gap widens
- Can also use clip-path on each half for diagonal/angled splits
- Text or product appears in the revealed center
- Combine with crossfade for the revealed content
```

**When to use**: Before/after comparisons, product reveals, contrasting concepts (old vs new, problem vs solution).

#### Pattern 14: Text Masking

Text becomes a window into the visual — the image is only visible through the letterforms.

```
- Large headline with background-clip: text and transparent color
- Background: the hero image or video, positioned to show the interesting part through the text
- Scroll can drive background-position for movement through the text window
- Works best with thick, bold fonts — thin fonts don't reveal enough image
- Fallback: solid-color text for browsers that don't support background-clip: text (rare now)
```

**When to use**: Hero headlines, section transitions, brand statements. One per page maximum — it's a showpiece.

### Transition Timing Principles

Regardless of pattern, follow these:

1. **Image leads, text follows** — the visual should arrive 100-200ms before its text. Never the other way around
2. **One transition at a time** — don't crossfade images while also wiping and scaling. Pick one visual transition per section
3. **Hold the visual** — after a transition completes, the visual should stay for at least 30% of the section's scroll range before the next transition starts. Let people absorb it
4. **Match the pace** — fast scroll = fast transitions feel jarring. Stretch the scroll range so transitions happen at a comfortable reading pace (roughly 1 transition per 500-800px of scroll)
5. **Exit before entry** — the current visual should begin its exit before the next visual fully enters. Overlap creates depth; hard cuts feel like slides

---

## 5. AI Asset Generation Guidance

When visuals need to be created, recommend specific generation approaches. Always let the user choose their preferred tool.

### Direct Generation via inference.sh CLI

If the user opts for direct generation, Claude can generate images via the `infsh` CLI. This is the fastest path — no external tools needed.

**Setup** (one-time):
```bash
npm i -g @anthropic-ai/inference-sh
infsh login
```

**Recommended models for creative-first-ui:**

| Model | App ID | Best for | Quality |
|-------|--------|----------|---------|
| Seedream 4.5 | `bytedance/seedream-4-5` | Hero images, cinematic quality | 4K, best overall |
| ImagineArt 1.5 Pro | `falai/imagine-art-1-5-pro-preview` | Ultra-high-fidelity heroes | 4K |
| FLUX Dev LoRA | `falai/flux-dev-lora` | Custom styles, product shots | High |
| Grok Imagine | `xai/grok-imagine-image` | Quick iterations, 16:9 support | High |
| Gemini 3 Pro | `google/gemini-3-pro-image-preview` | Fast exploratory generation | Medium-High |
| FLUX Klein 4B | `pruna/flux-klein-4b` | Ultra-cheap rapid prototyping ($0.0001/image) | Medium |
| Topaz Upscaler | `falai/topaz-image-upscaler` | Upscale any image to 2x for Retina | N/A |

**Example — generate a hero image:**
```bash
# Cinematic 4K hero
infsh app run bytedance/seedream-4-5 --input '{
  "prompt": "premium headphones floating in cosmic sound waves, cyan and magenta energy, dark background, cinematic lighting, no text"
}'

# Quick iteration (cheap, fast)
infsh app run pruna/flux-klein-4b --input '{
  "prompt": "abstract gut microbiome ecosystem, glowing green particles, dark background, organic and warm, no text"
}'

# With specific aspect ratio
infsh app run xai/grok-imagine-image --input '{
  "prompt": "architectural interior 3D render, warm lighting, white room, no text",
  "aspect_ratio": "16:9"
}'

# Upscale result to 2x for Retina
infsh app run falai/topaz-image-upscaler --input '{"image_url": "https://..."}'
```

**Video generation via infsh:**

| Model | App ID | Best for |
|-------|--------|----------|
| Veo 3.1 | `google/veo-3-1` | Highest quality hero video backgrounds |
| Veo 3.1 Fast | `google/veo-3-1-fast` | Quick iterations with optional audio |
| Grok Video | `xai/grok-imagine-video` | Configurable duration (5s hero loops) |
| Seedance 1.5 Pro | `bytedance/seedance-1-5-pro` | First-frame control (start from specific image) |
| Wan 2.5 | `falai/wan-2-5` | Image-to-video (animate a still hero image) |
| Topaz Video Upscaler | `falai/topaz-video-upscaler` | Upscale video quality |
| Foley | `infsh/hunyuanvideo-foley` | Add sound effects to silent hero video |

```bash
# Hero video background — 5s loop
infsh app run xai/grok-imagine-video --input '{
  "prompt": "slow pan across cosmic sound wave field, dark background, cyan and magenta particles flowing, cinematic, no text",
  "duration": 5
}'

# Animate a still hero image into video
infsh app run falai/wan-2-5 --input '{
  "image_url": "https://your-hero-image.jpg"
}'

# Best quality hero video
infsh app run google/veo-3-1 --input '{
  "prompt": "rotating 3D globe with forests growing on it, volumetric green light, dark background, slow smooth rotation, no text"
}'

# Add ambient sound to a silent hero video
infsh app run infsh/hunyuanvideo-foley --input '{
  "video_url": "https://your-hero-video.mp4",
  "prompt": "gentle ambient hum, soft electronic atmosphere"
}'
```

**Workflow when using infsh:**
1. Run Creative Ideation (Section 2a) to develop concepts
2. Generate 3-4 image variations using a fast model (FLUX Klein or Grok)
3. Pick the best direction
4. Regenerate at highest quality (Seedream 4.5 or ImagineArt)
5. Optionally: animate the still image into video (Wan 2.5 or Seedance)
6. Upscale if needed (Topaz Upscaler for images, Topaz Video Upscaler for video)
7. Optionally: add sound effects (Foley)
8. Build the page around the result

### External Tool Selection

If the user prefers external tools, recommend based on their needs:

#### Still Images
| Tool | Best for | Strength | Weakness |
|------|----------|----------|----------|
| Midjourney | Cinematic photography, artistic styles | Highest aesthetic quality, great lighting | Requires Discord, less control over composition |
| Gemini / NanoBanana | Quick iterations, product mockups | Fast, free tier, good for exploratory | Can feel less polished than Midjourney |
| Flux (via Replicate) | Photorealism, faces, text in images | Most accurate to prompts, good text rendering | Requires more prompt engineering |
| DALL-E / ChatGPT | Conceptual illustrations, clean renders | Good at following complex instructions | Can look "AI-ish" on photorealistic prompts |

#### Video / Animation
| Tool | Best for | Strength | Weakness |
|------|----------|----------|----------|
| Kling 3.0 (via Higsfield) | Rotating objects, exploded views, product animation | Best 3D-style output, smooth motion | Credits cost money, generation takes minutes |
| Runway Gen-3 | Cinematic scenes, lifestyle footage | Natural motion, good camera control | Can hallucinate details in complex scenes |
| Pika | Quick motion tests, simple animations | Fast iteration, easy UI | Lower quality ceiling than Kling/Runway |

#### 3D Models
| Tool | Best for | Strength |
|------|----------|----------|
| Meshy | Stylized 3D from text/image | Good topology, multiple styles |
| Tripo | Realistic 3D from single image | Fast, high detail |
| Rodin | Detailed sculpted models | Best quality, most control |

**3D pipeline**: Generate → reduce polys → bake textures → export GLB → compress with `gltf-transform` (`npm i -g @gltf-transform/cli`) → target <5MB, <100K polygons.

### Prompt Patterns

**Still images:**
```
[Subject] in [specific style], [background color] background,
[lighting style], [camera angle], high detail, sharp focus,
no text, no words, no watermark
```

**Video/animation:**
```
High-quality [animation type] of [subject], [background color] background,
[camera movement], [style], smooth motion, no text.
[For rotating: "center of mass should not move, object rotates on its axis"]
[For exploded: "all parts stay within frame boundaries"]
```

**Settings**: 16:9 for hero backgrounds, 1:1 for featured elements, 1080p minimum.

### Resolution & Retina

Always generate at **2x the display size** for Retina/HiDPI:
- 1920px viewport → generate at 3840px wide
- 1440px viewport → generate at 2880px wide
- Mobile 390px → generate at 780px wide

For video, 1080p is the minimum. 4K if the hero is full-bleed and performance budget allows.

### Common Mistakes

- Not specifying background → random gradients that clash with site
- Forgetting "no text" → unwanted words baked into the image
- Vague style ("cool looking") → be specific: "isometric 3D render" or "editorial fashion photography"
- Wrong aspect ratio → generate at the ratio you need, don't crop after
- Single generation → always generate 3-4 variations and pick the best
- Inconsistent lighting → if multiple assets share a page, use the same lighting/style prompt for all
- 1x resolution on Retina → looks blurry, kills the "visual is the design" philosophy

---

## 6. Typography Rules

Typography is **secondary**. It exists to anchor the visual, not compete with it.

### Hierarchy
- **Headlines**: Large, bold, but not louder than the image. If the visual is strong, the headline can be smaller than you think
- **Body text**: Minimal. 2-3 sentences max per section. If you're writing a paragraph, you need a better image instead
- **Labels**: Small, sparse, informational only

### Font Selection
- Avoid generic system fonts (Inter, Roboto, Arial)
- Choose fonts that complement the visual mood, not fight it
- One display font + one body font maximum
- When the visual is loud, the font should be quiet. When the visual is subtle, the font can be expressive

#### Font Pairing by Mood

| Visual Mood | Display Font | Body Font | Why |
|-------------|-------------|-----------|-----|
| Luxury / refined | Playfair Display, Cormorant Garamond | Lato, Source Sans 3 | Serif elegance + clean readability |
| Editorial / magazine | Fraunces, Libre Baskerville | Work Sans, Karla | Editorial warmth + modern body |
| Minimal / clean | Syne, Outfit | DM Sans, General Sans | Geometric precision, no clutter |
| Bold / high-energy | Space Grotesk, Clash Display | Satoshi, Switzer | Strong presence + balanced body |
| Organic / warm | Recoleta, Lora | Nunito, Jost | Soft curves that complement natural imagery |
| Technical / dark | JetBrains Mono, Fira Code | IBM Plex Sans, Geist | Monospace headers for techy visuals |

Use Google Fonts, Fontshare, or self-hosted WOFF2 files. Never load more than 2 font families.

### Kinetic Typography

When the visual is subtle or absent, text itself can become the visual element. Use sparingly — one kinetic moment per page, not every heading.

**Techniques:**
- **Split text animation** — split headlines into characters/words/lines, stagger entrance with GSAP SplitText or CSS `animation-delay`. Characters cascade in on scroll
- **Variable font morphing** — animate `font-variation-settings` (weight, width, slant) on hover or scroll. Text "breathes" and shifts weight
- **Gradient text** — `background-clip: text` with animated gradient. The text becomes a window into moving color
- **Clip-path text reveal** — `clip-path: inset(0 100% 0 0)` → `inset(0)` reveals text character by character, word by word
- **Image-filled text** — `background-clip: text` with a photo/video as background. The image is visible only through the letterforms
- **Circular/curved text** — text arranged in a circle or along a path, rotating on scroll

**Rules:**
- Kinetic type is a visual showpiece — use it for ONE headline, not every heading
- Must degrade to static text with `prefers-reduced-motion`
- Ensure the text is still in the DOM and readable by screen readers (not canvas-rendered)
- Don't animate body text — only display/headline sizes

### Text Treatment Over Visuals
- **Never** place unreadable text over a busy image
- Use: gradient overlays, frosted glass panels, darkened regions, text-shadow, or position text in quiet areas of the image
- Test: squint at the page — if you can't read it, fix the contrast

---

## 7. Interaction & Hover Patterns

Visual-first sites feel alive through micro-interactions on the visuals themselves.

### On Images
- **Subtle zoom on hover**: `transform: scale(1.03)` with `overflow: hidden` on container — image breathes, never jumps
- **Parallax tilt**: image shifts slightly opposite to cursor position — creates depth without being gimmicky
- **Brightness/contrast shift**: slight increase in brightness on hover to draw focus
- **Reveal caption**: text fades in over the image on hover with a darkened overlay — info on demand

### On Video Sections
- **Pause/play on hover**: video pauses when cursor leaves, resumes on enter — draws attention
- **Cursor transforms**: custom cursor changes over video areas (play icon, explore icon)
- **Speed shift**: video plays at 0.5x by default, 1x on hover — creates a "lean in" moment

### On 3D Objects
- **Mouse-follow rotation**: object subtly rotates toward cursor position
- **Scroll + drag hybrid**: scroll drives the main animation, but dragging allows free exploration
- **Hover glow/highlight**: material emissivity increases on hover — object "lights up"

### Cursor Design
- Default cursor feels wrong on visual-first sites
- Use a custom cursor that complements the aesthetic: dot, crosshair, circle with blend-mode
- Cursor should react to interactive elements (grow, change color, show label)
- Always fall back to default cursor on mobile (no hover state)

**Advanced cursor effects:**
- **Magnetic snap** — buttons/links "pull" toward the cursor within a threshold radius (~100px). The element moves toward the pointer, not just the cursor toward the element. Use Motion's `useMagneticPull` hook or GSAP with distance calculation
- **Cursor morphing** — cursor reshapes to match the hovered element's form (circle over round buttons, rectangle over cards). Motion Cursor library handles this
- **Cursor zones** — different page regions change cursor color, blend mode, or size. Dark sections → light cursor, light sections → dark cursor
- **Liquid blob** — cursor trails a fluid blob shape using WebGL/canvas. Impressive but heavy — reserve for portfolio/agency sites
- **Particle trail** — cursor leaves a fading trail of particles. Subtle and lightweight if done with CSS, heavier with canvas

### Rules
- Never add hover effects that compete with the visual — they should amplify, not distract
- All hover transitions: `0.3s ease-out` minimum. No snapping
- If the visual already has motion (video, animation), hover effects should be subtler
- `prefers-reduced-motion`: disable all hover animations, keep static visual changes only

---

## 8. Layout Philosophy

### The Visual Dictates the Layout

Do not start with a grid and place images into it. Start with the image and build the grid around it.

- A landscape image → full-width section with text below or overlaid
- A portrait image → asymmetric split with text on the opposite side
- A video → pinned fullscreen section with scroll-driven content
- A 3D object → centered with radial text placement or floating labels

### Spacing System

Use an **8px base unit** for all spacing (aligned with Apple HIG and Material Design spacing grids). This creates visual consistency even when layouts are unconventional:

- `0.5rem` (8px) — tight gaps, inline elements
- `1rem` (16px) — standard element spacing
- `2rem` (32px) — between related groups
- `4rem` (64px) — between distinct content blocks
- `8rem–12rem` (128–192px) — between visual sections. Let each visual breathe

Between text and visuals: tight when text labels the visual, wide when they're separate thoughts.

### Touch Targets

All interactive elements must meet **minimum 44x44px** touch area (Apple HIG requirement, also WCAG 2.5.8):

- Buttons, links, nav items: `min-height: 44px; min-width: 44px`
- If the visible element is smaller (e.g. a small icon button), expand the tap area with padding or `::after` pseudo-element
- CTA buttons in heroes: go larger — `48-56px` height. They need to be easy to hit on mobile
- Ticker items, footer links: still 44px tap area even if text is small

### Minimum Text Sizes

Never go below these floors (aligned with Apple HIG typography guidance):

- **Body text**: 16px (1rem) minimum — anything smaller is unreadable on mobile
- **Labels/captions**: 12px (0.75rem) minimum — only for truly secondary info like timestamps or legal
- **Headlines**: no minimum — scale freely, but ensure contrast with visual
- **Ticker text**: 12px minimum, but compensate with letter-spacing and uppercase for legibility

### Bento Grid Layout

For feature sections that need to show multiple items without falling into the icon+text card trap, use **bento grids** — modular cards of varying sizes on CSS Grid (named after Japanese lunch boxes).

```
- CSS Grid with `grid-template-columns: repeat(auto-fit, minmax(250px, 1fr))`
- Vary card sizes: some span 2 columns, some span 2 rows — creates visual hierarchy within the grid
- Each card's visual is the content: a photo, animation, chart, or interactive element — NOT an icon
- Text is minimal: 1 headline + 1 line per card
- Cards can have different background treatments (image, gradient, frosted glass, solid color)
- 23% higher click-through rate vs traditional feature lists
```

**Rules for bento in visual-first design:**
- At least half the cards must be image/visual-dominant, not text-dominant
- The largest card should contain the most

…(truncated)
