# Cinematography

> Cinematography for generated video

- Skill: `raphaelbgr/cinematography` (Agent Skill)
- Install (CLI): `npx skillmds@latest add raphaelbgr/cinematography`
- Raw SKILL.md: https://api.skillmd.com/api/skills/raphaelbgr/cinematography/raw
- Safety review: pending
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Coding & Dev Tools
- Author: raphaelbgr (https://skillmd.com/u/raphaelbgr)
- Updated: 2026-09-17
- Page: https://skillmd.com/skills/raphaelbgr/cinematography

---


# Cinematography for generated video

The camera and lighting decisions, with **what each one does to the viewer** — because a
video model will happily render a technically correct shot that lands emotionally wrong.

Full menus (shot sizes, angles, all camera moves, optics, composition, lighting setups,
palettes, film stocks, styles):
[`prompt-expander/references/visual-language.md`](../prompt-expander/references/visual-language.md).
This skill is the decision layer on top of them.

---

## 1. Shot, lens and aperture by purpose

Pick the row that matches the *job* of the shot, not the one that sounds most cinematic.

| Purpose | Shot | Lens | Aperture |
|---|---|---|---|
| Hook / emotion | Close-up | 85mm | f/1.8 |
| Peak intensity | Extreme close-up | 85-135mm | f/2.0 |
| Dialogue | Medium | 50mm | f/2.8 |
| Reveal / context | Medium wide | 35mm | f/4.0–5.6 |
| Two people | Two-shot | 50mm | f/3.5 |
| POV / environment | POV | 35mm | f/5.6 |
| Over-the-shoulder | OTS | 50mm | f/2.5 |

**A wide shot has two jobs, and the frame decides which.** Subject large in frame with
the space around them reads as *context*. Subject small and swallowed by the space reads
as *isolation*. Both are wides; only the second one is lonely. Say which you want.

**Lens to feel:** ultra-wide (≤24mm) expands space and exaggerates depth · wide (28–35mm)
holds subject *and* environment · normal (50mm) is human-eye natural · telephoto (85mm+)
compresses depth and flatters portraits.

**Depth of field:**

| Aperture | Effect | Use |
|---|---|---|
| f/1.8 | Very shallow, creamy bokeh | Close-ups, emotion |
| f/2.0–2.8 | Shallow, subject isolated | Dialogue, reactions |
| f/3.5–4.0 | Moderate, background reads | Two-shots, walking |
| f/5.6 | Deep, background visible | Reveals, environment |

---

## 2. Camera movement — pick for the effect, name exactly one

| Movement | Effect | Use case |
|---|---|---|
| Slow push-in | Building intensity | Hook, punchline |
| Tracking (parallel) | Energy, momentum | Walking scenes |
| Static handheld | Documentary honesty | Interviews, reactions |
| Handheld | Urgency | Someone moving with purpose |
| Shaky camera | Chaos | The moment control is lost |
| Dolly zoom | Shock | The ground going out from under someone |
| POV push-in | Curiosity, leaning forward | Observing chaos |
| Dolly-in | Hero moment | Triumphant reveal |
| Pull-out | Isolation, scale | Endings, reveals of scope |
| Orbit / arc | Emphasises importance | Dynamic reveal |
| Whip pan / crash zoom | Shock, comedy, transition | Pattern break |

**Rules that are not optional:**

- **Exactly one move per clip.** Two named moves, or a vague one, makes the model default
  to a near-static drift.
- **Some single moves have two-word names, and they are still one move.**
  `dolly zoom`, `slow push-in`, `slow pull-out`, `whip pan`, `crash zoom`. The rule above
  is about two *independent* moves ("pan while craning down"), never about word count.
- **Focus and duration are not moves.** `rack focus` (reveal), `long take` (suspense) and
  `slow motion` describe something other than where the camera goes. A clip carrying only
  one of them still has no camera move and still needs one, or `static shot`.
- **A locked frame must be stated:** `static shot`, optionally
  `static shot, no camera movement`. Omitting the camera line does *not* give you a still
  frame.
- **Never write "locked" at all** — not "locked tripod", not "locked shot", not "locked
  off". "Locked tripod" is misread: the model adds motion anyway and has been observed
  rendering a literal tripod in frame. `static shot` says the same thing and is safe.
- **Never combine contradictory terms** ("drone shot zooming into close-up"). Pick one
  perspective.

---

## 2b. Angle — the cheapest way to change how someone reads

Where the camera sits relative to the eyeline changes who has power in the frame, before
a word is spoken. An unstated angle renders eye-level: fine as a decision, a defect as an
accident.

| Angle | Reads as | Use |
|---|---|---|
| Eye level | Neutral, equal | Default; dialogue |
| Low angle | Authority, threat | Someone taking power |
| High angle | Vulnerable, small | Someone losing it |
| Dutch angle | Unease, instability | Something is wrong and nobody has said so yet |
| Overhead / top-down | Powerless, fated | A person as one dot among many |

The trap is using the angle that contradicts the beat — a low angle on someone being
humiliated makes them look like they are winning. Pick the angle for who should have
power in that second, not for what looks impressive.

## 3. The over-the-shoulder rule (stops the wrong face appearing)

In a dialogue scene between two characters, the listener is where generated clips go
wrong — the model renders their face, badly, in the corner of frame.

- **Close-ups (85mm, f/1.8):** only the speaking character is visible. Do not mention
  the other one at all.
- **Medium (50mm/35mm) and wider:** show the **back of the other character's head** at
  the edge of frame. Describe them *only* from behind — "back of head", "back of
  shoulders", "shirt collar seen from behind". **Never** describe their face,
  expression, eyes or glasses.

Wrong — this renders a face:

> Man frozen in shock visible at frame edge left in soft focus. Balding with
> semi-rimless glasses. Completely stunned expression, gaze fixed on her.

Right — this renders an over-the-shoulder:

> The back of his head and shoulders visible at frame edge left in soft focus. His
> balding crown and the back of his light blue-grey shirt collar seen from behind. He is
> facing away from camera toward her.

Same idea generalises: **in a close-up, keep exactly one character in focus.** Others
are soft-blurred, off-screen, or speaking from off-screen.

---

## 4. Lighting

| Style | Description | Use |
|---|---|---|
| Flat documentary | Even, overhead fluorescent | Hallways, offices |
| Split warm/cool | Two colour temperatures meeting | Doorway scenes, transitions |
| Soft diffused | Gentle, flattering | Interview subjects |
| Harsh overhead | Hard shadows | Raw documentary feel |
| Window natural | Warm, cosy | Home interiors |
| Chiaroscuro | Strong side light, extreme contrast | Dramatic character moments |

**Colour temperature:**

| Setting | Temp | Mood |
|---|---|---|
| Fluorescent office | 4000–4100K | Cool, corporate |
| Afternoon daylight | 5200–5600K | Natural, bright |
| Golden hour | 4000–4500K | Warm, heroic |
| Indoor apartment | 5000K | Cosy, domestic |
| Mixed (doorway) | split 3500K / 4100K | Transition |

**Standard lighting line:**

```
Lighting: [quality] [source] at [angle]. [detail]. Color temperature [X]K. [grade]. 24fps.
```

Examples:

```
Lighting: Bright harsh fluorescent office lighting from multiple ceiling panels.
Color temperature 4100K cool white. Very slight green cast typical of office
fluorescents. 24fps.

Lighting: Warm natural afternoon sunlight from a window at camera-left. Soft golden key
on the face. Color temperature 5000K. Warm cosy domestic grade. 24fps.
```

**One lighting motif per clip.** And never write a mid-clip lighting transition into the
look line — a light that changes is a **cut**, not a description.

---

## 5. Framing for the aspect ratio

**9:16 (vertical):** centre the subject · leave headroom at the top for platform UI ·
keep everything important inside the centre 80% of frame · text overlays read best at
top and bottom.

**16:9 (landscape):** rule of thirds · more room for environmental context · better for
wide shots and scenery.

| Platform | Ratio |
|---|---|
| TikTok · Instagram Reels · YouTube Shorts | 9:16 |
| YouTube (standard) | 16:9 |
| Facebook / LinkedIn | 16:9 or 1:1 |

If the interface already sets the ratio, do **not** also write it in the prompt text.

---

## 5b. Composition — where the subject sits inside the frame

| Composition | Reads as | Use |
|---|---|---|
| Centre frame | Control, stability | Someone in command of the room |
| Negative space | Anticipation | Something is about to enter the empty half |
| Frame within a frame | Trapped | Shooting through a doorway, window, mirror, windscreen |

**In 9:16 you buy centre framing whether you want it or not.** Section 5 tells you to
centre the subject for platform reasons — under-the-UI safety, not meaning — and the
frame still reads as control. To signal instability in vertical you have to work for it:
a dutch angle, or the subject pushed hard to one side against deliberate negative space.

**Frame-within-a-frame is easy to use by accident.** Anything shot through a window,
a doorway or a windscreen is one, and it says *trapped* even when the scene is about
freedom. Check that the meaning is the one you want before shooting the whole film
through glass.

## 6. Where the shot fits the story

One clip is **one beat**. Choose the coverage to match:

| Beat | Coverage |
|---|---|
| Reveal | slow push-in as they see it |
| Reaction | close-up; face shifts from calm to shock |
| Tension build | rising stillness, held breath, static frame |
| Payoff | release — exhale, sudden motion, pull-out |
| Establishing | wide, deep DoF, then cut in |

Standard sequence: **establishing wide → medium → detail/close.** Reveal, then specify.

Identity holds best in calm close and medium shots and drifts in fast action and wide
shots — so put identity-critical beats in calm coverage and let action shots run loose.
See [`identity-and-likeness`](../identity-and-likeness/SKILL.md).

---

## 7. Emotion -> technique (reverse index)

Every table above runs technique -> effect. This one runs the way a writer actually
arrives: knowing how the beat should feel and needing the camera that delivers it.

| The beat should feel | Reach for |
|---|---|
| Anticipation | Negative space |
| Authority | Low angle |
| Chaos | Shaky camera |
| Conflict | Over-the-shoulder |
| Connection | Close-up, 85mm f/1.8 |
| Control | `static shot`; centre frame |
| Immersion | POV |
| Intensity | Extreme close-up |
| Isolation | Wide, subject small in frame |
| Loneliness | Slow pull-out |
| Momentum | Tracking shot |
| Powerless | Overhead / top-down |
| Reveal | Rack focus |
| Shock | Dolly zoom |
| Suspense | Long take |
| Tension | Slow push-in |
| Trapped | Frame within a frame |
| Unease | Dutch angle |
| Urgency | Handheld |
| Vulnerable | High angle |

**[FIELD PRACTICE].** These pairings are live-action editorial convention, not vendor
documentation, and nothing in this repo yet proves a video model renders "low angle" as
*authority* rather than merely as a low camera. Use them to choose, then let a shoot
confirm. Source and the gap analysis they came from:
`docs/camera-agent/emotion-technique-map.md`.

