# Director Workflow

> Linear video production workflow — start workspace, write screenplay, plan shots, generate, review, and merge into final video.

- Skill: `jd-opensource/director-workflow-3` (Agent Skill)
- Install (CLI): `npx skillmds@latest add jd-opensource/director-workflow-3`
- Raw SKILL.md: https://api.skillmd.com/api/skills/jd-opensource/director-workflow-3/raw
- Safety review: pending
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Productivity
- Author: jd-opensource (https://skillmd.com/u/jd-opensource)
- Updated: 2026-09-17
- Page: https://skillmd.com/skills/jd-opensource/director-workflow-3

---


# Director Workflow

Execute the video production workflow as a linear pipeline. Each phase completes before the next begins. Only pause for user confirmation at designated interaction points.

## Global Rules

- Use director tools as the sole interface to workspace state. Never manually edit `state.json`, `story.md`, `story_profile.json`, or shot JSON files with filesystem tools.
- Never paste full screenplay, shot specs, or shot lists into chat. Write them with tools, report only brief status.
- Advance through phases sequentially. Do not skip ahead or route between sub-skills.
- When writing story or shot content, call `get_guidance` to load the relevant reference for quality guidance.
- Keep user-facing text minimal: short status, concise questions, and final results only.
- Reply in the same language the user is using. Do not mix languages in a single response.
- Language policy is story-wide and mandatory. Unless the user explicitly requests another language, write all natural-language `story_profile` content in the user's conversational language: summary, every `beats[].summary`, character anchors, scene anchors, every `beats[].dialogue_intent`, and `shot_to_content` prose. Write `story_md`, every complete shot caption, all dialogue, and every user-facing Agent reply in that same selected language. When the user's messages mix languages without an explicit choice, follow the primary language of the latest substantive user instruction. Keep only required technical tokens such as `ID_A`, `shotN:`, and `OCR` unchanged. Unless the user explicitly requests a silent, wordless, no-dialogue, or person-free outer shot, every outer shot must contain spoken dialogue. Do not depend on a `speaks` field or any other structured speech flag to enforce this.
- Voice-anchor policy is binary and language-matched: every speaker has exactly one stable voice anchor, using `ID_X's voice is ...` for English captions or `ID_X的声音是...` for Chinese captions, plus a heightened scene-specific delivery before English `ID_X says` or Chinese `ID_X说`. A non-speaker has no voice anchor or voice-quality description. Keep the voice anchor roughly one third of the former verbose length: 8 to 14 English words or 12 to 25 Chinese characters.
- Never mix Chinese and English natural-language prose in one caption. Necessary technical tokens such as `ID_A`, `shot1:`, and `OCR` may remain unchanged; Chinese speech uses `ID_X说`, never `ID_X says`.
- Never write, imply, or preserve a sigh, audible breathing or breath sound, inhale, exhale, intake of breath, panting, gasping, sniff, snort, nasal hum, nasal grunt, breathy/airy voice quality, microphone-caught respiration, or nasal voice quality, even when the user explicitly requests one. Replace it with facial expression, gaze, posture, action, or permitted spoken dialogue.
- Once the workflow enters Phase 4 (Create Shot Prompts) or later, `shot_count` is locked and cannot be modified. If the user requests to add or reduce shots during Phase 4, 5, or 6, reply: "当前阶段已不可以修改镜头数。" Do not change `shot_count` or `story_profile` beats after Phase 3 is complete.
- When you need the user to pick among a few concrete choices, call `ask_user` instead of listing options as plain text or raw JSON in chat.
  - Put the short intro in `ask_user.content` and each choice card in `ask_user.questions` with `question`, `options`, and optional `allow_custom: true`.
  - Each card starts with `status: "pending"`; the WebUI persists the user's tap as `status: "answered"` so refresh keeps the selected state.
  - When using `ask_user`, do NOT also write the same question or similar content as plain text in chat. All question content must go through `ask_user` only — no duplication.
  - Each response contains at most ONE `ask_user` card.
- When presenting multiple options to the user, always use `ask_user` card. Never output options as a plain text bullet list in chat.
- **Critical: When calling `ask_user`, the `ask_user` card IS the entire message. Do not generate any text output in the same response. No intro text, no duplicate question, no bullet list — only the `ask_user` tool call.**

---

## Phase 1: Start Workspace

**Tool:** `start_director(goal=<user's video idea>, continue_policy="ask")`

This phase is silent — do not tell the user you are opening a workspace or report tool execution status.

1. Call `start_director` with the user's stated goal.
2. If the tool returns `needs_confirmation` (existing unfinished project found):
   - **ASK USER:** Continue the existing project or start a new one?
   - Wait for reply before proceeding.
3. Call `get_workplace_status` to load current state.
4. If a story already exists, call `get_story` to load context.

**Exit condition:** Workspace is active and current state is known.

---

## Phase 1.5: First-Frame Reference Image Gate

Read `get_workplace_status` / fact fields. Do not guess frontend upload state.

- `reference_image_present`: whether a persisted reference image exists
- `reference_image_locked`: true after `shot_count` is confirmed; then the image cannot change
- `auto_generate`: when true, skip this entire phase and skip later confirmation cards
- `reference_image_needs_story_rewrite`: true when the persisted image changed since the last `write_story`

**Skip this phase when `auto_generate=true`.** If a reference image is already present, use it immediately for screenplay writing. Then go to Phase 2.

After Phase 1, if `auto_generate=false` and `reference_image_present=false`, call `ask_user` with question 「在开始构思之前，是否要上传首帧参考图？」 and options exactly: `需要上传,已上传完毕` / `不上传`.

**Backend owns the match loop.** Do not issue mismatch cards yourself. If the user text contains `REFERENCE_IMAGE_GATE match=true`, continue conceiving. A second consecutive `不上传` while an image is present is deleted by the backend.

When the user picks `我想修改/增删参考图`:
1. Call `ask_user` with question 「是否已完成参考图修改？」 and the only option `[是]`.
2. After they answer `是`, call `get_workplace_status`. If `reference_image_needs_story_rewrite=true` or the current image differs from the last story, rewrite the **entire** screenplay from the CURRENT image (or text-only if none). Then `write_story(..., confirmed=false)` and show the story-direction card.

When the user picks `可以，按这个来`: ignore `reference_image_needs_story_rewrite`. Do not rewrite. Lock the current `story_md` with `write_story(..., confirmed=true)` and ask shot count. Rewrite only after they answer `是` on 「是否已完成参考图修改？」.

---

## Phase 2: Write Story

**Tools:** `write_story`, `get_story`

**Reference:** Before writing, call `get_guidance(topic="shot-sequence-patterns")` for structural guidance.

### 2a. Gather Requirements (max 3 rounds)

Ask only the smallest set of questions needed to write the screenplay. Prefer asking over advising, and ask for facts or choices that unblock writing.

1. If the user has not provided any story idea at all — only expressed intent to create (e.g. "我想创作一个故事", "帮我拍个视频") — ask what story they want to tell and offer a few concrete story premises as options via `ask_user`. Always include the extra option `我想修改/增删参考图` on this story-direction card.
2. If the user has given a recognizable story idea (e.g. "孙悟空大闹天宫", "一个人深夜在便利店遇到老朋友"), skip step 1. If critical information is still missing to write the screenplay, ask **one question** about the most critical gap via `ask_user`.
3. Once you have enough detail to write the screenplay, stop asking and go straight to drafting.
4. If `auto_generate=true` and `goal.shot_count` is still empty, skip remaining requirement questions and draft immediately. Call `set_director_goal(shot_count=<auto_generate_shot_count>, shot_duration_sec=10)` when locking the story. If `goal.shot_count` is already locked, keep that value and do not overwrite it with a duration-tier default. Do not ask the user to confirm the screenplay or shot count.

Rules:
- Max 3 rounds total. If the user already provided enough detail, skip remaining rounds and go straight to drafting.
- Each `ask_user` call has at most ONE question card.
- When the user responds to an `ask_user` card (by tapping an option or typing a reply), treat that response as an answered question. Do not re-ask the same question or ask for the same information again. Advance to the next step immediately.


### 2b. Draft Screenplay

1. Create the displayed screenplay in the user's conversational language unless the user explicitly requests a different screenplay display language; when the user mixes Chinese and English without such a request, follow the primary language of the latest substantive user instruction. `story_md` should contain only the story itself: synopsis. For a Chinese-speaking user this means Chinese `story_md`; for an English-speaking user this means English `story_md`. Do NOT include shot breakdowns or scene-by-scene shot rhythm.
2. Call `write_story(story_md=..., story_profile=..., summary=..., confirmed=false)`.
3. `story_profile` must be compact and use the same selected language for all natural-language prose: `summary`, non-empty provisional `beats`, character anchors, scene anchors, `shot_to_content`, and every `dialogue_intent`. `content_to_shots` may retain technical shot IDs. It MUST store `caption_language`, exactly `Simplified Chinese` or `English`, chosen from an explicit request or otherwise the user's conversational language. It MUST also store the aligned `dialogue_language`, `Mandarin Chinese` for a Chinese caption or `English` for an English caption. Lock both across the story; never create mixed-language natural-language output. Unless the user explicitly requests silence, every provisional beat has a non-empty `dialogue_intent`.
4. Plan visibly distinct outer-shot scenes. Keep one scene for at most 1 to 4 consecutive outer shots, then move to a clearly different location, layout, time, lighting, weather, or story situation unless the user explicitly requires the same scene to continue. Apply any user-requested new scene immediately; outer-shot scene changes are separate from internal `shotN:` segments.
5. If `auto_generate=true`, call `write_story(..., confirmed=true)` and do not ask the user to confirm the screenplay. Otherwise via `ask_user`, ask the user to confirm or revise the screenplay. This is the story-direction card even when Phase 2a was skipped because the user already gave a story idea. Options MUST include, in this order: `可以，按这个来`, `需要修改`, `我想修改/增删参考图`. `allow_custom=true` is OK. Do not omit the reference-image option. The card `content` MUST match the `story_md` just written — not an earlier draft. Do not add any chat text — the `ask_user` card is the only output. Do not mention "确认剧本", workspace buttons, or "下一步".

### 2c. Revision Loop

If the user requests changes:
1. `get_story` to read current version.
2. Apply changes internally.
3. `write_story(..., confirmed=false)` with updated content.
4. Reply with a short change note.

Repeat until user is satisfied.

### 2d. Lock Screenplay → Phase 3: Set Director Goal

This is a two-step sequence that spans two user turns. Follow it exactly.

**Turn 1 — User confirms screenplay:**

If `auto_generate=true`, skip this card. Call `write_story(..., confirmed=true)`, then `set_director_goal(shot_count=<locked goal.shot_count, else auto_generate_shot_count from get_workplace_status>, shot_duration_sec=10)`, update beats, and stop. Do not ask about shot count. Do not mention workplace buttons.

When `auto_generate=false` and the user confirms the screenplay, this response must contain ONLY:
1. Call `write_story(..., confirmed=true)` (silent, no output). Ignore `reference_image_needs_story_rewrite` on this confirm turn. Do not rewrite the story.
2. Call `ask_user` to suggest shot count. **Hard rule for this first recommendation:** every option MUST use ~10s per shot (`shot_duration_sec=10`). Never recommend 4s, 6s, 8s, or any duration other than ~10s. Example: question="这个故事用几个镜头来讲比较合适？", options=["4个镜头，每个约10秒", "6个镜头，每个约10秒"], allow_custom=true.

Zero text output in this turn. No "已确认". No "确认剧本". No "下一步". Only the `ask_user` card.

**Turn 2 — User confirms shot count:**

When `auto_generate=true`, skip any confirmation. If `goal.shot_count` is already locked, keep it. After `set_director_goal` and beat alignment, stop. Do not output workplace-button guidance. The auto-generate pipeline continues on its own.

When `auto_generate=false` and the user confirms shot count:
1. Validate `shot_duration_sec` does not exceed 10 seconds. If the user requests more than 10s per shot, inform them the maximum is 10 seconds and ask them to choose again via `ask_user`.
2. Call `set_director_goal(shot_count=..., shot_duration_sec=...)`.
3. Call `write_story` to update `story_profile` beats to match the confirmed `shot_count`. Preserve both locked `caption_language` and `dialogue_language`. Unless the user explicitly requested silence for a beat or the whole story, every beat must produce speech and carry a non-empty `dialogue_intent`; do not depend on a `speaks` field surviving in stored state. Preserve the 1-to-4-outer-shot scene grouping and keep later scenes visibly distinct.

After these tools succeed, do not write extra chat and do not mention 「进入逐镜打磨」 or 「确认并一键成片」 — those are Workplace 02 buttons only. The system will tell the user to click 「下一步」 to preview the storyboard.

**Exit condition:** `set_director_goal` shot_count is not null and `story_profile` beats are aligned.

---

## Phase 4: Create Shot Prompts

**Tool:** `create_shot_prompt` (called once per shot)

**Reference:** Before writing shots, call `get_guidance(topic="shot-prompt-writer")` for caption format rules.

1. Read the confirmed story with `get_story`.
2. For each shot (1 to `shot_count`):
   - Pass the current beat, the relevant user conversation, and any explicit silence or dialogue-language instruction into the selected PE guidance. Do not require a `speaks` field: unless the user explicitly requested silence, the finished outer shot must contain speech. Preserve the locked story language; if stored metadata is absent, recover the same language from the explicit request or user-conversation fallback instead of guessing from names, ethnicity, appearance, or location.
   - Create the shot spec in the locked full-caption language with realistic/live-action style by default, stable character anchors with voice/no-voice declarations, temporal action sequence, explicit BGM/SFX declarations, the language-matched no-OCR declaration required by the selected PE, and `num_frames`. Give each ID its detailed description only on its first outer-shot appearance and compact continuity anchors later. For internal cuts, keep a character description in every segment where that ID is visible: use the full or outer-shot-appropriate introduction at the ID's first internal appearance, then roughly one-third-length continuity descriptions without repeating formal anchor starts. In a speaking action/dialogue passage, place English `ID_X says` or Chinese `ID_X说` between pre-speech and post-speech action rather than at the end, without assigning it to a fixed `shotN:` segment. Across all internal segments combined, target 1500 to 1800 Unicode characters and use fewer than 2000 characters total (maximum 1999). Compose concisely and rewrite before submission when needed; never truncate.
   - Call `create_shot_prompt` with the complete shot spec.
3. Report which shots were created (IDs only, no specs in chat).
4. **ASK USER:** Briefly mention one or two shots you think turned out well, then ask if they want to adjust anything before generating.

If the user requests changes to specific shots:
1. `get_shot(shot_id=...)` to read current spec.
2. Apply changes.
3. `create_shot_prompt` with updated spec.
4. Report the update briefly.

**Exit condition:** All shots created and user confirms readiness for generation.

---

## Phase 5: Generate & Review

**Tools:** `generate_echo_shot`, `review_shot`, `get_shot`, `create_shot_prompt`

### 5a. Queue All Shots

1. Call `get_workplace_status(include_shots=true)` to get current shot states.
2. For each ready shot that is not already queued/generated/approved:
   - Determine `reference_shot_ids`:
     - Shot 1 may use an empty list.
     - All other shots MUST reference at least one earlier shot for visual continuity.
     - Include prior shots that share characters, environment, props, or wardrobe.
     - For continuous shots (`cut=false`), always include the immediately previous shot.
   - Call `generate_echo_shot(shot_id=..., reference_shot_ids=[...])`.
   - `reference_shot_ids` is narrative/context metadata only. Never treat it as
     permission to add a Memory slot.
3. Submit all shots without waiting for earlier ones to finish.
4. Briefly report how many shots were submitted.

### 5b. Review Loop

When shots are generated (via callback or status check):

- **Accepted / Approved:** Keep the shot, move on.
- **Revise / review_fail:** 
  1. `get_shot(shot_id=...)` to read current spec.
  2. Update the shot prompt based on feedback via `create_shot_prompt`.
  3. Recalculate `reference_shot_ids` for the updated context.
  4. `generate_echo_shot` again for that shot.
  5. Briefly report that the shot has been updated and requeued.

If human review is needed:
- **ASK USER:** Present the shot for review, ask which shots need changes and what to change.

Repeat until all required shots are accepted.

### 5c. Build Memory for the next shot

After an accepted shot has produced profiled Memory Workspace assets, read only
the `memory_assets` text catalog from `get_workplace_status`. Ignore assets with
an empty profile. Use `reference_type`, `reference_label`, `identity_ids`, and
`profile_text` to match each asset's assigned role, then call
`set_shot_memory_recommendations` with zero to seven
ordered draft slots and reasons. The human owns final approval in Build Memory;
only `approved_memory_slots` may reach R2V, and `reference_shot_ids` must never be
silently converted into slots.

**Exit condition:** Every shot has status `accepted` or `approved`.

---

## Phase 6: Merge

**Tool:** `merge_shot`

1. Do **not** ask whether to merge via `ask_user` or chat. Stepwise merge confirmation is the Workplace 03 「下一步」 button; one-click already merges automatically.
2. Do not call `merge_shot` until the workplace injects `workplace_workflow_start_merge`.
3. If all shots are approved and the user has not clicked 「下一步」 yet, you may briefly tell them to use that button. No option card.
4. When merge is injected: call `get_workplace_status(include_shots=true, include_jobs=true)` to verify no `review_fail` shots and no pending generations, then `merge_shot(shot_ids=[...in timeline order...])`.
5. When the merge callback returns, provide the final video link to the user.

**Exit condition:** Final video delivered to user.

## Anti-Patterns

- 不要跳过剧本确认直接进入镜头生成
- 不要在有未通过的 shots 时执行合成
- 不要手动编辑 director state 文件，所有状态变更通过 director tools
- 不要在故事发生变化后跳过 `write_story` 调用
- 不要在执行期向用户描述内部工具调用逻辑、队列机制或 reference 选择细节
- 不要收到 card feedback 后只回复确认而不实际更新 shot prompt 并重新生成
- 不要忽略 workspace 状态凭记忆继续工作
- 永远不要让用户去工作区"确认剧本"——工作区没有这个按钮。剧本确认通过 chat 完成，agent 调用 `write_story(confirmed=true)` 即可
- 不要在用户确认镜头数之前提及"下一步"
- 镜头数确认后不要再下发「进入逐镜打磨 / 确认并一键成片」ask_user 卡片，这两项只出现在右侧工作区 02
- 全部镜头接受后不要再下发「合成最终视频 / 暂不合成」ask_user 卡片，逐步合成入口只在右侧 03「下一步」
- 不要用纯文本 bullet list 向用户展示多个选项——所有多选项必须通过 `ask_user` 卡片呈现

