Talking Head Video Editing
What this skill covers
Required input: an existing talking-head / 口播 video uploaded to the project. If the user wants to start without one (e.g., generate a fresh talking-head from scratch), this skill doesn't apply.
When the user enters this workflow without a source video uploaded yet, ask only for missing treatment or preference decisions in ordinary WorkBuddy chat. Load asset-import for the source media and follow its WorkBuddy upload strategy, including its editor upload fallback when local automation is unavailable.
If this workflow creates, targets, or opens a ChatCut project, follow the expert's editor handoff rules before nontrivial edits and again before final delivery when the visible editor may no longer match the project.
Independent treatments that can be applied to talking-head videos. Pick the ones that match what the user wants — not all are needed every time.
- A-roll editing (中文称 语音剪辑 / 含 去口癖、停顿、重复) — transcript-based speech editing. Common operations include cleanup, highlight extraction, restructure, opening hook, and others as needed for the aligned outcome.
- Motion graphics overlay (英文展示给用户时写全称 Motion Graphics,不要缩成 "MG";中文产品术语固定为 MG 动画——不要叫"动效""字幕条""动态字幕"等其它说法) — reinforce key information, structured content, and topic transitions with on-screen motion graphics
- B-roll (industry term — keep as "B-roll" in any language, do not translate) — cover jump cuts or visualize what's being said
- Background music (中文 背景音乐) — set mood and smooth micro-gaps
- Captions (中文 字幕) — on-screen text for accessibility
- AI Voice Isolation (中文 AI 人声隔离) — clean or isolate spoken human voice with DeepFilterNet3, picture untouched. Use the visible
isolate_voice tool when available.
用户语言为中文时,在 widget options / choices options / 对话文案里严格使用上面括号里的产品术语——别自己再翻译一遍,会跟产品其它地方对不上。
What shapes the edit
Beyond picking treatments, a talking-head edit is shaped by several orthogonal variables. When the user's ask is vague, these are what's worth clarifying first:
- Target — platform (YouTube / TikTok / Shorts / ...), desired length, and final canvas aspect ratio
- Which treatments to apply — the treatments above are optional; don't assume all of them apply
- Pacing / tone — tight / energetic / formal / casual; brand or voice preferences if stated. (For MG visual style, follow the active Motion Graphics skill/workflow.)
When more than one of these variables is missing, ask them together in one concise WorkBuddy message. Do not run a fixed questionnaire or ask for decisions already visible in the project or source material.
Resolve and apply the final canvas before placing the first visual item. An explicit user ratio wins; otherwise use the platform convention when it is unambiguous, or the primary source ratio when neither platform nor ratio is specified. Update the target timeline with manage_timelines and read it back before visual assembly.
Order of execution
When multiple treatments have been aligned with the user, they depend on each other and must be finalized in dependency order. This section is only relevant after alignment — it doesn't tell you what to start with on a fresh request.
The speech timing (set by A-roll editing) anchors everything downstream — MG placement, B-roll cut-covers, music duration, and caption sync all reference the final speech timeline.
So: finalize A-roll editing before committing any visual, audio, or text layer. Don't write captions against pre-edit speech, don't cut music to pre-edit length, don't place MG against timing that will shift.
You must confirm the result with the user after each major step before starting the next, unless the user has explicitly asked to run end-to-end without stopping. Key checkpoints when multiple treatments apply: after A-roll editing finalizes the speech timing; before MG creation (confirm style and direction, and, when it isn't obvious, whether it sits over the video as an overlay or takes the whole frame); after MG placement; same pattern for B-roll, music, and captions. Don't bundle multiple checkpoints into one response — confirm each step separately. An upstream mistake forces redoing everything downstream (for example, an MG placed against pre-cleanup timing must be repositioned after the timeline shifts).
A-roll editing
Scenario
In a talking-head workflow, the first step is usually A-roll editing: editing the original spoken footage.
A-roll edits are ultimately applied to the timeline and change what the viewer actually hears and sees. However, the editing decisions should usually start from the transcript, because the core question is: what spoken content should the viewer hear, and what should be removed, compressed, or reordered?
Common A-roll tasks
A-roll editing is not only cleanup. First decide what spoken-content task the user is asking for, then choose the editing strategy and tools.
Common tasks:
- Cleanup — remove mistakes, repeated attempts, verbal habits, filler words, and meaningless pauses so the speech becomes clearer and more natural.
- Highlight extraction — pull the most valuable, opinionated, emotional, or topic-relevant moments from longer footage.
- Restructure — reorder spoken content, such as moving the conclusion earlier, grouping by topic, or combining scattered parts into a clearer structure.
- Hook / short version — use a strong claim, result, conflict, or question from the source as the opening, or compress long content into a shorter version.
- Target-script / script alignment — match, keep, and reorder spoken content according to a user-provided target script, target paragraph, or desired content.
Cleanup is the most common task and the one most likely to fail from bad boundary decisions. It is described in detail below. Other tasks get shorter rules, but still follow the shared A-roll principles: complete meaning, clear boundaries, and natural listening flow.
Shared A-roll principles
These principles apply to all A-roll tasks, not only cleanup.
- Decide the task before choosing the tool. Do not let tool availability change the editing strategy.
- Edit by complete semantic units. Whenever possible, move/delete/keep complete sentences, complete ideas, complete answers, or complete steps. Do not cut out a half-sentence just because a few words match.
- When the task names what to keep, trim to that boundary. The inverse of the rule above, for any task that specifies which content to keep — restoring a specific sentence, matching a target script, pulling a named highlight, building a version: keep exactly the requested span. Trim the kept range to start and end at the requested words and drop the off-script head/tail of the source
[sN] segment it sits in; keeping a whole segment for one requested sentence is over-keeping that drags in unrequested speech. This applies only when the task names what to keep — never to open-ended cleanup, where you keep complete units (above).
- Do not stitch unfinished fragments across retakes. Do not combine incomplete pieces from different attempts into one artificial sentence. This does not make the earlier attempt disposable: keep a complete useful lead-in, setup, contrast, category, evaluation, or context if it is not repeated later and can naturally connect to the later complete retake.
- Preserve connective tissue. List labels, contrast words, subjects, verbs, and adjacent source words are not filler when removing them makes a kept idea ungrammatical, abrupt, or misleading. Trim the smallest span that keeps the line speakable.
- Keep listening flow natural. The result should still have natural phrasing and breathing room. Do not make sentences feel glued together just to make them "clean."
- Be conservative when boundaries are uncertain. If unsure whether a cut harms meaning, logic, or listening flow, keep it or make a smaller cut.
- Confirm complex changes first. For complex restructuring, aggressive shortening, structural changes, or generated hooks, confirm target length, structure direction, and what to preserve with the user before editing.
- Explain content, never indices. You MUST NOT explain edits to the user with internal addresses such as
[sN], [cN], [gap], word indices, clip ids, or segment ids. The user cannot see those addresses and will not understand what they mean. Use the actual spoken content, a short quote, or a plain-language description of the edit.
- Never name a screen position for a panel. When you invite the user to review or fine-tune the result, call it "the Transcript panel" (中文「文字稿面板」) — never a direction (left / right / side / 左侧 / 右侧). The layout is rearrangeable and the panel does not sit in a fixed corner.
Cleanup goals and decisions
What good cleanup means
Good cleanup does not mean making the video as short as possible, and it does not mean rewriting the speaker into a different script.
Good cleanup means:
- The logic stays coherent
- The expression becomes clearer
- The audio feels natural
- Obvious mistakes, repeated attempts, meaningless stalls, and filler are removed
- The speaker's intent, tone, and natural rhythm are preserved
Bad cleanup usually falls into two failure modes:
- Under-cleaning: obvious mistakes, repetition, long pauses, or filler remain.
- Over-cleaning: sentences are cut off, meaning is missing, rhythm becomes too hard, or the result sounds stitched together.
Default principle: remove defects without changing meaning; make speech smoother, not harder; prefer small local cuts over whole-sentence or whole-segment deletion; when unsure whether a cut harms meaning, keep it.
How to judge common cleanup cases
Below are the common cleanup categories and how to make editing decisions for each.
Meaningless filler words
Fillers fall into two categories.
The first category is clearly meaningless hesitation sounds. These are usually safe to remove:
When they do not carry special meaning, use clean_script first for bulk cleanup.
The second category depends on context and must not be removed by word list alone:
so
like
然后
就是
嗯
啊
那个
那
对
所以
但是
How to decide:
- If the word is only hesitation or padding, remove it.
- If it carries sequence, continuation, contrast, cause, reference, response, emphasis, or natural tone, keep it.
- If removing it makes the surrounding words sound hard-spliced, keep it or only compress the pause.
- If unsure, keep it.
Examples:
um, I think this solves the main problem -> remove um.
It works like a checklist -> keep like; it is a comparison.
The upload failed, so we retried it -> keep so; it carries cause/result.
right after the call, send the recap -> keep right; it modifies timing.
然后我们再看第二点 -> keep 然后; it marks sequence.
Retakes and repeated attempts
A retake is when the speaker retries the same intended idea because they misspoke, got stuck, forgot words, or restarted. Retake cleanup is not "delete repeated text." The goal is to keep one complete, natural, logically coherent version of the intended idea.
Use this decision path:
- Decide whether it is really a retake.
Treat it as a retake only when multiple attempts are trying to say the same intended idea. Do not treat it as a normal retake when the repetition is intentional emphasis, a rhetorical beat, a structural marker, or a second pass that adds new information or tone.
- Define the complete version to keep.
A complete version may include more than the main content sentence. It may need a lead-in, connector, section marker, topic setup, contrast, qualifier, subject, object, or conclusion. These are not filler when the kept content depends on them.
- Cut only the failed or covered part.
Remove only words that are wrong, dangling, abandoned, or fully covered by the kept version. The cut boundary starts at the repeated or failed idea, not automatically at the earlier transition, setup, or continuous speech. If earlier speech contains useful context that the kept version does not repeat, keep it.
- Choose the best complete attempt.
If several attempts are complete, usually prefer the later one because it is often closer to the speaker's intended take. But do not choose the last attempt mechanically. If the later attempt is missing needed context, structure, subject, object, or conclusion, keep the more complete version or preserve the missing lead-in from the earlier attempt.
A repeated lead-in is redundant only when another equivalent lead-in remains naturally connected to the kept content. If removing every copy makes the result lose structure or sound abrupt, keep one natural copy and remove only the extra restarts. Do not stitch unfinished fragments from different attempts into one artificial sentence.
Examples are patterns, not a closed list:
- Local false start inside a kept sentence:
There, there's no After Effects, no Premiere, no DaVinci Resolve learning.
Keep the complete sentence, but remove the abandoned restart:
There's no After Effects, no Premiere, no DaVinci Resolve learning.
Do not keep the stray first word just because the full sentence is otherwise useful.
- Repeated structural lead-in:
And secondly, ... and secondly, we're introducing a brand new UI.
Remove the extra restart, but keep one natural lead-in attached to the kept content:
And secondly, we're introducing a brand new UI.
Do not delete every structural marker and leave only:
We're introducing a brand new UI.
- Useful setup before a failed ending:
Then the next one is different from comedy. It is popular on Disney Plus. It is called...
Later retake:
It is a popular Disney Plus show called Love Story.
Keep useful setup that the later retake does not repeat, and cut from the failure point:
Then the next one is different from comedy. It is a popular Disney Plus show called Love Story.
False starts and unfinished fragments
Use false starts / unfinished fragments for this category. False start is the more natural editing/transcription term for a speaker beginning a phrase and then restarting or abandoning it; unfinished fragment makes the dangling half-sentence case explicit.
Only remove a fragment when it clearly does not form useful information.
Safe to remove:
- The speaker abandons the thought and a complete version appears later.
- The segment is only a dangling phrase, such as "this is actually..." with no completion.
- It is clearly the leftover beginning of a failed attempt.
Do not remove:
- A sentence that is imperfect but contains useful information.
- A lead-in that provides the subject, object, or context needed later.
- Content that provides setup, contrast, conclusion, emotion, or tone.
If only part of a sentence or segment is wrong, do not delete the useful content around it. Remove only the bad word, phrase, or pause; if a local cut cannot sound natural, keep the segment.
Pauses and breaths
Pause cleanup should default to compression, not zeroing out. Spoken video needs natural breathing room.
Default rules:
- Obvious long pauses over 0.8-1s: usually compress to about 0.3s.
- Between sentences: keep about 0.3-0.5s so listeners can hear natural phrasing.
- Around topic shifts, contrast, or emphasis: keep slightly longer pauses when needed; do not make the delivery too rushed.
- Short breaths inside one sentence: if they are normal breathing, do not remove them.
- Clear long pauses inside one sentence: compress them, but not so tightly that adjacent words sound glued together.
- Long pauses before a retake: if the failed attempts around it are removed, remove the pause with them.
- If the user provides explicit thresholds, follow them. For example, if the user says "only process pauses over 0.8s and keep at least 0.3s", do not process natural pauses under 0.8s.
How to operate on pauses:
- For batch pause cleanup across the timeline or track, use
clean_script. This is the default path for compressing many long pauses.
- Translate common user wording into
clean_script pause rules:
- "Tighter breaths" / "compress pauses" / "compress anything over 0.3s to 0.3s" →
silence: "compress:300" (or the requested cap).
- "Restore some breathing room" / "do not make it too rushed" / "keep at least 0.5s" →
silence: "restore:500" (or the requested minimum).
- "Make all pauses around 0.5s" →
silence: "normalize:500".
- "Keep pauses between 0.3s and 0.8s" →
silence: "range:300-800".
Any rule that makes a pause longer — restore, normalize, or the lower bound in range — never invents new silence. It only recovers pause time that already existed at that exact spot in the original recording. If the original pause was shorter than the requested value, it stops at the original pause length.
- You do not need to call
read_script({ showSilence: true }) before batch pause cleanup. By default, timeline.md hides silence markers, but clean_script can still detect and rewrite silences internally.
- Use
read_script({ showSilence: true }) only when you need to inspect or manually adjust a specific pause. Then edit the visible marker: ~~[silence=0.8s]~~ to fully cut it, [silence=0.8s→0.2s] to compress it, or leave it untouched to keep it.
- After semantic edits, review the final clean
timeline.md. If the final pacing still has many long pauses, run clean_script>; if only one or two pauses feel wrong, use showSilence: true and adjust those manually.
Script gap primitive note:
- Do not create an accidental
[gap] on the primary video track as a pacing pause. A Script [gap] means no source is playing; on the only visible video track it renders as black. If pacing needs breathing room, preserve or restore source silence with clean_script / [silence=...], cover the moment with B-roll/MG/a full-frame visual beat, or intentionally declare the black beat in the plan.
Other A-roll task guidance
Highlight extraction
Highlight extraction is not about making the content as short as possible. It is about selecting the most valuable spoken content according to the user's criteria.
Rules:
- First identify the highlight standard: opinion, conclusion, story, emotion, conflict, tutorial step, data point, or a specific topic.
- Each highlight should be understandable on its own. Do not remove the subject, setup, question, or conclusion needed to understand it.
- Do not keep only a short punchy sentence if the surrounding context is required for it to make sense.
- If the user asks for a specific topic, remove other topics. If the user asks for the "best" or "most exciting" moments, prioritize information density and expression strength.
- After extracting highlights, usually clean up the kept segments so the final result is polished.
Restructure
Restructure means changing the order of spoken content. It does not mean freely breaking sentences apart.
Rules:
- First confirm the target structure: chronological, by topic, by question, conclusion-first, tutorial steps, or short-form pacing.
- Move complete semantic units: complete sentences, ideas, answers, or steps.
- Do not split one sentence so the first half appears in one place and the second half elsewhere.
- After moving content, check whether connectors still work, such as "so," "but," "next," or "this."
- If the user asks for major restructuring without specifying the target structure, confirm before editing.
Hook / short version
Hook / short version work aims to make the opening more compelling or compress long content into a shorter but still complete version.
Rules:
- Prefer pulling the hook from the original footage: a strong claim, result, conflict, question, counterintuitive statement, or emotionally strong moment.
- If a new hook or new narration must be generated, confirm the direction with the user first.
- For short versions, do not cut only by duration. First identify the main line to preserve: problem, core point, key reasons, and conclusion.
- Short versions can remove examples, repetition, and setup, but must keep the logic needed for the point to hold.
- If the user gives a target duration, try to match it. If duration and semantic completeness conflict, explain the tradeoff.
Target-script / script alignment
Target-script / script alignment means cutting the final spoken content according to a user-provided script, target paragraph, or desired content.
Rules:
- The target script is the main constraint: prioritize content that matches the target meaning.
- Natural spoken paraphrases are acceptable, but do not include surrounding content that the target does not ask for.
- If the source has multiple similar versions, choose the most complete, natural, and target-aligned version.
- If target order differs from source order, reorder as needed, but move complete semantic units.
- If the target script omits source context, follow the target. Do not add long surrounding context unless the result would be incomprehensible without it.
Building versions, highlights, and excerpts — stay on Script
Highlight, short version, excerpt, hook, restructure, and making several versions are all transcript-content tasks: drive them through Script (read_script → edit timeline.md → apply_script), never by looking up timestamps and placing source clips manually.
- Pick the starting point by where the content comes from. Versions on the current timeline: trim or reorder
timeline.md and apply_script. A version on its own timeline (the user asked for separate timelines, or wants each version independently editable/exportable): manage_timelines action=duplicate — the copy carries the content and its script, so you immediately read_script → trim → apply_script on it. Building fresh from library assets: manage_timelines action=create, add the source asset, then drive it through Script.
- To bring in source content the current cut no longer shows (a hook line, a segment needed for another version), read
library/<filename>.md, copy the needed [sN] line(s) into timeline.md where they belong, and apply_script. This is how you pull source content onto the timeline — through Script.
- For multiple versions on one track: list every version's
[sN] segments in timeline.md in version order, one version after another, then apply_script once. Reuse is just repetition — the same [sN] segment may appear in more than one version, and repeating the line replays that source range again.
- Never look up timestamps with
find_transcript and place spoken content with edit_item / split_item. If you are converting transcript segments into source frame or second ranges, you are off the editing surface — return to Script. edit_item / find_transcript are only for non-transcript placement such as MG overlays and B-roll visual timing.
Check each version against its request. After assembling a version, highlight, or excerpt, re-read the result end to end and confirm every requested sentence is present, in the requested order, with no extra source carried in. Fix any dropped, duplicated, or out-of-order content before finishing.
A-roll / transcript-based editing workflow
Use this flow for any A-roll task driven by transcript meaning.
- Start with orientation. Call
read_script, then read timeline.md once to understand the user's goal, the content structure, and whether fixed fillers or long pauses are present. If you will run clean_script, do not build the full semantic edit from this pre-clean read.
- For cleanup tasks, run the mechanical cleanup pass before semantic editing when fixed fillers or long pauses are present. Use
clean_script for fixed hesitation sounds (um, uh, er, ah, 呃, 额) and batch pause compression. If both are present, use the default clean_script pass so both are handled together. Do not use this step for context-dependent fillers, retakes, repeated sentences, or anything that needs meaning.
- After
clean_script, always read the refreshed clean timeline.md before semantic editing. Use this refreshed file as the source of truth; clean_script changes the canonical timeline and rematerializes the script, so previously read text may be stale. Do not edit from memory based on the pre-clean script. Then edit timeline.md with semantic judgment: choose the best retake, clean false starts, remove repeated or failed attempts, preserve useful setup and context, reorder content when needed, and keep the speech natural. For long transcripts, work one clear section at a time if that improves judgment accuracy.
- Apply the edit with
apply_script. If apply fails, fix the markdown error or stale state, re-read the current timeline.md if needed, and apply again.
- Review the edited result. After a real
apply_script, read the regenerated clean timeline.md and check what the viewer will actually hear: broken logic, missing context, over-deletion, missed cleanup, wrong order, or pauses that feel too tight or too long. Fix clear problems only. If the final result still needs batch pause adjustment, use clean_script>. Use read_script({ showSilence: true }) only for manual adjustment of specific pauses.
What transcript editing actually changes
Editing timeline.md is not just changing displayed text. It describes which source media ranges should play on the timeline.
[sN] rows are ASR segments, not semantic units. A complete sentence, idea, retake, or transition may span several [sN] rows, and one [sN] row may contain only part of a sentence. Before deciding what to delete or keep, mentally reconstruct the complete spoken sentence or idea across adjacent rows.
- Each spoken-text line maps to a playable source range.
- Inline
~~...~~ removes the corresponding audible audio range.
- Deleting a whole line removes that whole spoken segment.
- Moving/reordering lines changes playback order.
apply_script applies the result back to the timeline.
- Start/end trims may remain as one trimmed clip.
- Deleting words or pauses in the middle of a sentence splits the original clip into multiple new clips: one kept range before the deletion and one kept range after it.
- Moving spoken content also creates a new clip at the destination.
- More clips after middle deletions or moves are expected and usually correct. Do not describe that as a "fragmentation problem" or as proof that word deletion is unsupported.
Tool boundaries
Choose the editing goal and content boundaries first, then choose the tool. Do not let tool availability change the editing strategy.
clean_script: use for mechanical first-pass cleanup: bulk removal of fixed meaningless fillers and batch silence compression/adjustment. It can process silence even when timeline.md is currently rendered without silence markers. Do not use it for context-dependent fillers, retakes, repeated sentences, or semantic decisions.
read_script + apply_script: the main transcript-based editing surface. Use it for real semantic editing: deleting words, sentences, pauses, reordering, or pulling library content onto the timeline.
manage_transcript action fix: only fixes ASR mistakes or speaker attribution. It does not cut audio and does not change what the viewer hears.
- Caption SEGMENTATION (分句 / where pages break) is INDEPENDENT of the transcript and controlled by two per-word primitives only: to SPLIT one card into two, set
display_text forcePageBreak:true on the word that should START the new card; to MERGE a card up into the previous one, set display_text keepWithPrevious:true on that card's FIRST word (works for any break — no box resizing, no wordsPerPage fiddling). To drop a repeated/false-start word, use display_text hidden:true. Box width / fontSize / wordsPerPage are style & density knobs, NOT per-boundary segmentation levers — do not widen the box or raise wordsPerPage to merge or split a specific card. NEVER edit the transcript to fix a caption line break — manage_transcript fix is only for an ASR-misheard WORD (content), not layout. read_captions shows each page's break= reason and per-word keys for these edits.
find_transcript: only locates when a phrase is spoken. It does not edit. If the next step is cutting spoken content, return to Script.
Edit / Write: use these to modify timeline.md. The edit only reaches the timeline after apply_script.
Script details to preserve:
read_script materializes timeline.md (current cut) and library/<filename>.md (full read-only source transcripts) in the workspace.
- Single-word audible deletion is supported with inline strike syntax, such as
[s1] 过去~~呢~~一个月.
- Silence markers are hidden by default. Use
clean_script for batch pause cleanup. Use read_script({ showSilence: true }) only to expose [silence=Ns] markers for precise manual edits such as ~~[silence=0.8s]~~ or [silence=0.8s→0.2s].
find_transcript can locate a phrase for visual timing; it is not the editing surface. Do not use find_transcript + split_item / edit_item to cut, place, or assemble transcript-based clips — this includes highlights, hooks, excerpts, and multi-version cuts. All spoken-content selection, placement, and reuse happens in Script (read_script → edit timeline.md → apply_script).
MG Overlay
Goal
Motion graphics layered into A-roll reinforce what the speaker is conveying — deepening the audience's impression of the key points and helping them grasp content that's hard to land through speech alone. Complete A-roll editing first; MG timing is based on the post-edit timeline.
This section only adds talking-head timing, frame-composition, subject/caption protection, and review constraints. For visual style alignment, MG creation or authoring, implementation constraints, editable properties, asset sizing, and verification, use the active Motion Graphics skill/workflow available in the current ChatCut environment.
MG workflow
For talking-head MG work, treat the video as one edited piece, not as isolated graphics.
- Understand the video — read the transcript and representative frames to learn topic, audience / platform, visual tone, and speaker layout.
- Set the visual language — use the active Design Style, the user's style / reference, a clarified direction, or visual presets from the active MG workflow.
- Choose useful MG moments — add MG only where a visual layer improves comprehension, emphasis, orientation, or pacing.
- Prepare each moment — decide the viewer job, content, visual mechanism, speech span, settled frame, read time, form, background, and composition relationship before creating the MG.
- Create through the active MG workflow — pass the talking-head context into the current environment's MG creation/authoring path. Different viewer jobs, information structures, or visual forms should usually become distinct MGs; reuse only intentionally recurring components.
- Place, review, confirm, then extend — check face, captions, readability, size, and composition. After the first real MG is placed in frame, confirm the effect with the user before expanding, unless they explicitly asked you to finish end-to-end.
Visual identity
Design Style is the video's confirmed visual language. It gives MGs a shared tone, color logic, typography logic, visual density, and motion language. It keeps different MGs in one family without forcing them into the same shape. It does not decide which MGs are useful, when they appear, where they sit, or whether they are transparent / opaque; those remain per-MG editing decisions.
Resolve the visual language before planning MG moments. Use the active MG workflow for the actual style-alignment interaction and implementation details:
- Active Design Style — use it unless the user asks to change the overall style. If Project Context names an active Design Style but does not show details, inspect it once with
manage_design_style action="get" before planning MG moments.
- Specific user style / reference — follow it. If it is custom and not yet confirmed for a batch, use a real planned MG as the sample when the user needs to approve the look.
- Generic or vague direction — quality words such as clean, premium, modern, professional, polished, or YouTube-style emphasis are goals, not a visual language. Follow the active MG workflow's style-alignment gate: prefer visual preset options, or use one representative MG for confirmation when the direction is textual / custom.
- No visual direction — use the active MG workflow to show relevant visual preset options. Talking-head can be used as a catalog filter when available.
- "Directly do" / "don't ask" — choose a concrete temporary direction from the transcript and footage, then continue without user style confirmation. Do not create or update a Design Style from this unconfirmed guess.
Picker is a visual Design Style selector. It shows preset thumbnails so the user can choose a visual direction by sight, instead of describing style in words.
- Call
manage_design_style with action: "list_presets", scenario: "talking-head" when clear, and the user's locale. Use the scenario as a catalog filter, then choose reasonable visual options by preset descriptions and the actual video context.
- Render reasonable returned presets as visual options using the active form/widget route; do not replace thumbnails with text-only style names when visual thumbnails are available.
- The picker is a turn boundary: after showing it, stop and wait for the user's submitted selection.
- When the user picks an option, call
manage_design_style with action: "apply_preset" and the selected presetId, then inspect the applied Design Style with action: "get" before authoring.
- If the user responds with text instead of picking, treat it as user direction and continue with the custom direction path.
Persist only confirmed visual language:
- Picked preset — the user confirmed it by choosing the visual option. Call
manage_design_style action="apply_preset".
- Custom direction — after the user accepts the sample, treat it as the confirmed direction for the current MG work. If the current environment supports saving project Design Styles and the user accepts it as the shared project style, save/apply it with the agreed style facts.
- Unconfirmed guess — do not create or update a Design Style, including when the user said "directly do it".
After applying a preset or confirming a custom direction as the project style, tell the user in one or two natural sentences that this is now the video's visual style, future MGs in this video will follow it by default, and it can be changed or adjusted later.
Where MG is useful
MG meaningfully helps comprehension or orientation when the content has:
- Identity / context labels — speaker name, role, product name, date, source, or a small persistent section label.
- Key information / quotes — a core concept, definition, statistic, conclusion, or key sentence worth emphasizing.
- Structured information — multiple points, steps, comparisons, rankings, lists, or processes.
- Chapter / topic markers — opening titles, section titles, topic transitions, or visual dividers between sections.
- Abstract concepts — cause-effect relationships, cycles, systems, frameworks, or other ideas that are hard to follow verbally.
Repeated Components
One video should usually have one visual language, but not one universal MG shape.
Reuse a Motion Graphic asset only for intentionally recurring instances of the same component: same viewer task, same information structure, same visual form, and content changed through properties. Repeated chapter markers, recurring section labels, or a repeated status badge can share one asset. Different jobs such as an opening title, chapter marker, quote, list, diagram, and CTA should usually be separate assets that share palette, typography, motion tone, spacing, and material treatment.
An accepted first MG proves the visual language works in frame. It is not automatically a template for unrelated MGs.
Per-MG decisions
For talking-head videos, do not start MG creation from transcript timing alone. Inspect the target frame first: transcript tells you what and when; the frame tells you form, placement, and background.
Before creating the MG, make four linked editor decisions. They prepare the active MG workflow and the later timeline placement.
| Decision |
Question |
Output |
| Content |
What idea deserves a visual layer? |
Message or visual fact expressed by the MG. |
| Timing |
When should it land with the speech? |
Timeline start, duration, read time, and internal motion beats. |
| Form and placement |
What kind of MG is it, and where can it live safely? |
MG form / size, then timeline placement after asset creation. |
| Background |
Is this an overlay on the talking-head shot, or its own moment? |
Transparent overlay or opaque / full-screen beat. |
1. Content
Choose what the MG expresses, not just what text it repeats. The content may be a speaker identity, distilled quote, key term, statistic, list, comparison, relationship diagram, chapter marker, or another visual representation of the point.
2. Timing
Choose the timeline anchor first. The MG should land with the relevant speech beat or section boundary, not trail after the speaker has already made the point. Use find_transcript; pass includeWordTimestamps: true when the MG has internal rhythm such as list items appearing one by one or multi-step reveals.
Write internal timing values relative to the MG's own start time. The timeline item start is the absolute video position; internal timing is the MG-internal rhythm after that start. Exit when the point is fully made.
3. Form and placement
Choose the MG form and likely placement region before creating the asset. The active MG workflow creates the graphic; place the finished asset on the video canvas afterward.
Placement principles:
- Protect the subject and safe zones. Avoid the speaker's face, head, hair, glasses, mouth, chin, important products or objects, relevant hand gestures, captions/subtitles, and existing on-screen elements.
- Keep the caption/subtitle area clear. If captions may appear, bottom overlays must sit above the caption band, not compete with or cover subtitles.
- Separate overlays from full-screen MGs. Subject/safe-zone protection applies to overlays on top of A-roll. A full-screen MG is an intentional visual beat that replaces the A-roll for its duration, so it may cover the speaker and background.
- Keep the composition intentional. The MG should support the speaker and message. It should not look like a random sticker, compete with the face, or make the frame feel unbalanced.
Common forms and areas:
| Content type | Common form | Common area |
| ---------------------------- | ---------------------------------------------------
…(truncated)
1---2name: talking-head-guide3description: Guide for editing videos where the primary content is people talking — talking-head / 口播, interview / 访谈, lecture, tutorial, podcast, course content, and similar talking-driven formats. Use when the user wants speech editing on a talking video (剪口播 / 口播剪辑 / 去口癖 / clean up fillers / smooth speech), motion graphics layered onto talking video (口播加 MG / 加动画), or B-roll on a talking video (加 B-roll / add B-roll). For motion graphics specifically, use this together with the active Motion Graphics skill/workflow available in the current ChatCut environment — this skill adds talking-specific guidance (speech-rhythm timing, frame-aware placement, subject/caption protection, placement verification).4---5
6# Talking Head Video Editing
7
8## What this skill covers
9
10**Required input**: an existing talking-head / 口播 video uploaded to the project. If the user wants to start without one (e.g., generate a fresh talking-head from scratch), this skill doesn't apply.
11
12**When the user enters this workflow without a source video uploaded yet, ask only for missing treatment or preference decisions in ordinary WorkBuddy chat.** Load `asset-import` for the source media and follow its WorkBuddy upload strategy, including its editor upload fallback when local automation is unavailable.
13
14If this workflow creates, targets, or opens a ChatCut project, follow the expert's editor handoff rules before nontrivial edits and again before final delivery when the visible editor may no longer match the project.
15
16Independent treatments that can be applied to talking-head videos. Pick the ones that match what the user wants — not all are needed every time.
17
18- **A-roll editing** (中文称 **语音剪辑** / 含 **去口癖、停顿、重复**) — transcript-based speech editing. Common operations include cleanup, highlight extraction, restructure, opening hook, and others as needed for the aligned outcome.
19- **Motion graphics overlay** (英文展示给用户时写全称 **Motion Graphics**,不要缩成 "MG";中文产品术语固定为 **MG 动画**——不要叫"动效""字幕条""动态字幕"等其它说法) — reinforce key information, structured content, and topic transitions with on-screen motion graphics
20- **B-roll** (industry term — keep as "B-roll" in any language, do not translate) — cover jump cuts or visualize what's being said
21- **Background music** (中文 **背景音乐**) — set mood and smooth micro-gaps
22- **Captions** (中文 **字幕**) — on-screen text for accessibility
23- **AI Voice Isolation** (中文 **AI 人声隔离**) — clean or isolate spoken human voice with DeepFilterNet3, picture untouched. Use the visible `isolate_voice` tool when available.
24
25> 用户语言为中文时,在 widget options / choices options / 对话文案里**严格使用上面括号里的产品术语**——别自己再翻译一遍,会跟产品其它地方对不上。
26
27## What shapes the edit
28
29Beyond picking treatments, a talking-head edit is shaped by several orthogonal variables. When the user's ask is vague, these are what's worth clarifying first:
30
31- **Target** — platform (YouTube / TikTok / Shorts / ...), desired length, and final canvas aspect ratio
32- **Which treatments to apply** — the treatments above are optional; don't assume all of them apply
33- **Pacing / tone** — tight / energetic / formal / casual; brand or voice preferences if stated. (For MG visual style, follow the active Motion Graphics skill/workflow.)
34
35When more than one of these variables is missing, ask them together in one concise WorkBuddy message. Do not run a fixed questionnaire or ask for decisions already visible in the project or source material.
36
37Resolve and apply the final canvas before placing the first visual item. An explicit user ratio wins; otherwise use the platform convention when it is unambiguous, or the primary source ratio when neither platform nor ratio is specified. Update the target timeline with `manage_timelines` and read it back before visual assembly.
38
39## Order of execution
40
41When multiple treatments have been aligned with the user, they depend on each other and must be finalized in dependency order. This section is **only relevant after alignment** — it doesn't tell you what to start with on a fresh request.
42
43The speech timing (set by A-roll editing) anchors everything downstream — MG placement, B-roll cut-covers, music duration, and caption sync all reference the final speech timeline.
44
45So: finalize A-roll editing before committing any visual, audio, or text layer. Don't write captions against pre-edit speech, don't cut music to pre-edit length, don't place MG against timing that will shift.
46
47**You must confirm the result with the user after each major step before starting the next**, unless the user has explicitly asked to run end-to-end without stopping. Key checkpoints when multiple treatments apply: after A-roll editing finalizes the speech timing; before MG creation (confirm style and direction, and, when it isn't obvious, whether it sits over the video as an overlay or takes the whole frame); after MG placement; same pattern for B-roll, music, and captions. **Don't bundle multiple checkpoints into one response — confirm each step separately.** An upstream mistake forces redoing everything downstream (for example, an MG placed against pre-cleanup timing must be repositioned after the timeline shifts).
48
49---
50
51## A-roll editing
52
53### Scenario
54
55In a talking-head workflow, the first step is usually A-roll editing: editing the original spoken footage.
56
57A-roll edits are ultimately applied to the timeline and change what the viewer actually hears and sees. However, the editing decisions should usually start from the transcript, because the core question is: what spoken content should the viewer hear, and what should be removed, compressed, or reordered?
58
59### Common A-roll tasks
60
61A-roll editing is not only cleanup. First decide what spoken-content task the user is asking for, then choose the editing strategy and tools.
62
63Common tasks:
64
65- **Cleanup** — remove mistakes, repeated attempts, verbal habits, filler words, and meaningless pauses so the speech becomes clearer and more natural.
66- **Highlight extraction** — pull the most valuable, opinionated, emotional, or topic-relevant moments from longer footage.
67- **Restructure** — reorder spoken content, such as moving the conclusion earlier, grouping by topic, or combining scattered parts into a clearer structure.
68- **Hook / short version** — use a strong claim, result, conflict, or question from the source as the opening, or compress long content into a shorter version.
69- **Target-script / script alignment** — match, keep, and reorder spoken content according to a user-provided target script, target paragraph, or desired content.
70
71Cleanup is the most common task and the one most likely to fail from bad boundary decisions. It is described in detail below. Other tasks get shorter rules, but still follow the shared A-roll principles: complete meaning, clear boundaries, and natural listening flow.
72
73### Shared A-roll principles
74
75These principles apply to all A-roll tasks, not only cleanup.
76
77- **Decide the task before choosing the tool.** Do not let tool availability change the editing strategy.
78- **Edit by complete semantic units.** Whenever possible, move/delete/keep complete sentences, complete ideas, complete answers, or complete steps. Do not cut out a half-sentence just because a few words match.
79- **When the task names what to keep, trim to that boundary.** The inverse of the rule above, for any task that specifies which content to keep — restoring a specific sentence, matching a target script, pulling a named highlight, building a version: keep exactly the requested span. Trim the kept range to start and end at the requested words and drop the off-script head/tail of the source `[sN]` segment it sits in; keeping a whole segment for one requested sentence is over-keeping that drags in unrequested speech. This applies only when the task names what to keep — never to open-ended cleanup, where you keep complete units (above).
80- **Do not stitch unfinished fragments across retakes.** Do not combine incomplete pieces from different attempts into one artificial sentence. This does not make the earlier attempt disposable: keep a complete useful lead-in, setup, contrast, category, evaluation, or context if it is not repeated later and can naturally connect to the later complete retake.
81- **Preserve connective tissue.** List labels, contrast words, subjects, verbs, and adjacent source words are not filler when removing them makes a kept idea ungrammatical, abrupt, or misleading. Trim the smallest span that keeps the line speakable.
82- **Keep listening flow natural.** The result should still have natural phrasing and breathing room. Do not make sentences feel glued together just to make them "clean."
83- **Be conservative when boundaries are uncertain.** If unsure whether a cut harms meaning, logic, or listening flow, keep it or make a smaller cut.
84- **Confirm complex changes first.** For complex restructuring, aggressive shortening, structural changes, or generated hooks, confirm target length, structure direction, and what to preserve with the user before editing.
85- **Explain content, never indices.** You MUST NOT explain edits to the user with internal addresses such as `[sN]`, `[cN]`, `[gap]`, word indices, clip ids, or segment ids. The user cannot see those addresses and will not understand what they mean. Use the actual spoken content, a short quote, or a plain-language description of the edit.
86- **Never name a screen position for a panel.** When you invite the user to review or fine-tune the result, call it "the Transcript panel" (中文「文字稿面板」) — never a direction (left / right / side / 左侧 / 右侧). The layout is rearrangeable and the panel does not sit in a fixed corner.
87
88### Cleanup goals and decisions
89
90#### What good cleanup means
91
92Good cleanup does not mean making the video as short as possible, and it does not mean rewriting the speaker into a different script.
93
94Good cleanup means:
95
96- The logic stays coherent
97- The expression becomes clearer
98- The audio feels natural
99- Obvious mistakes, repeated attempts, meaningless stalls, and filler are removed
100- The speaker's intent, tone, and natural rhythm are preserved
101
102Bad cleanup usually falls into two failure modes:
103
104- Under-cleaning: obvious mistakes, repetition, long pauses, or filler remain.
105- Over-cleaning: sentences are cut off, meaning is missing, rhythm becomes too hard, or the result sounds stitched together.
106
107Default principle: remove defects without changing meaning; make speech smoother, not harder; prefer small local cuts over whole-sentence or whole-segment deletion; when unsure whether a cut harms meaning, keep it.
108
109#### How to judge common cleanup cases
110
111Below are the common cleanup categories and how to make editing decisions for each.
112
113##### Meaningless filler words
114
115Fillers fall into two categories.
116
117The first category is clearly meaningless hesitation sounds. These are usually safe to remove:
118
119- `um`
120- `uh`
121- `er`
122- `ah`
123- `呃`
124- `额`
125
126When they do not carry special meaning, use `clean_script` first for bulk cleanup.
127
128The second category depends on context and must not be removed by word list alone:
129
130- `so`
131- `like`
132- `然后`
133- `就是`
134- `嗯`
135- `啊`
136- `那个`
137- `那`
138- `对`
139- `所以`
140- `但是`
141
142How to decide:
143
144- If the word is only hesitation or padding, remove it.
145- If it carries sequence, continuation, contrast, cause, reference, response, emphasis, or natural tone, keep it.
146- If removing it makes the surrounding words sound hard-spliced, keep it or only compress the pause.
147- If unsure, keep it.
148
149Examples:
150
151- `um, I think this solves the main problem` -> remove `um`.
152- `It works like a checklist` -> keep `like`; it is a comparison.
153- `The upload failed, so we retried it` -> keep `so`; it carries cause/result.
154- `right after the call, send the recap` -> keep `right`; it modifies timing.
155- `然后我们再看第二点` -> keep `然后`; it marks sequence.
156
157##### Retakes and repeated attempts
158
159A retake is when the speaker retries the same intended idea because they misspoke, got stuck, forgot words, or restarted. Retake cleanup is not "delete repeated text." The goal is to keep one complete, natural, logically coherent version of the intended idea.
160
161Use this decision path:
162
1631. Decide whether it is really a retake.
164 Treat it as a retake only when multiple attempts are trying to say the same intended idea. Do not treat it as a normal retake when the repetition is intentional emphasis, a rhetorical beat, a structural marker, or a second pass that adds new information or tone.
1652. Define the complete version to keep.
166 A complete version may include more than the main content sentence. It may need a lead-in, connector, section marker, topic setup, contrast, qualifier, subject, object, or conclusion. These are not filler when the kept content depends on them.
1673. Cut only the failed or covered part.
168 Remove only words that are wrong, dangling, abandoned, or fully covered by the kept version. The cut boundary starts at the repeated or failed idea, not automatically at the earlier transition, setup, or continuous speech. If earlier speech contains useful context that the kept version does not repeat, keep it.
1694. Choose the best complete attempt.
170 If several attempts are complete, usually prefer the later one because it is often closer to the speaker's intended take. But do not choose the last attempt mechanically. If the later attempt is missing needed context, structure, subject, object, or conclusion, keep the more complete version or preserve the missing lead-in from the earlier attempt.
171
172A repeated lead-in is redundant only when another equivalent lead-in remains naturally connected to the kept content. If removing every copy makes the result lose structure or sound abrupt, keep one natural copy and remove only the extra restarts. Do not stitch unfinished fragments from different attempts into one artificial sentence.
173
174Examples are patterns, not a closed list:
175
176- Local false start inside a kept sentence:
177 `There, there's no After Effects, no Premiere, no DaVinci Resolve learning.`
178 Keep the complete sentence, but remove the abandoned restart:
179 `There's no After Effects, no Premiere, no DaVinci Resolve learning.`
180 Do not keep the stray first word just because the full sentence is otherwise useful.
181- Repeated structural lead-in:
182 `And secondly, ... and secondly, we're introducing a brand new UI.`
183 Remove the extra restart, but keep one natural lead-in attached to the kept content:
184 `And secondly, we're introducing a brand new UI.`
185 Do not delete every structural marker and leave only:
186 `We're introducing a brand new UI.`
187- Useful setup before a failed ending:
188 `Then the next one is different from comedy. It is popular on Disney Plus. It is called...`
189 Later retake:
190 `It is a popular Disney Plus show called Love Story.`
191 Keep useful setup that the later retake does not repeat, and cut from the failure point:
192 `Then the next one is different from comedy. It is a popular Disney Plus show called Love Story.`
193
194##### False starts and unfinished fragments
195
196Use `false starts / unfinished fragments` for this category. `False start` is the more natural editing/transcription term for a speaker beginning a phrase and then restarting or abandoning it; `unfinished fragment` makes the dangling half-sentence case explicit.
197
198Only remove a fragment when it clearly does not form useful information.
199
200Safe to remove:
201
202- The speaker abandons the thought and a complete version appears later.
203- The segment is only a dangling phrase, such as "this is actually..." with no completion.
204- It is clearly the leftover beginning of a failed attempt.
205
206Do not remove:
207
208- A sentence that is imperfect but contains useful information.
209- A lead-in that provides the subject, object, or context needed later.
210- Content that provides setup, contrast, conclusion, emotion, or tone.
211
212If only part of a sentence or segment is wrong, do not delete the useful content around it. Remove only the bad word, phrase, or pause; if a local cut cannot sound natural, keep the segment.
213
214##### Pauses and breaths
215
216Pause cleanup should default to compression, not zeroing out. Spoken video needs natural breathing room.
217
218Default rules:
219
220- Obvious long pauses over 0.8-1s: usually compress to about 0.3s.
221- Between sentences: keep about 0.3-0.5s so listeners can hear natural phrasing.
222- Around topic shifts, contrast, or emphasis: keep slightly longer pauses when needed; do not make the delivery too rushed.
223- Short breaths inside one sentence: if they are normal breathing, do not remove them.
224- Clear long pauses inside one sentence: compress them, but not so tightly that adjacent words sound glued together.
225- Long pauses before a retake: if the failed attempts around it are removed, remove the pause with them.
226- If the user provides explicit thresholds, follow them. For example, if the user says "only process pauses over 0.8s and keep at least 0.3s", do not process natural pauses under 0.8s.
227
228How to operate on pauses:
229
230- For batch pause cleanup across the timeline or track, use `clean_script`. This is the default path for compressing many long pauses.
231- Translate common user wording into `clean_script` pause rules:
232 - "Tighter breaths" / "compress pauses" / "compress anything over 0.3s to 0.3s" → `silence: "compress:300"` (or the requested cap).
233 - "Restore some breathing room" / "do not make it too rushed" / "keep at least 0.5s" → `silence: "restore:500"` (or the requested minimum).
234 - "Make all pauses around 0.5s" → `silence: "normalize:500"`.
235 - "Keep pauses between 0.3s and 0.8s" → `silence: "range:300-800"`.
236 Any rule that makes a pause longer — `restore`, `normalize`, or the lower bound in `range` — never invents new silence. It only recovers pause time that already existed at that exact spot in the original recording. If the original pause was shorter than the requested value, it stops at the original pause length.
237- You do not need to call `read_script({ showSilence: true })` before batch pause cleanup. By default, `timeline.md` hides silence markers, but `clean_script` can still detect and rewrite silences internally.
238- Use `read_script({ showSilence: true })` only when you need to inspect or manually adjust a specific pause. Then edit the visible marker: `~~[silence=0.8s]~~` to fully cut it, `[silence=0.8s→0.2s]` to compress it, or leave it untouched to keep it.
239- After semantic edits, review the final clean `timeline.md`. If the final pacing still has many long pauses, run `clean_script only="silence"`; if only one or two pauses feel wrong, use `showSilence: true` and adjust those manually.
240
241Script gap primitive note:
242
243- Do not create an accidental `[gap]` on the primary video track as a pacing pause. A Script `[gap]` means no source is playing; on the only visible video track it renders as black. If pacing needs breathing room, preserve or restore source silence with `clean_script` / `[silence=...]`, cover the moment with B-roll/MG/a full-frame visual beat, or intentionally declare the black beat in the plan.
244
245### Other A-roll task guidance
246
247#### Highlight extraction
248
249Highlight extraction is not about making the content as short as possible. It is about selecting the most valuable spoken content according to the user's criteria.
250
251Rules:
252
253- First identify the highlight standard: opinion, conclusion, story, emotion, conflict, tutorial step, data point, or a specific topic.
254- Each highlight should be understandable on its own. Do not remove the subject, setup, question, or conclusion needed to understand it.
255- Do not keep only a short punchy sentence if the surrounding context is required for it to make sense.
256- If the user asks for a specific topic, remove other topics. If the user asks for the "best" or "most exciting" moments, prioritize information density and expression strength.
257- After extracting highlights, usually clean up the kept segments so the final result is polished.
258
259#### Restructure
260
261Restructure means changing the order of spoken content. It does not mean freely breaking sentences apart.
262
263Rules:
264
265- First confirm the target structure: chronological, by topic, by question, conclusion-first, tutorial steps, or short-form pacing.
266- Move complete semantic units: complete sentences, ideas, answers, or steps.
267- Do not split one sentence so the first half appears in one place and the second half elsewhere.
268- After moving content, check whether connectors still work, such as "so," "but," "next," or "this."
269- If the user asks for major restructuring without specifying the target structure, confirm before editing.
270
271#### Hook / short version
272
273Hook / short version work aims to make the opening more compelling or compress long content into a shorter but still complete version.
274
275Rules:
276
277- Prefer pulling the hook from the original footage: a strong claim, result, conflict, question, counterintuitive statement, or emotionally strong moment.
278- If a new hook or new narration must be generated, confirm the direction with the user first.
279- For short versions, do not cut only by duration. First identify the main line to preserve: problem, core point, key reasons, and conclusion.
280- Short versions can remove examples, repetition, and setup, but must keep the logic needed for the point to hold.
281- If the user gives a target duration, try to match it. If duration and semantic completeness conflict, explain the tradeoff.
282
283#### Target-script / script alignment
284
285Target-script / script alignment means cutting the final spoken content according to a user-provided script, target paragraph, or desired content.
286
287Rules:
288
289- The target script is the main constraint: prioritize content that matches the target meaning.
290- Natural spoken paraphrases are acceptable, but do not include surrounding content that the target does not ask for.
291- If the source has multiple similar versions, choose the most complete, natural, and target-aligned version.
292- If target order differs from source order, reorder as needed, but move complete semantic units.
293- If the target script omits source context, follow the target. Do not add long surrounding context unless the result would be incomprehensible without it.
294
295#### Building versions, highlights, and excerpts — stay on Script
296
297Highlight, short version, excerpt, hook, restructure, and making several versions are all transcript-content tasks: drive them through Script (`read_script` → edit `timeline.md` → `apply_script`), never by looking up timestamps and placing source clips manually.
298
299- Pick the starting point by where the content comes from. Versions on the current timeline: trim or reorder `timeline.md` and `apply_script`. A version on its own timeline (the user asked for separate timelines, or wants each version independently editable/exportable): `manage_timelines` action=duplicate — the copy carries the content and its script, so you immediately `read_script` → trim → `apply_script` on it. Building fresh from library assets: `manage_timelines` action=create, add the source asset, then drive it through Script.
300- To bring in source content the current cut no longer shows (a hook line, a segment needed for another version), read `library/<filename>.md`, copy the needed `[sN]` line(s) into `timeline.md` where they belong, and `apply_script`. This is how you pull source content onto the timeline — through Script.
301- For multiple versions on one track: list every version's `[sN]` segments in `timeline.md` in version order, one version after another, then `apply_script` once. Reuse is just repetition — the same `[sN]` segment may appear in more than one version, and repeating the line replays that source range again.
302- Never look up timestamps with `find_transcript` and place spoken content with `edit_item` / `split_item`. If you are converting transcript segments into source frame or second ranges, you are off the editing surface — return to Script. `edit_item` / `find_transcript` are only for non-transcript placement such as MG overlays and B-roll visual timing.
303
304**Check each version against its request.** After assembling a version, highlight, or excerpt, re-read the result end to end and confirm every requested sentence is present, in the requested order, with no extra source carried in. Fix any dropped, duplicated, or out-of-order content before finishing.
305
306### A-roll / transcript-based editing workflow
307
308Use this flow for any A-roll task driven by transcript meaning.
309
3101. Start with orientation. Call `read_script`, then read `timeline.md` once to understand the user's goal, the content structure, and whether fixed fillers or long pauses are present. If you will run `clean_script`, do not build the full semantic edit from this pre-clean read.
3112. For cleanup tasks, run the mechanical cleanup pass before semantic editing when fixed fillers or long pauses are present. Use `clean_script` for fixed hesitation sounds (`um`, `uh`, `er`, `ah`, `呃`, `额`) and batch pause compression. If both are present, use the default `clean_script` pass so both are handled together. Do not use this step for context-dependent fillers, retakes, repeated sentences, or anything that needs meaning.
3123. After `clean_script`, always read the refreshed clean `timeline.md` before semantic editing. Use this refreshed file as the source of truth; `clean_script` changes the canonical timeline and rematerializes the script, so previously read text may be stale. Do not edit from memory based on the pre-clean script. Then edit `timeline.md` with semantic judgment: choose the best retake, clean false starts, remove repeated or failed attempts, preserve useful setup and context, reorder content when needed, and keep the speech natural. For long transcripts, work one clear section at a time if that improves judgment accuracy.
3134. Apply the edit with `apply_script`. If apply fails, fix the markdown error or stale state, re-read the current `timeline.md` if needed, and apply again.
3145. Review the edited result. After a real `apply_script`, read the regenerated clean `timeline.md` and check what the viewer will actually hear: broken logic, missing context, over-deletion, missed cleanup, wrong order, or pauses that feel too tight or too long. Fix clear problems only. If the final result still needs batch pause adjustment, use `clean_script only="silence"`. Use `read_script({ showSilence: true })` only for manual adjustment of specific pauses.
315
316### What transcript editing actually changes
317
318Editing `timeline.md` is not just changing displayed text. It describes which source media ranges should play on the timeline.
319
320`[sN]` rows are ASR segments, not semantic units. A complete sentence, idea, retake, or transition may span several `[sN]` rows, and one `[sN]` row may contain only part of a sentence. Before deciding what to delete or keep, mentally reconstruct the complete spoken sentence or idea across adjacent rows.
321
322- Each spoken-text line maps to a playable source range.
323- Inline `~~...~~` removes the corresponding audible audio range.
324- Deleting a whole line removes that whole spoken segment.
325- Moving/reordering lines changes playback order.
326- `apply_script` applies the result back to the timeline.
327- Start/end trims may remain as one trimmed clip.
328- Deleting words or pauses in the middle of a sentence splits the original clip into multiple new clips: one kept range before the deletion and one kept range after it.
329- Moving spoken content also creates a new clip at the destination.
330- More clips after middle deletions or moves are expected and usually correct. Do not describe that as a "fragmentation problem" or as proof that word deletion is unsupported.
331
332### Tool boundaries
333
334Choose the editing goal and content boundaries first, then choose the tool. Do not let tool availability change the editing strategy.
335
336- `clean_script`: use for mechanical first-pass cleanup: bulk removal of fixed meaningless fillers and batch silence compression/adjustment. It can process silence even when `timeline.md` is currently rendered without silence markers. Do not use it for context-dependent fillers, retakes, repeated sentences, or semantic decisions.
337- `read_script` + `apply_script`: the main transcript-based editing surface. Use it for real semantic editing: deleting words, sentences, pauses, reordering, or pulling library content onto the timeline.
338- `manage_transcript` action `fix`: only fixes ASR mistakes or speaker attribution. It does not cut audio and does not change what the viewer hears.
339- Caption SEGMENTATION (分句 / where pages break) is INDEPENDENT of the transcript and controlled by two per-word primitives only: to SPLIT one card into two, set `display_text` `forcePageBreak:true` on the word that should START the new card; to MERGE a card up into the previous one, set `display_text` `keepWithPrevious:true` on that card's FIRST word (works for any break — no box resizing, no wordsPerPage fiddling). To drop a repeated/false-start word, use `display_text` `hidden:true`. Box width / fontSize / `wordsPerPage` are style & density knobs, NOT per-boundary segmentation levers — do not widen the box or raise wordsPerPage to merge or split a specific card. NEVER edit the transcript to fix a caption line break — `manage_transcript fix` is only for an ASR-misheard WORD (content), not layout. `read_captions` shows each page's `break=` reason and per-word keys for these edits.
340- `find_transcript`: only locates when a phrase is spoken. It does not edit. If the next step is cutting spoken content, return to Script.
341- `Edit` / `Write`: use these to modify `timeline.md`. The edit only reaches the timeline after `apply_script`.
342
343Script details to preserve:
344
345- `read_script` materializes `timeline.md` (current cut) and `library/<filename>.md` (full read-only source transcripts) in the workspace.
346- Single-word audible deletion is supported with inline strike syntax, such as `[s1] 过去~~呢~~一个月`.
347- Silence markers are hidden by default. Use `clean_script` for batch pause cleanup. Use `read_script({ showSilence: true })` only to expose `[silence=Ns]` markers for precise manual edits such as `~~[silence=0.8s]~~` or `[silence=0.8s→0.2s]`.
348- `find_transcript` can locate a phrase for visual timing; it is not the editing surface. Do not use `find_transcript` + `split_item` / `edit_item` to cut, place, or assemble transcript-based clips — this includes highlights, hooks, excerpts, and multi-version cuts. All spoken-content selection, placement, and reuse happens in Script (`read_script` → edit `timeline.md` → `apply_script`).
349
350---
351
352## MG Overlay
353
354### Goal
355
356Motion graphics layered into A-roll reinforce what the speaker is conveying — deepening the audience's impression of the key points and helping them grasp content that's hard to land through speech alone. Complete A-roll editing first; MG timing is based on the post-edit timeline.
357
358This section only adds talking-head timing, frame-composition, subject/caption protection, and review constraints. For visual style alignment, MG creation or authoring, implementation constraints, editable properties, asset sizing, and verification, use the active Motion Graphics skill/workflow available in the current ChatCut environment.
359
360### MG workflow
361
362For talking-head MG work, treat the video as one edited piece, not as isolated graphics.
363
3641. **Understand the video** — read the transcript and representative frames to learn topic, audience / platform, visual tone, and speaker layout.
3652. **Set the visual language** — use the active Design Style, the user's style / reference, a clarified direction, or visual presets from the active MG workflow.
3663. **Choose useful MG moments** — add MG only where a visual layer improves comprehension, emphasis, orientation, or pacing.
3674. **Prepare each moment** — decide the viewer job, content, visual mechanism, speech span, settled frame, read time, form, background, and composition relationship before creating the MG.
3685. **Create through the active MG workflow** — pass the talking-head context into the current environment's MG creation/authoring path. Different viewer jobs, information structures, or visual forms should usually become distinct MGs; reuse only intentionally recurring components.
3696. **Place, review, confirm, then extend** — check face, captions, readability, size, and composition. After the first real MG is placed in frame, confirm the effect with the user before expanding, unless they explicitly asked you to finish end-to-end.
370
371### Visual identity
372
373Design Style is the video's confirmed visual language. It gives MGs a shared tone, color logic, typography logic, visual density, and motion language. It keeps different MGs in one family without forcing them into the same shape. It does not decide which MGs are useful, when they appear, where they sit, or whether they are transparent / opaque; those remain per-MG editing decisions.
374
375Resolve the visual language before planning MG moments. Use the active MG workflow for the actual style-alignment interaction and implementation details:
376
377- **Active Design Style** — use it unless the user asks to change the overall style. If Project Context names an active Design Style but does not show details, inspect it once with `manage_design_style action="get"` before planning MG moments.
378- **Specific user style / reference** — follow it. If it is custom and not yet confirmed for a batch, use a real planned MG as the sample when the user needs to approve the look.
379- **Generic or vague direction** — quality words such as clean, premium, modern, professional, polished, or YouTube-style emphasis are goals, not a visual language. Follow the active MG workflow's style-alignment gate: prefer visual preset options, or use one representative MG for confirmation when the direction is textual / custom.
380- **No visual direction** — use the active MG workflow to show relevant visual preset options. Talking-head can be used as a catalog filter when available.
381- **"Directly do" / "don't ask"** — choose a concrete temporary direction from the transcript and footage, then continue without user style confirmation. Do not create or update a Design Style from this unconfirmed guess.
382
383Picker is a visual Design Style selector. It shows preset thumbnails so the user can choose a visual direction by sight, instead of describing style in words.
384
3851. Call `manage_design_style` with `action: "list_presets"`, `scenario: "talking-head"` when clear, and the user's `locale`. Use the scenario as a catalog filter, then choose reasonable visual options by preset descriptions and the actual video context.
3862. Render reasonable returned presets as visual options using the active form/widget route; do not replace thumbnails with text-only style names when visual thumbnails are available.
3873. The picker is a turn boundary: after showing it, stop and wait for the user's submitted selection.
3884. When the user picks an option, call `manage_design_style` with `action: "apply_preset"` and the selected `presetId`, then inspect the applied Design Style with `action: "get"` before authoring.
3895. If the user responds with text instead of picking, treat it as user direction and continue with the custom direction path.
390
391Persist only confirmed visual language:
392
393- **Picked preset** — the user confirmed it by choosing the visual option. Call `manage_design_style action="apply_preset"`.
394- **Custom direction** — after the user accepts the sample, treat it as the confirmed direction for the current MG work. If the current environment supports saving project Design Styles and the user accepts it as the shared project style, save/apply it with the agreed style facts.
395- **Unconfirmed guess** — do not create or update a Design Style, including when the user said "directly do it".
396
397After applying a preset or confirming a custom direction as the project style, tell the user in one or two natural sentences that this is now the video's visual style, future MGs in this video will follow it by default, and it can be changed or adjusted later.
398
399### Where MG is useful
400
401MG meaningfully helps comprehension or orientation when the content has:
402
403- **Identity / context labels** — speaker name, role, product name, date, source, or a small persistent section label.
404- **Key information / quotes** — a core concept, definition, statistic, conclusion, or key sentence worth emphasizing.
405- **Structured information** — multiple points, steps, comparisons, rankings, lists, or processes.
406- **Chapter / topic markers** — opening titles, section titles, topic transitions, or visual dividers between sections.
407- **Abstract concepts** — cause-effect relationships, cycles, systems, frameworks, or other ideas that are hard to follow verbally.
408
409### Repeated Components
410
411One video should usually have one visual language, but not one universal MG shape.
412
413Reuse a Motion Graphic asset only for intentionally recurring instances of the same component: same viewer task, same information structure, same visual form, and content changed through properties. Repeated chapter markers, recurring section labels, or a repeated status badge can share one asset. Different jobs such as an opening title, chapter marker, quote, list, diagram, and CTA should usually be separate assets that share palette, typography, motion tone, spacing, and material treatment.
414
415An accepted first MG proves the visual language works in frame. It is not automatically a template for unrelated MGs.
416
417### Per-MG decisions
418
419For talking-head videos, do not start MG creation from transcript timing alone. Inspect the target frame first: transcript tells you what and when; the frame tells you form, placement, and background.
420
421Before creating the MG, make four linked editor decisions. They prepare the active MG workflow and the later timeline placement.
422
423| Decision | Question | Output |
424| ---------------------- | --------------------------------------------------------------- | --------------------------------------------------------------- |
425| **Content** | What idea deserves a visual layer? | Message or visual fact expressed by the MG. |
426| **Timing** | When should it land with the speech? | Timeline start, duration, read time, and internal motion beats. |
427| **Form and placement** | What kind of MG is it, and where can it live safely? | MG form / size, then timeline placement after asset creation. |
428| **Background** | Is this an overlay on the talking-head shot, or its own moment? | Transparent overlay or opaque / full-screen beat. |
429
430#### 1. Content
431
432Choose what the MG expresses, not just what text it repeats. The content may be a speaker identity, distilled quote, key term, statistic, list, comparison, relationship diagram, chapter marker, or another visual representation of the point.
433
434#### 2. Timing
435
436Choose the timeline anchor first. The MG should land with the relevant speech beat or section boundary, not trail after the speaker has already made the point. Use `find_transcript`; pass `includeWordTimestamps: true` when the MG has internal rhythm such as list items appearing one by one or multi-step reveals.
437
438Write internal timing values relative to the MG's own start time. The timeline item start is the absolute video position; internal timing is the MG-internal rhythm after that start. Exit when the point is fully made.
439
440#### 3. Form and placement
441
442Choose the MG form and likely placement region before creating the asset. The active MG workflow creates the graphic; place the finished asset on the video canvas afterward.
443
444Placement principles:
445
446- **Protect the subject and safe zones.** Avoid the speaker's face, head, hair, glasses, mouth, chin, important products or objects, relevant hand gestures, captions/subtitles, and existing on-screen elements.
447- **Keep the caption/subtitle area clear.** If captions may appear, bottom overlays must sit above the caption band, not compete with or cover subtitles.
448- **Separate overlays from full-screen MGs.** Subject/safe-zone protection applies to overlays on top of A-roll. A full-screen MG is an intentional visual beat that replaces the A-roll for its duration, so it may cover the speaker and background.
449- **Keep the composition intentional.** The MG should support the speaker and message. It should not look like a random sticker, compete with the face, or make the frame feel unbalanced.
450
451Common forms and areas:
452
453| Content type | Common form | Common area |
454| ---------------------------- | ---------------------------------------------------
455
456…(truncated)