Finish Interview
Use canonical timeline metadata and the shared renderer. Do not handwrite a
one-off ffmpeg graph unless the project's downstream overlay format cannot be
represented by the timeline.
Workflow
- Read
05_timeline/timeline.json, 06_review/review_patch.json, the latest
preview, and any downstream card/caption artifacts.
- Measure before changing:
- create a representative frame contact sheet;
- measure integrated LUFS, true peak, and LRA;
- probe video/audio start times and durations;
- record
avg_frame_rate, r_frame_rate, and time_base; treat a phone
source whose cadence cannot be proven CFR as VFR for sync review.
- when choosing framing automatically, run
.build/debug/videoos-studio-cli interview-reframe --source=<video> --in-us=<n> --out-us=<n>;
- treat its face/eye/yaw/hand result as a proposal, then inspect representative frames.
- Keep editorial timing unchanged unless the user separately requests a cut.
For a titled J-cut opening, stage information instead of presenting exterior,
title, and dialogue at the same instant: establish the exterior, let the title
read, then fade dialogue in before the title exits. Enter the interview on a
moving source frame, never a frozen bridge frame. Preserve the complete final
utterance and a short moving-picture tail; do not extend the ending with a
freeze frame.
- Write review patch operations:
- one
change_visual_transform per affected video clip;
- one timeline-level
change_audio_finish operation.
Studio can author both operations from the selected clip's
インタビュー仕上げ inspector. Adjusted framing appears immediately in the
Viewer; press 画角を保留 and MAを保留 before Apply & Preview.
- Apply the patch with
scripts/compile-timeline.ts --patch and require schema
validation to pass.
- Render the clean assembly first. Burn captions and insert question cards
afterward so zoom/crop never scales authored graphics.
- Verify the final preview by full decode, loudness measurement, visual frame
sampling, audio/video duration comparison, and dialogue placement QA. The
duration comparison alone is insufficient: package QA must show
dialogue_timeline_alignment_valid as well as av_drift_valid.
Patch pattern
{
"timeline_version": "1",
"operations": [
{
"op": "change_visual_transform",
"target_clip_id": "CLP_0001",
"visual_transform": {
"zoom": 1.15,
"position": { "x": -144, "y": -39 }
},
"reason": "Increase portrait presence while preserving look room"
},
{
"op": "change_audio_finish",
"audio_finish": {
"preset": "dialogue-clean",
"loudness_target_lufs": -16,
"true_peak_target_dbtp": -1.5
},
"reason": "Improve speech intelligibility and delivery loudness"
}
]
}
position is measured in output pixels. Negative x shifts the image left;
negative y shifts it upward. Start around zoom: 1.10–1.18 for a restrained
talking-head punch-in and inspect hand gestures across the whole cut.
For zoom > 1, position pans inside the available zoom overscan and clamps at
the source boundary. Do not recreate the old post-crop pad/translate graph; it
exposes black edges. The automatic planner normally caps an interview punch-in
at 1.18, reduces it when hands are detected, and allows stronger zoom only when
the detected face is genuinely small.
Audio presets
dialogue-clean: 70 Hz high-pass, gentle noise reduction, mud/presence EQ,
3:1 compression, and measured two-pass loudness normalization.
loudness-only: measured two-pass normalization without cleanup or EQ.
none: disables the declared finish preset.
Defaults target -16 LUFS, LRA 7, and -1.5 dBTP. The dialogue preset keeps
0.3 dB of codec headroom before AAC encoding.
QA contract
- No cropped head, hands, or interviewer look room in representative samples.
- Captions and question cards retain their authored size and position.
- Integrated loudness is within 0.5 LU of target.
- Encoded true peak does not exceed the declared target by more than 0.2 dB.
- Audio/video start at zero and end within one frame.
- Dialogue-only signal stays inside the timeline windows that own dialogue.
adelay=...,atrim=start=0 on one input branch is forbidden because the trim
removes the inserted lead-in; source atrim must precede adelay, and only
the final mixed output may be duration-trimmed.
- For VFR phone footage, first choose the intended video frame, then measure the
video and audio source offsets independently. If frame quantization leaves a
sub-frame residual, keep the chosen picture frame and offset audio by the
measured residual. Do not chase sync by repeatedly moving the video trim
across frame boundaries.
- A proxy or alternate encode made to diagnose a player is diagnostic only. It
must not become the canonical creative output or constrain title/J-cut timing.
- Full ffmpeg decode exits successfully.
- Keep preview output distinct from final package output until caption and
review gates are approved.
1---2name: finish-interview3description: Applies dialogue MA, portrait reframing, and loudness/sync QA to an interview or talking-head rough cut using canonical timeline metadata and a shared renderer.4---56# Finish Interview78Use canonical timeline metadata and the shared renderer. Do not handwrite a9one-off ffmpeg graph unless the project's downstream overlay format cannot be10represented by the timeline.1112## Workflow13141. Read `05_timeline/timeline.json`, `06_review/review_patch.json`, the latest15 preview, and any downstream card/caption artifacts.162. Measure before changing:17 - create a representative frame contact sheet;18 - measure integrated LUFS, true peak, and LRA;19 - probe video/audio start times and durations;20 - record `avg_frame_rate`, `r_frame_rate`, and `time_base`; treat a phone21 source whose cadence cannot be proven CFR as VFR for sync review.22 - when choosing framing automatically, run23 `.build/debug/videoos-studio-cli interview-reframe --source=<video> --in-us=<n> --out-us=<n>`;24 - treat its face/eye/yaw/hand result as a proposal, then inspect representative frames.253. Keep editorial timing unchanged unless the user separately requests a cut.26 For a titled J-cut opening, stage information instead of presenting exterior,27 title, and dialogue at the same instant: establish the exterior, let the title28 read, then fade dialogue in before the title exits. Enter the interview on a29 moving source frame, never a frozen bridge frame. Preserve the complete final30 utterance and a short moving-picture tail; do not extend the ending with a31 freeze frame.324. Write review patch operations:33 - one `change_visual_transform` per affected video clip;34 - one timeline-level `change_audio_finish` operation.35 Studio can author both operations from the selected clip's36 **インタビュー仕上げ** inspector. Adjusted framing appears immediately in the37 Viewer; press **画角を保留** and **MAを保留** before Apply & Preview.385. Apply the patch with `scripts/compile-timeline.ts --patch` and require schema39 validation to pass.406. Render the clean assembly first. Burn captions and insert question cards41 afterward so zoom/crop never scales authored graphics.427. Verify the final preview by full decode, loudness measurement, visual frame43 sampling, audio/video duration comparison, and dialogue placement QA. The44 duration comparison alone is insufficient: package QA must show45 `dialogue_timeline_alignment_valid` as well as `av_drift_valid`.4647## Patch pattern4849```json50{51 "timeline_version": "1",52 "operations": [53 {54 "op": "change_visual_transform",55 "target_clip_id": "CLP_0001",56 "visual_transform": {57 "zoom": 1.15,58 "position": { "x": -144, "y": -39 }59 },60 "reason": "Increase portrait presence while preserving look room"61 },62 {63 "op": "change_audio_finish",64 "audio_finish": {65 "preset": "dialogue-clean",66 "loudness_target_lufs": -16,67 "true_peak_target_dbtp": -1.568 },69 "reason": "Improve speech intelligibility and delivery loudness"70 }71 ]72}73```7475`position` is measured in output pixels. Negative `x` shifts the image left;76negative `y` shifts it upward. Start around `zoom: 1.10–1.18` for a restrained77talking-head punch-in and inspect hand gestures across the whole cut.7879For `zoom > 1`, position pans inside the available zoom overscan and clamps at80the source boundary. Do not recreate the old post-crop pad/translate graph; it81exposes black edges. The automatic planner normally caps an interview punch-in82at 1.18, reduces it when hands are detected, and allows stronger zoom only when83the detected face is genuinely small.8485## Audio presets8687- `dialogue-clean`: 70 Hz high-pass, gentle noise reduction, mud/presence EQ,88 3:1 compression, and measured two-pass loudness normalization.89- `loudness-only`: measured two-pass normalization without cleanup or EQ.90- `none`: disables the declared finish preset.9192Defaults target `-16 LUFS`, LRA 7, and `-1.5 dBTP`. The dialogue preset keeps930.3 dB of codec headroom before AAC encoding.9495## QA contract9697- No cropped head, hands, or interviewer look room in representative samples.98- Captions and question cards retain their authored size and position.99- Integrated loudness is within 0.5 LU of target.100- Encoded true peak does not exceed the declared target by more than 0.2 dB.101- Audio/video start at zero and end within one frame.102- Dialogue-only signal stays inside the timeline windows that own dialogue.103 `adelay=...,atrim=start=0` on one input branch is forbidden because the trim104 removes the inserted lead-in; source `atrim` must precede `adelay`, and only105 the final mixed output may be duration-trimmed.106- For VFR phone footage, first choose the intended video frame, then measure the107 video and audio source offsets independently. If frame quantization leaves a108 sub-frame residual, keep the chosen picture frame and offset audio by the109 measured residual. Do not chase sync by repeatedly moving the video trim110 across frame boundaries.111- A proxy or alternate encode made to diagnose a player is diagnostic only. It112 must not become the canonical creative output or constrain title/J-cut timing.113- Full ffmpeg decode exits successfully.114- Keep preview output distinct from final package output until caption and115 review gates are approved.