Interactive POV Vlog
Mission
用于第一人称互动短片:用严格 POV 约束设计陪伴、宠物、虚拟角色或双主体喜剧互动短片,重点控制镜头身份、可见主体、台词/SFX、跨段连续和合规边界。
This workflow is about viewpoint and interaction. It can be combined with flova-immersive-camera-language when detailed POV/handheld camera grammar is needed.
JarvisHub Execution Model
- 主 Agent 负责编排:读取画布事实、整理任务 brief、按阶段同步 TodoWrite 进度,并把已确认的文本成果写入画布;媒体生成、等待、拼接和评审交给具备相应能力的执行 agent 或当前可用工具。TodoWrite 不是画布写入前置条件。
- 图像、视频和拼接交给具备媒体能力的执行 agent:brief 给出稳定输出身份、用途、真实参考 URL 和关键约束;需要下游引用时必须等待真实
imageUrl / videoUrl,不能把提交态当完成态。
- 多模态验收交给
critic sub-agent:只读取真实媒体并评审,不生成、不补素材。
- 当前已知 canvas 工具集没有通用音频生成、音频驱动口型、末帧抽取或视频直改工具;旁白、BGM、字幕、末帧承接和精准 lip sync 只能作为后期合成计划或
blocked 项,除非本轮工具列表明确暴露对应能力。
When To Use
Use this skill for:
- first-person interactive vlog,
- companion POV scene,
- pet or mascot interacting with the camera,
- two-subject caught-in-the-act comedy,
- adult relationship POV with safe non-explicit affection,
- handheld low-angle POV,
- multi-segment POV continuity.
Do not use it for generic cinematic narratives, third-person dialogue scenes, or explicit sexual content.
Canvas-Native Boundary
Follow the JarvisHub Execution Model for reference analysis, character/scene assets, video clips, covers, and assembly. Treat audio and start-frame carryover as available only when exposed by this turn's tool list. If exact cross-segment frame carryover is unavailable, return explicit continuity prompts and reference-frame instructions.
Intake
Collect:
- subject type: adult person, pet, mascot/IP-like original character, or two-subject comedy,
- reference images or text description,
- scene/theme,
- duration: 15s / 30s / 45s / 60s,
- aspect ratio,
- output language,
- whether dialogue, off-camera voice, SFX, cover images, or social copy are needed,
- interaction boundaries and forbidden actions.
If the user provides a real person reference for affectionate/romantic POV, confirm the scene is non-explicit and consensual. Do not use public figures for intimate content.
Modes
| Mode |
Use For |
Core Rule |
single_subject_pov |
one person/pet/character interacts with camera |
only one main subject visible |
two_subject_comedy_pov |
pet + child-like fictional character, two pets, two mascots |
camera is adult/observer POV; observer body not visible |
multi_segment_pov |
30s+ across 2-4 segments |
plan full arc before splitting |
POV Rules
For single_subject_pov:
- The camera is the observer's eyes or phone.
- Only one main subject appears in frame.
- Observer may appear only as partial hand/sleeve if necessary.
- No full second person, no observer face/body, no third-person objective shot.
For two_subject_comedy_pov:
- Camera is the adult/observer viewpoint.
- Two subjects remain visible and interact in the scene.
- Observer body should not appear.
- Off-camera adult voice may be included if requested or central to the gag.
Content Boundaries
Allowed:
- warm daily affection,
- playful teasing,
- safe closeness,
- pet/mascot interaction,
- light comedy conflict,
- caught-in-the-act reactions.
Not allowed:
- explicit sexual content,
- nudity,
- coercion, intoxicated incapacity, threats, or non-consensual framing,
- underage romantic/sexual implication,
- public figure intimate imitation,
- realistic harm to children, pets, or vulnerable subjects.
Rewrite risky user phrasing into safe, affectionate, daily-life interaction before generating prompts.
Workflow
- Analyze references and classify subject type.
- Lock final POV spec.
- Generate or bind character/subject references.
- Plan full story arc.
- Draft segment storyboard.
- Generate scene/prop references if needed.
- Generate POV clips.
- Extract/carry final frame for next segment when supported.
- Generate cover/social assets if requested.
- Assemble and QA.
Pause after spec, reference character sheet, storyboard, scene/prop assets, first clip/batch, cover, and final assembly.
Subject References
For a person:
- extract face, hair, outfit, style, light mood, but avoid oversexualized body detail,
- create/bind a stable character sheet if needed,
- preserve adult, consensual, non-explicit framing.
For a pet:
- extract species, fur/coat, face, size, temperament,
- prefer natural interaction: pet approaches, looks up, gets fed, plays, tilts head.
For mascot/original IP-like character:
- extract shape, material, palette, texture,
- avoid protected character names/logos unless user owns/permits them.
For two-subject comedy:
- create separate subject cards plus a scale relationship card if same-frame proportion matters.
Story Arc
For 15s single-subject POV:
- 0-2s: establish environment and subject noticing camera,
- 2-5s: first interaction,
- 5-10s: escalation or turn,
- 10-15s: emotional/comedy landing.
For 30s+:
- plan the full timeline before writing segment details,
- each 15s segment must have a complete mini-beat,
- adjacent segments need clear continuity: pose, prop state, subject position, emotional state.
Two-Subject Comedy Structure
Before storyboard, answer:
- What visible trouble happened?
- Where is the evidence?
- What does the off-camera observer react to?
A good gag has:
- visible physical evidence,
- two subjects' conflicting reactions,
- a final caught-in-the-act freeze or reaction.
Examples:
- food scattered,
- toy dismantled,
- forbidden place occupied,
- blanket pulled apart,
- spilled flour,
- stolen snack.
Keep it harmless and light. No injury or danger.
Segment Fields
Each segment/shot must include:
segment_id,
duration,
pov_type,
subject_ids,
scene_element_id,
visible_action_timeline,
camera_behavior,
dialogue_or_offscreen_voice,
sfx,
continuity_in,
continuity_out,
references.
If dialogue is generated in-video, write exact lines and timing. If exact lip-sync is not supported, mark dialogue for audio/assembly.
Camera Language
Single-subject:
- handheld phone POV,
- eye-level or seated POV,
- gentle push-in,
- small reactions from camera,
- close framing but no incoherent body occlusion.
Pet/mascot:
- low-angle handheld at subject height,
- small forward/backward camera reactions,
- hand enters only for feeding/petting/toy interaction.
Two-subject comedy:
- low observer viewpoint,
- action-driven pan/tilt/follow,
- quick push-in at the caught moment,
- visible evidence in the same frame as subject reaction.
Avoid aimless drifting. Camera movement must respond to the subject.
Prompt Rules
Use reference placeholders. Keep prompts concrete and visible.
Single-subject template:
<<<image_1>>> first-person POV from the observer, only the main subject visible, [scene], [subject action and expression], [camera movement], [dialogue/SFX if any], no subtitles, no random text, no full second person, no third-person shot.
Two-subject template:
<<<image_1>>> <<<image_2>>> low-angle handheld POV from the adult observer, observer not visible, both subjects in frame, [visible trouble evidence], [interaction beats], [off-camera line/SFX], no subtitles, no random text, no harm.
For multi-segment continuity:
- include previous final state in the next prompt,
- use previous final frame/start frame if tool supports it,
- do not reset clothing, prop state, or emotion between segments.
Audio Rules
Use in-video audio only when the selected tool supports it reliably.
Otherwise separate:
- SFX layer,
- off-camera voice,
- dialogue/VO,
- BGM if requested.
Do not add background music by default for POV realism unless user requests it.
QA
Check:
- POV perspective not broken,
- observer does not appear beyond allowed partial hand/sleeve,
- only intended subjects appear,
- content stays safe and non-explicit,
- two-subject comedy has visible evidence,
- segment continuity holds,
- no random subtitles/text,
- audio/dialogue aligns with action,
- references remain consistent.
Fix the smallest failed unit: reference sheet, single segment prompt, one audio line, or one transition.
1---2name: flova-interactive-pov-vlog3description: Use when 用户要第一人称互动短片、陪伴感 POV、宠物/角色互动、双主体抓包喜剧、低角度手持 POV 或多段沉浸式 Vlog。4---56# Interactive POV Vlog78## Mission910用于第一人称互动短片:用严格 POV 约束设计陪伴、宠物、虚拟角色或双主体喜剧互动短片,重点控制镜头身份、可见主体、台词/SFX、跨段连续和合规边界。1112This workflow is about viewpoint and interaction. It can be combined with `flova-immersive-camera-language` when detailed POV/handheld camera grammar is needed.1314## JarvisHub Execution Model1516- 主 Agent 负责编排:读取画布事实、整理任务 brief、按阶段同步 TodoWrite 进度,并把已确认的文本成果写入画布;媒体生成、等待、拼接和评审交给具备相应能力的执行 agent 或当前可用工具。TodoWrite 不是画布写入前置条件。17- 图像、视频和拼接交给具备媒体能力的执行 agent:brief 给出稳定输出身份、用途、真实参考 URL 和关键约束;需要下游引用时必须等待真实 `imageUrl` / `videoUrl`,不能把提交态当完成态。18- 多模态验收交给 `critic` sub-agent:只读取真实媒体并评审,不生成、不补素材。19- 当前已知 canvas 工具集没有通用音频生成、音频驱动口型、末帧抽取或视频直改工具;旁白、BGM、字幕、末帧承接和精准 lip sync 只能作为后期合成计划或 `blocked` 项,除非本轮工具列表明确暴露对应能力。2021## When To Use2223Use this skill for:2425- first-person interactive vlog,26- companion POV scene,27- pet or mascot interacting with the camera,28- two-subject caught-in-the-act comedy,29- adult relationship POV with safe non-explicit affection,30- handheld low-angle POV,31- multi-segment POV continuity.3233Do not use it for generic cinematic narratives, third-person dialogue scenes, or explicit sexual content.3435## Canvas-Native Boundary3637Follow the JarvisHub Execution Model for reference analysis, character/scene assets, video clips, covers, and assembly. Treat audio and start-frame carryover as available only when exposed by this turn's tool list. If exact cross-segment frame carryover is unavailable, return explicit continuity prompts and reference-frame instructions.3839## Intake4041Collect:4243- subject type: adult person, pet, mascot/IP-like original character, or two-subject comedy,44- reference images or text description,45- scene/theme,46- duration: 15s / 30s / 45s / 60s,47- aspect ratio,48- output language,49- whether dialogue, off-camera voice, SFX, cover images, or social copy are needed,50- interaction boundaries and forbidden actions.5152If the user provides a real person reference for affectionate/romantic POV, confirm the scene is non-explicit and consensual. Do not use public figures for intimate content.5354## Modes5556| Mode | Use For | Core Rule |57| --- | --- | --- |58| `single_subject_pov` | one person/pet/character interacts with camera | only one main subject visible |59| `two_subject_comedy_pov` | pet + child-like fictional character, two pets, two mascots | camera is adult/observer POV; observer body not visible |60| `multi_segment_pov` | 30s+ across 2-4 segments | plan full arc before splitting |6162## POV Rules6364For `single_subject_pov`:6566- The camera is the observer's eyes or phone.67- Only one main subject appears in frame.68- Observer may appear only as partial hand/sleeve if necessary.69- No full second person, no observer face/body, no third-person objective shot.7071For `two_subject_comedy_pov`:7273- Camera is the adult/observer viewpoint.74- Two subjects remain visible and interact in the scene.75- Observer body should not appear.76- Off-camera adult voice may be included if requested or central to the gag.7778## Content Boundaries7980Allowed:8182- warm daily affection,83- playful teasing,84- safe closeness,85- pet/mascot interaction,86- light comedy conflict,87- caught-in-the-act reactions.8889Not allowed:9091- explicit sexual content,92- nudity,93- coercion, intoxicated incapacity, threats, or non-consensual framing,94- underage romantic/sexual implication,95- public figure intimate imitation,96- realistic harm to children, pets, or vulnerable subjects.9798Rewrite risky user phrasing into safe, affectionate, daily-life interaction before generating prompts.99100## Workflow1011021. Analyze references and classify subject type.1032. Lock final POV spec.1043. Generate or bind character/subject references.1054. Plan full story arc.1065. Draft segment storyboard.1076. Generate scene/prop references if needed.1087. Generate POV clips.1098. Extract/carry final frame for next segment when supported.1109. Generate cover/social assets if requested.11110. Assemble and QA.112113Pause after spec, reference character sheet, storyboard, scene/prop assets, first clip/batch, cover, and final assembly.114115## Subject References116117For a person:118119- extract face, hair, outfit, style, light mood, but avoid oversexualized body detail,120- create/bind a stable character sheet if needed,121- preserve adult, consensual, non-explicit framing.122123For a pet:124125- extract species, fur/coat, face, size, temperament,126- prefer natural interaction: pet approaches, looks up, gets fed, plays, tilts head.127128For mascot/original IP-like character:129130- extract shape, material, palette, texture,131- avoid protected character names/logos unless user owns/permits them.132133For two-subject comedy:134135- create separate subject cards plus a scale relationship card if same-frame proportion matters.136137## Story Arc138139For 15s single-subject POV:140141- 0-2s: establish environment and subject noticing camera,142- 2-5s: first interaction,143- 5-10s: escalation or turn,144- 10-15s: emotional/comedy landing.145146For 30s+:147148- plan the full timeline before writing segment details,149- each 15s segment must have a complete mini-beat,150- adjacent segments need clear continuity: pose, prop state, subject position, emotional state.151152## Two-Subject Comedy Structure153154Before storyboard, answer:155156- What visible trouble happened?157- Where is the evidence?158- What does the off-camera observer react to?159160A good gag has:161162- visible physical evidence,163- two subjects' conflicting reactions,164- a final caught-in-the-act freeze or reaction.165166Examples:167168- food scattered,169- toy dismantled,170- forbidden place occupied,171- blanket pulled apart,172- spilled flour,173- stolen snack.174175Keep it harmless and light. No injury or danger.176177## Segment Fields178179Each segment/shot must include:180181- `segment_id`,182- `duration`,183- `pov_type`,184- `subject_ids`,185- `scene_element_id`,186- `visible_action_timeline`,187- `camera_behavior`,188- `dialogue_or_offscreen_voice`,189- `sfx`,190- `continuity_in`,191- `continuity_out`,192- `references`.193194If dialogue is generated in-video, write exact lines and timing. If exact lip-sync is not supported, mark dialogue for audio/assembly.195196## Camera Language197198Single-subject:199200- handheld phone POV,201- eye-level or seated POV,202- gentle push-in,203- small reactions from camera,204- close framing but no incoherent body occlusion.205206Pet/mascot:207208- low-angle handheld at subject height,209- small forward/backward camera reactions,210- hand enters only for feeding/petting/toy interaction.211212Two-subject comedy:213214- low observer viewpoint,215- action-driven pan/tilt/follow,216- quick push-in at the caught moment,217- visible evidence in the same frame as subject reaction.218219Avoid aimless drifting. Camera movement must respond to the subject.220221## Prompt Rules222223Use reference placeholders. Keep prompts concrete and visible.224225Single-subject template:226227```text228<<<image_1>>> first-person POV from the observer, only the main subject visible, [scene], [subject action and expression], [camera movement], [dialogue/SFX if any], no subtitles, no random text, no full second person, no third-person shot.229```230231Two-subject template:232233```text234<<<image_1>>> <<<image_2>>> low-angle handheld POV from the adult observer, observer not visible, both subjects in frame, [visible trouble evidence], [interaction beats], [off-camera line/SFX], no subtitles, no random text, no harm.235```236237For multi-segment continuity:238239- include previous final state in the next prompt,240- use previous final frame/start frame if tool supports it,241- do not reset clothing, prop state, or emotion between segments.242243## Audio Rules244245Use in-video audio only when the selected tool supports it reliably.246247Otherwise separate:248249- SFX layer,250- off-camera voice,251- dialogue/VO,252- BGM if requested.253254Do not add background music by default for POV realism unless user requests it.255256## QA257258Check:259260- POV perspective not broken,261- observer does not appear beyond allowed partial hand/sleeve,262- only intended subjects appear,263- content stays safe and non-explicit,264- two-subject comedy has visible evidence,265- segment continuity holds,266- no random subtitles/text,267- audio/dialogue aligns with action,268- references remain consistent.269270Fix the smallest failed unit: reference sheet, single segment prompt, one audio line, or one transition.