Video Understanding Capability
When to Use
Use this capability to analyze video files or links for hook, pacing, scene structure, retention devices, CTA, captions, audio direction, visual motifs, and reusable creative patterns.
Preferred Implementation
- Use video metadata extraction, frame sampling, transcript extraction, speech-to-text, and multimodal analysis when available.
- Use manual visual review only when automated extraction is insufficient.
- Use Computer Use only to view videos that cannot be accessed via file/API/browser automation.
Inputs
- Video file or link.
- Platform.
- Campaign objective.
- Target audience.
- Questions to answer.
- Usage rights or constraints.
Outputs
- Video Structure Breakdown.
- Hook and Retention Analysis.
- CTA and Offer Analysis.
- Scene or Frame Notes.
- Reusable Creative Template.
- Risk and rights notes.
Security and Ethics
- Do not download or reuse protected video assets without permission.
- Do not claim access to private analytics unless the user provides them.
- Do not produce deepfake, deceptive, or undisclosed synthetic media guidance.