AI Avatar Production (Global) — Pipeline 3-Tier, 4 Workflows, QA Score 100
Flagship skill of the AI Content cluster. Covers the full pipeline from zero to publish, voice clone, anti-detection, and region-specific disclosure law.
For newbies
What is an AI Avatar?
An AI Avatar is a video that shows your face (or a stand-in) but uses AI-generated voice and motion. You provide one photo or a short selfie video; the AI produces a final video with natural-looking speech, gestures, and expressions. No filming crew, no studio, no actor required.
What do you need to start?
| Method |
Requirement |
Quality |
| Portrait photo |
1 forward-facing photo, clean background, 1024x1024+ |
Medium — mouth less natural |
| Selfie video |
30s video, looking at the lens, speaking naturally |
Good — better lipsync |
| Custom avatar |
2-5 min recording with teleprompter + lavalier mic |
Excellent — near photo-real |
Minimum gear: Phone with HD front camera + lavalier mic (or headset mic).
How long does it take?
- One single video (60s): 30-60 min (script + render)
- Batch of 10: 1-2 days
- Batch of 30: 4-5 days (with optimized process)
What does it cost?
| Tier |
USD/month |
Output |
| Free |
$0 |
1-3 videos, watermark |
| Pro |
$30-100 |
10-30 videos, no watermark |
| Enterprise |
$200-500+ |
30+ videos, custom avatar, API |
5 common newbie mistakes
- Lipsync drift: Script too fast or voice mismatch -> slow speech 10-15%, use voice clone instead of default voice.
- Voice doesn't sound like you: Sample too short or noisy -> re-record 3-5 minutes in a quiet room with phonetically varied script.
- Video flagged as "AI content": Platform pattern detection -> see Anti-detection section below.
- Blurry / pixelated output: Low-quality input -> use 1024x1024+ photo, natural lighting, no filters.
- Slow render: Free tier queue -> render off-peak (early morning in your timezone = US night) or upgrade to Pro.
Information collection (4 questions max)
Ask up to 4 questions before starting:
- Primary use case? Brand awareness / Sales / Education / Internal training?
- Primary platform? TikTok / YouTube / Facebook / Instagram / LinkedIn / X / Threads?
- Budget tier? Free ($0) / Pro ($30-100/mo) / Enterprise ($200+/mo)?
- Videos per month target? 1-5 / 10-30 / 30+?
Based on the 4 answers, auto-select Tier + Workflow.
If the user has already uploaded reference images, do not ask a long intake form first; classify the images, create the setup/prompt, then ask only for missing assets.
Tier decision — Tools and pricing
| Tier |
Suggested tool |
Price/month |
Quality |
Limit |
Fits |
| Free |
Captions Free, HeyGen Trial, D-ID Trial |
$0 |
6/10 — watermark, limited duration |
1-5 videos, max 60s/video |
Personal test, new freelancers |
| Pro |
HeyGen Creator ($29), Synthesia Starter ($29), ElevenLabs Pro ($22) |
$30-100 |
8/10 — no watermark, HD |
10-30 videos, max 5 min/video |
SME, small agency, content creator |
| Enterprise |
HeyGen Business ($89+), Synthesia Enterprise (custom) |
$200-500+ |
9.5/10 — custom avatar, API, priority render |
30+ videos, unlimited |
Large agency, large brand, e-learning |
Quick recommendations:
- Just starting: HeyGen Trial (1 video free, full experience)
- Serious but budget-limited: Captions Pro ($10/mo) for lipsync + ElevenLabs Starter ($5) for voice
- Scale fast: HeyGen Creator + ElevenLabs Pro = best price/quality combo
- Enterprise: Synthesia Enterprise + ElevenLabs Scale
Workflow 1: Single Avatar Production
One video, end-to-end in 30-60 minutes.
6-step process
| Step |
Task |
Tool |
Time |
| 1. Script |
150-300 words for a 60s video |
Skill 04-script-video-global |
10 min |
| 2. Voice |
Generate or use voice clone |
ElevenLabs / HeyGen Voice |
5 min |
| 3. Avatar |
Pick stock avatar or upload your media |
HeyGen / Synthesia / D-ID |
3 min |
| 4. Render |
Combine voice + avatar, choose background, gestures |
Tool from step 3 |
5-15 min (render) |
| 5. QA |
QA Score 100 review (see section below) |
Manual review |
5 min |
| 6. Publish |
Export MP4 -> post to platform |
Manual / Scheduler |
2 min |
Script template for AI Avatar (60s)
[HOOK — 3s] Curiosity hook, frame the problem
[PROBLEM — 10s] Describe the customer pain
[SOLUTION — 25s] Your solution, 2-3 key points
[PROOF — 12s] Numbers, testimonial, result
[CTA — 10s] Concrete action: "Link in bio for..."
Workflow 2: Multi-language translate
One source video -> many languages for global rollout. Use cases: DTC brand expanding markets, multi-language courses, multi-country agency work.
Tool comparison
| Tool |
Languages |
Price |
Notes |
| Rask AI |
130+ |
$50/mo (Pro) |
Best for translate today |
| HeyGen Translate |
40+ |
Included Creator+ |
Built-in, convenient |
| Synthesia Translate |
35+ |
Included Enterprise |
Best for e-learning |
Process
- Create source video (Workflow 1)
- Upload to translate tool (Rask AI recommended)
- Pick target language — tool auto-translates and lipsyncs
- Review with a native speaker
- Export and publish per market
Caveat: Tonal languages (Mandarin, Vietnamese, Thai) have weaker lipsync. Workaround: produce native voice clone + native avatar per language.
See full disclosure law per region in the variant files.
Workflow 3: Batch Production
30 videos in 5 days — assembly-line process.
Detailed timeline
| Day |
Task |
Output |
Tool |
| Day 1 |
Script batch — write 10 scripts from template |
10 scripts (.md) |
Skill 04-script-video-global + AI assist |
| Day 2 |
Voice batch — render 10 audio files |
10 audio (.mp3) |
ElevenLabs API |
| Day 3 |
Avatar batch — upload audio + avatar, queue render |
10 videos rendering |
HeyGen Batch / Synthesia |
| Day 4 |
QA batch — review 10 videos, fix issues, re-render |
10 QA'd videos |
Manual + QA Score |
| Day 5 |
Publish batch — export, add captions, schedule |
10 videos published |
Buffer / Later / Manual |
Repeat 3 weeks = 30 videos. Or scale Days 1-2 to 15 scripts/week.
Cost estimate batch 30 videos/month
| Tier |
Tool combo |
Monthly cost |
Per-video cost |
| Free |
HeyGen Trial + Captions Free |
$0 (limited 3-5 videos) |
$0 (watermark) |
| Pro |
HeyGen Creator + ElevenLabs Pro |
~$51 |
~$1.70 |
| Enterprise |
HeyGen Business + ElevenLabs Scale |
~$189 |
~$6.30 |
Batch optimization tips
- Templated scripts: 3-5 frameworks, swap the core content
- Voice consistency: One voice clone for the entire series
- Off-peak rendering: Queue overnight to skip the queue
- QA checklist: Print the QA Score, check videos like an assembly line
Workflow 4: Hybrid Real + AI
Real face for trust + AI body for speed.
Use cases
- Real face intro 5s + AI body 55s (save filming time)
- AI video weekdays + Real video weekly (balance quality/effort)
- Real talking head + AI B-roll (studio-grade output)
Assembly + tools
- Film real intro 5-10s (eye contact, natural greeting); use Captions for lipsync fixes
- Create AI for the rest with same outfit/background (HeyGen / Synthesia)
- Edit in CapCut / Premiere (precise cuts, smooth transitions)
- Color match AI to real footage (LUT or DaVinci Resolve free)
Trust gain: Real face up front -> 20-35% more engagement than full-AI.
Voice Clone Protocol
Voice sample requirements
| Criterion |
Requirement |
| Duration |
3-5 minutes |
| Quality |
WAV/FLAC, 44.1kHz+, mono, quiet room |
| Script content |
Phonetically varied passages (all vowels, hard consonants) |
| Emotion |
Read normal, natural, not acted |
Tool comparison
| Tool |
Price |
Quality |
Notes |
| ElevenLabs |
From $5/mo |
9/10 |
Best overall, 30+ languages |
| HeyGen Voice |
Included Creator+ |
6/10 |
Convenient if using HeyGen |
| Resemble AI |
From $99/mo |
7/10 |
Strong API |
| PlayHT |
From $39/mo |
7/10 |
Good for narration |
Consent form template
MANDATORY before cloning anyone's voice.
VOICE USAGE CONSENT
I, [FULL NAME], consent to [COMPANY] using my voice for: [SPECIFIC PURPOSE].
Term: [X months / Until revoked]
Date: [YYYY-MM-DD]
Signature: _______________
Reference: See references/voice-clone-prompts-global.md
Avatar Setup Checklist
Before recording / uploading photo or video for an AI avatar:
Reference Image -> Avatar Prompt Director
Use this when the user drops one or more reference images and wants to create an avatar, replace a face, adapt brand colors, add a logo, or create the prompt before uploading assets into a tool.
Classify Input Images
| Image type |
Role |
Requirement |
| Style ref |
Mood, lighting, background, outfit, camera angle |
Do not use as identity unless requested |
| Face ref |
Identity preservation / face replacement |
1-3 clear face images, no filter, front + 3/4 angle |
| Selfie video |
Better custom avatar / natural lipsync |
30s-2 min, looking at camera, speaking naturally |
| Logo/palette |
Personal/company brand adaptation |
PNG/SVG logo + 2-4 hex colors |
| Product/location |
Prop or avatar environment |
Clear product label or location/background image |
Multiple Images = Multiple Flows
## Avatar Flows
| Flow | Input image | Role | Suggested tool | Missing assets |
|------|-------------|------|----------------|----------------|
| A | style-01 | style/background | Design Master -> HeyGen | face ref, logo |
| B | face-01 | identity | HeyGen custom avatar | script, voice sample |
- If every image is a different style direction, create a separate prompt for each flow.
- If images support one avatar, group by role: style + face + logo + palette + product.
- Ask for each next asset explicitly: face image, selfie video, logo, hex colors, script, voice sample.
Prompt Setup Output
## Avatar Prompt Setup — Flow A
- Style ref:
- Face ref:
- Brand assets:
- Target platform:
- Tool route:
## Copy-Paste Visual Prompt
[English prompt for avatar/source image generation]
## Upload Next
- Face/selfie video:
- Logo:
- Brand colors:
- Voice sample:
- Script:
For a static personal avatar only, route to 30-design-master-global personal-brand mode. For talking-head video, continue this workflow.
Anti-detection for FB / IG / TikTok / YouTube
5 detection signals and fixes
| Signal |
Platforms flagging |
Fix |
| Stiff face, no natural blinking |
FB, IG |
Use selfie video over photo; pick avatars with micro-expressions |
| Monotone voice, no natural pauses |
TikTok, FB |
Use voice clone (natural pacing) over default TTS |
| Fully static background |
FB, IG |
Add slight noise/grain, or use real-world background |
| Isolated motion (only mouth moves) |
TikTok |
Pick avatars with gesture (hands, head); use HeyGen v3+ |
| Metadata flagged as AI tool |
YouTube (monetize) |
Re-export through CapCut (strips metadata); add color grade |
Techniques to add "human feel"
- Add film grain / noise: 2-5% in CapCut or Premiere
- Zoom and crop: 5-10% crop with subtle motion (Ken Burns)
- Color grade: Apply film LUT or manually grade — avoid "too clean"
- Text overlay: Add subtitles, callouts, stickers to cover AI weak spots
- B-roll insert: Drop 2-3 b-roll clips (product, lifestyle) every 15-20s
- Sound design: Background music + light SFX (immersion + masks AI voice)
Per platform
- TikTok: Most lenient — content quality wins over AI checks
- Facebook / Instagram: Moderate scrutiny — anti-detection matters
- LinkedIn: Practically no detection — best fit for AI avatars
- YouTube: Strict for monetized videos — must disclose per YPP policy
CRITICAL: NEVER use AI avatars to impersonate real people without consent. This is illegal in most jurisdictions and grounds for permanent platform bans.
Ethics and Disclosure — Region selector
Disclosure laws differ dramatically by region. Pick the matching variant:
| Region |
Variant file |
Key law |
| US / Canada |
variants/01-us.md |
FTC Endorsement Guides (16 CFR Part 255), 2023 update |
| EU / EEA / UK |
variants/02-eu.md |
EU AI Act Article 50 (always disclose) + UCPD + GDPR |
| Southeast Asia |
variants/03-sea.md |
Per-country: ASAS (SG), AKARI (ID), DTI (PH), MCMC (MY), TH |
| Latin America |
variants/04-latam.md |
CONAR + LGPD (BR), PROFECO (MX), AAIP (AR), per-country |
ALWAYS read the matching variant BEFORE publishing AI avatar content in that region. Penalties range from warning to multi-thousand-USD fines per influencer (US) and can stack under EU AI Act + GDPR.
Universal disclosure rule of thumb
When in doubt, disclose. Disclosure is rarely penalized; non-disclosure can be.
"This video uses AI Avatar technology for visuals and voice."
Placement: video description, first 3 seconds on-screen text, OR platform "AI-generated" tag (where available — Meta, TikTok, YouTube all now support this).
QA Score — 100 points
Scorecard
| # |
Criterion |
Points |
Description |
| 1 |
Lipsync |
/10 |
Mouth tracks speech within 0.2s |
| 2 |
Voice match |
/10 |
Voice sounds like the speaker (if clone) or natural (if TTS) |
| 3 |
Visual quality |
/10 |
Sharp image, no artifacts, no blur |
| 4 |
Background |
/10 |
Background suits context, no render glitches |
| 5 |
Lighting |
/10 |
Even light, no harsh shadows, matches background |
| 6 |
Gesture |
/10 |
Natural, no jitters, hand/head movement present |
| 7 |
Script flow |
/10 |
Hook -> Problem -> Solution -> CTA |
| 8 |
Disclosure |
/10 |
AI disclosure compliant with region (see variant) |
| 9 |
Platform fit |
/10 |
Correct aspect ratio, duration, format for platform |
| 10 |
CTA |
/10 |
Clear call-to-action, easy to execute |
Action thresholds
| Tier |
Score |
Action |
| Excellent |
90-100 |
Publish now |
| Good |
70-89 |
Publish, note improvements for next round |
| Needs fix |
50-69 |
Fix items scoring under 7, then re-render |
| Redo |
<50 |
Rebuild from script + voice + avatar |
Output template
# AI Avatar Video — [Title] | [Region variant] | [Date]
1. Workflow used: [Single / Translate / Batch / Hybrid]
2. Script: [Content, 150-300 words]
3. Voice: [Tool] — [Voice ID / clone name] — Consent: [Yes / N/A]
4. Avatar: [Tool] — [Avatar ID / custom]
5. QA Score: [X]/100 (10 criteria)
6. Disclosure (per region variant): [Text + placement]
7. Publish: [Platform] — [Aspect ratio] — [Link]
Quality checklist
Related skills
25-voice-clone-podcast-global — voice clone deep-dive + podcast pipeline
04-script-video-global — script writing for AI avatar
26-thought-leadership-content-global — content strategy for personal brand
references/ai-video-disclosure-global — full legal reference
references/voice-clone-prompts-global — voice clone training prompts
Global Skill 24 (AI Avatar Production) | Over Powers Agency | v1.1.0
1---2name: 24-ai-avatar-production-global3description: Use when a PERSONAL brand needs AI avatar video at scale — three tool tiers, four workflows for single avatar, translation, batch, and hybrid, reference image intake, face, style, logo, and palette replacement, voice clone pairing, anti-detection, and a QA score, with disclosure-law variants for US FTC, EU AI Act, SEA, and LATAM, covering HeyGen and Synthesia. Trigger on 'AI avatar', 'HeyGen video', 'Synthesia', 'talking head AI video', 'translate my videos with AI', 'I cannot be on camera every day'. Not for — the words the avatar says, see `04-script-video-global`; audio-only voice clone and podcast, see `25-voice-clone-podcast-global`; a company product video edit, see `44-video-editor-brief-global`.4license: MIT5---6
7# AI Avatar Production (Global) — Pipeline 3-Tier, 4 Workflows, QA Score 100
8
9> Flagship skill of the AI Content cluster. Covers the full pipeline from zero to publish, voice clone, anti-detection, and region-specific disclosure law.
10
11---
12
13## For newbies
14
15### What is an AI Avatar?
16
17An AI Avatar is a video that shows your face (or a stand-in) but uses AI-generated voice and motion. You provide one photo or a short selfie video; the AI produces a final video with natural-looking speech, gestures, and expressions. No filming crew, no studio, no actor required.
18
19### What do you need to start?
20
21| Method | Requirement | Quality |
22|--------|-------------|---------|
23| Portrait photo | 1 forward-facing photo, clean background, 1024x1024+ | Medium — mouth less natural |
24| Selfie video | 30s video, looking at the lens, speaking naturally | Good — better lipsync |
25| Custom avatar | 2-5 min recording with teleprompter + lavalier mic | Excellent — near photo-real |
26
27**Minimum gear:** Phone with HD front camera + lavalier mic (or headset mic).
28
29### How long does it take?
30
31- **One single video (60s):** 30-60 min (script + render)
32- **Batch of 10:** 1-2 days
33- **Batch of 30:** 4-5 days (with optimized process)
34
35### What does it cost?
36
37| Tier | USD/month | Output |
38|------|-----------|--------|
39| Free | $0 | 1-3 videos, watermark |
40| Pro | $30-100 | 10-30 videos, no watermark |
41| Enterprise | $200-500+ | 30+ videos, custom avatar, API |
42
43### 5 common newbie mistakes
44
451. **Lipsync drift:** Script too fast or voice mismatch -> slow speech 10-15%, use voice clone instead of default voice.
462. **Voice doesn't sound like you:** Sample too short or noisy -> re-record 3-5 minutes in a quiet room with phonetically varied script.
473. **Video flagged as "AI content":** Platform pattern detection -> see Anti-detection section below.
484. **Blurry / pixelated output:** Low-quality input -> use 1024x1024+ photo, natural lighting, no filters.
495. **Slow render:** Free tier queue -> render off-peak (early morning in your timezone = US night) or upgrade to Pro.
50
51---
52
53## Information collection (4 questions max)
54
55Ask up to 4 questions before starting:
56
571. **Primary use case?** Brand awareness / Sales / Education / Internal training?
582. **Primary platform?** TikTok / YouTube / Facebook / Instagram / LinkedIn / X / Threads?
593. **Budget tier?** Free ($0) / Pro ($30-100/mo) / Enterprise ($200+/mo)?
604. **Videos per month target?** 1-5 / 10-30 / 30+?
61
62> Based on the 4 answers, auto-select Tier + Workflow.
63> If the user has already uploaded reference images, do not ask a long intake form first; classify the images, create the setup/prompt, then ask only for missing assets.
64
65---
66
67## Tier decision — Tools and pricing
68
69| Tier | Suggested tool | Price/month | Quality | Limit | Fits |
70|------|----------------|-------------|---------|-------|------|
71| **Free** | Captions Free, HeyGen Trial, D-ID Trial | $0 | 6/10 — watermark, limited duration | 1-5 videos, max 60s/video | Personal test, new freelancers |
72| **Pro** | HeyGen Creator ($29), Synthesia Starter ($29), ElevenLabs Pro ($22) | $30-100 | 8/10 — no watermark, HD | 10-30 videos, max 5 min/video | SME, small agency, content creator |
73| **Enterprise** | HeyGen Business ($89+), Synthesia Enterprise (custom) | $200-500+ | 9.5/10 — custom avatar, API, priority render | 30+ videos, unlimited | Large agency, large brand, e-learning |
74
75**Quick recommendations:**
76- **Just starting:** HeyGen Trial (1 video free, full experience)
77- **Serious but budget-limited:** Captions Pro ($10/mo) for lipsync + ElevenLabs Starter ($5) for voice
78- **Scale fast:** HeyGen Creator + ElevenLabs Pro = best price/quality combo
79- **Enterprise:** Synthesia Enterprise + ElevenLabs Scale
80
81---
82
83## Workflow 1: Single Avatar Production
84
85> One video, end-to-end in 30-60 minutes.
86
87### 6-step process
88
89| Step | Task | Tool | Time |
90|------|------|------|------|
91| 1. Script | 150-300 words for a 60s video | Skill `04-script-video-global` | 10 min |
92| 2. Voice | Generate or use voice clone | ElevenLabs / HeyGen Voice | 5 min |
93| 3. Avatar | Pick stock avatar or upload your media | HeyGen / Synthesia / D-ID | 3 min |
94| 4. Render | Combine voice + avatar, choose background, gestures | Tool from step 3 | 5-15 min (render) |
95| 5. QA | QA Score 100 review (see section below) | Manual review | 5 min |
96| 6. Publish | Export MP4 -> post to platform | Manual / Scheduler | 2 min |
97
98### Script template for AI Avatar (60s)
99
100```
101[HOOK — 3s] Curiosity hook, frame the problem
102[PROBLEM — 10s] Describe the customer pain
103[SOLUTION — 25s] Your solution, 2-3 key points
104[PROOF — 12s] Numbers, testimonial, result
105[CTA — 10s] Concrete action: "Link in bio for..."
106```
107
108---
109
110## Workflow 2: Multi-language translate
111
112> One source video -> many languages for global rollout. Use cases: DTC brand expanding markets, multi-language courses, multi-country agency work.
113
114### Tool comparison
115
116| Tool | Languages | Price | Notes |
117|------|-----------|-------|-------|
118| Rask AI | 130+ | $50/mo (Pro) | Best for translate today |
119| HeyGen Translate | 40+ | Included Creator+ | Built-in, convenient |
120| Synthesia Translate | 35+ | Included Enterprise | Best for e-learning |
121
122### Process
123
1241. Create source video (Workflow 1)
1252. Upload to translate tool (Rask AI recommended)
1263. Pick target language — tool auto-translates and lipsyncs
1274. Review with a native speaker
1285. Export and publish per market
129
130**Caveat:** Tonal languages (Mandarin, Vietnamese, Thai) have weaker lipsync. Workaround: produce native voice clone + native avatar per language.
131
132> **See full disclosure law per region in the variant files.**
133
134---
135
136## Workflow 3: Batch Production
137
138> 30 videos in 5 days — assembly-line process.
139
140### Detailed timeline
141
142| Day | Task | Output | Tool |
143|-----|------|--------|------|
144| **Day 1** | Script batch — write 10 scripts from template | 10 scripts (.md) | Skill `04-script-video-global` + AI assist |
145| **Day 2** | Voice batch — render 10 audio files | 10 audio (.mp3) | ElevenLabs API |
146| **Day 3** | Avatar batch — upload audio + avatar, queue render | 10 videos rendering | HeyGen Batch / Synthesia |
147| **Day 4** | QA batch — review 10 videos, fix issues, re-render | 10 QA'd videos | Manual + QA Score |
148| **Day 5** | Publish batch — export, add captions, schedule | 10 videos published | Buffer / Later / Manual |
149
150> Repeat 3 weeks = 30 videos. Or scale Days 1-2 to 15 scripts/week.
151
152### Cost estimate batch 30 videos/month
153
154| Tier | Tool combo | Monthly cost | Per-video cost |
155|------|-----------|--------------|----------------|
156| Free | HeyGen Trial + Captions Free | $0 (limited 3-5 videos) | $0 (watermark) |
157| Pro | HeyGen Creator + ElevenLabs Pro | ~$51 | ~$1.70 |
158| Enterprise | HeyGen Business + ElevenLabs Scale | ~$189 | ~$6.30 |
159
160### Batch optimization tips
161
162- **Templated scripts:** 3-5 frameworks, swap the core content
163- **Voice consistency:** One voice clone for the entire series
164- **Off-peak rendering:** Queue overnight to skip the queue
165- **QA checklist:** Print the QA Score, check videos like an assembly line
166
167---
168
169## Workflow 4: Hybrid Real + AI
170
171> Real face for trust + AI body for speed.
172
173### Use cases
174
175- Real face intro 5s + AI body 55s (save filming time)
176- AI video weekdays + Real video weekly (balance quality/effort)
177- Real talking head + AI B-roll (studio-grade output)
178
179### Assembly + tools
180
1811. Film real intro 5-10s (eye contact, natural greeting); use Captions for lipsync fixes
1822. Create AI for the rest with same outfit/background (HeyGen / Synthesia)
1833. Edit in CapCut / Premiere (precise cuts, smooth transitions)
1844. Color match AI to real footage (LUT or DaVinci Resolve free)
185
186> **Trust gain:** Real face up front -> 20-35% more engagement than full-AI.
187
188---
189
190## Voice Clone Protocol
191
192### Voice sample requirements
193
194| Criterion | Requirement |
195|-----------|-------------|
196| Duration | 3-5 minutes |
197| Quality | WAV/FLAC, 44.1kHz+, mono, quiet room |
198| Script content | Phonetically varied passages (all vowels, hard consonants) |
199| Emotion | Read normal, natural, not acted |
200
201### Tool comparison
202
203| Tool | Price | Quality | Notes |
204|------|-------|---------|-------|
205| ElevenLabs | From $5/mo | 9/10 | Best overall, 30+ languages |
206| HeyGen Voice | Included Creator+ | 6/10 | Convenient if using HeyGen |
207| Resemble AI | From $99/mo | 7/10 | Strong API |
208| PlayHT | From $39/mo | 7/10 | Good for narration |
209
210### Consent form template
211
212> **MANDATORY** before cloning anyone's voice.
213
214```
215VOICE USAGE CONSENT
216
217I, [FULL NAME], consent to [COMPANY] using my voice for: [SPECIFIC PURPOSE].
218Term: [X months / Until revoked]
219Date: [YYYY-MM-DD]
220Signature: _______________
221```
222
223> **Reference:** See `references/voice-clone-prompts-global.md`
224
225---
226
227## Avatar Setup Checklist
228
229Before recording / uploading photo or video for an AI avatar:
230
231- [ ] **Lighting:** Natural light or softbox; no harsh shadows on the face
232- [ ] **Background:** Solid (white / gray) or real environment (office, store)
233- [ ] **Wardrobe:** On-brand; avoid small busy patterns (AI moire)
234- [ ] **Framing:** Chest up; eyes on the upper-third line
235- [ ] **Eye contact:** Look directly at the lens (not the screen)
236- [ ] **Gestures:** Natural; hands can rest or do light gestures
237- [ ] **Resolution:** Minimum 1080p (1920x1080); 4K preferred
238- [ ] **Aspect ratio:** 9:16 (TikTok / Reels), 16:9 (YouTube), 1:1 (Feed)
239- [ ] **File format:** MP4 (H.264) for video, PNG / JPG for photo
240- [ ] **Backup:** Keep originals on cloud (Google Drive / OneDrive) before uploading to the tool
241
242---
243
244## Reference Image -> Avatar Prompt Director
245
246Use this when the user drops one or more reference images and wants to create an avatar, replace a face, adapt brand colors, add a logo, or create the prompt before uploading assets into a tool.
247
248### Classify Input Images
249
250| Image type | Role | Requirement |
251|------------|------|-------------|
252| **Style ref** | Mood, lighting, background, outfit, camera angle | Do not use as identity unless requested |
253| **Face ref** | Identity preservation / face replacement | 1-3 clear face images, no filter, front + 3/4 angle |
254| **Selfie video** | Better custom avatar / natural lipsync | 30s-2 min, looking at camera, speaking naturally |
255| **Logo/palette** | Personal/company brand adaptation | PNG/SVG logo + 2-4 hex colors |
256| **Product/location** | Prop or avatar environment | Clear product label or location/background image |
257
258### Multiple Images = Multiple Flows
259
260```markdown
261## Avatar Flows
262
263| Flow | Input image | Role | Suggested tool | Missing assets |
264|------|-------------|------|----------------|----------------|
265| A | style-01 | style/background | Design Master -> HeyGen | face ref, logo |
266| B | face-01 | identity | HeyGen custom avatar | script, voice sample |
267```
268
269- If every image is a different style direction, create a separate prompt for each flow.
270- If images support one avatar, group by role: style + face + logo + palette + product.
271- Ask for each next asset explicitly: face image, selfie video, logo, hex colors, script, voice sample.
272
273### Prompt Setup Output
274
275```markdown
276## Avatar Prompt Setup — Flow A
277
278- Style ref:
279- Face ref:
280- Brand assets:
281- Target platform:
282- Tool route:
283
284## Copy-Paste Visual Prompt
285[English prompt for avatar/source image generation]
286
287## Upload Next
288- Face/selfie video:
289- Logo:
290- Brand colors:
291- Voice sample:
292- Script:
293```
294
295For a static personal avatar only, route to `30-design-master-global` personal-brand mode. For talking-head video, continue this workflow.
296
297---
298
299## Anti-detection for FB / IG / TikTok / YouTube
300
301### 5 detection signals and fixes
302
303| Signal | Platforms flagging | Fix |
304|--------|-------------------|-----|
305| Stiff face, no natural blinking | FB, IG | Use selfie video over photo; pick avatars with micro-expressions |
306| Monotone voice, no natural pauses | TikTok, FB | Use voice clone (natural pacing) over default TTS |
307| Fully static background | FB, IG | Add slight noise/grain, or use real-world background |
308| Isolated motion (only mouth moves) | TikTok | Pick avatars with gesture (hands, head); use HeyGen v3+ |
309| Metadata flagged as AI tool | YouTube (monetize) | Re-export through CapCut (strips metadata); add color grade |
310
311### Techniques to add "human feel"
312
3131. **Add film grain / noise:** 2-5% in CapCut or Premiere
3142. **Zoom and crop:** 5-10% crop with subtle motion (Ken Burns)
3153. **Color grade:** Apply film LUT or manually grade — avoid "too clean"
3164. **Text overlay:** Add subtitles, callouts, stickers to cover AI weak spots
3175. **B-roll insert:** Drop 2-3 b-roll clips (product, lifestyle) every 15-20s
3186. **Sound design:** Background music + light SFX (immersion + masks AI voice)
319
320### Per platform
321
322- **TikTok:** Most lenient — content quality wins over AI checks
323- **Facebook / Instagram:** Moderate scrutiny — anti-detection matters
324- **LinkedIn:** Practically no detection — best fit for AI avatars
325- **YouTube:** Strict for monetized videos — must disclose per YPP policy
326
327> **CRITICAL:** NEVER use AI avatars to impersonate real people without consent. This is illegal in most jurisdictions and grounds for permanent platform bans.
328
329---
330
331## Ethics and Disclosure — Region selector
332
333Disclosure laws differ dramatically by region. Pick the matching variant:
334
335| Region | Variant file | Key law |
336|--------|-------------|---------|
337| US / Canada | `variants/01-us.md` | FTC Endorsement Guides (16 CFR Part 255), 2023 update |
338| EU / EEA / UK | `variants/02-eu.md` | **EU AI Act Article 50** (always disclose) + UCPD + GDPR |
339| Southeast Asia | `variants/03-sea.md` | Per-country: ASAS (SG), AKARI (ID), DTI (PH), MCMC (MY), TH |
340| Latin America | `variants/04-latam.md` | CONAR + LGPD (BR), PROFECO (MX), AAIP (AR), per-country |
341
342> ALWAYS read the matching variant BEFORE publishing AI avatar content in that region. Penalties range from warning to multi-thousand-USD fines per influencer (US) and can stack under EU AI Act + GDPR.
343
344### Universal disclosure rule of thumb
345
346When in doubt, disclose. Disclosure is rarely penalized; non-disclosure can be.
347
348```
349"This video uses AI Avatar technology for visuals and voice."
350```
351
352Placement: video description, first 3 seconds on-screen text, OR platform "AI-generated" tag (where available — Meta, TikTok, YouTube all now support this).
353
354---
355
356## QA Score — 100 points
357
358### Scorecard
359
360| # | Criterion | Points | Description |
361|---|-----------|--------|-------------|
362| 1 | Lipsync | /10 | Mouth tracks speech within 0.2s |
363| 2 | Voice match | /10 | Voice sounds like the speaker (if clone) or natural (if TTS) |
364| 3 | Visual quality | /10 | Sharp image, no artifacts, no blur |
365| 4 | Background | /10 | Background suits context, no render glitches |
366| 5 | Lighting | /10 | Even light, no harsh shadows, matches background |
367| 6 | Gesture | /10 | Natural, no jitters, hand/head movement present |
368| 7 | Script flow | /10 | Hook -> Problem -> Solution -> CTA |
369| 8 | Disclosure | /10 | AI disclosure compliant with region (see variant) |
370| 9 | Platform fit | /10 | Correct aspect ratio, duration, format for platform |
371| 10 | CTA | /10 | Clear call-to-action, easy to execute |
372
373### Action thresholds
374
375| Tier | Score | Action |
376|------|-------|--------|
377| **Excellent** | 90-100 | Publish now |
378| **Good** | 70-89 | Publish, note improvements for next round |
379| **Needs fix** | 50-69 | Fix items scoring under 7, then re-render |
380| **Redo** | <50 | Rebuild from script + voice + avatar |
381
382---
383
384## Output template
385
386```markdown
387# AI Avatar Video — [Title] | [Region variant] | [Date]
388
3891. Workflow used: [Single / Translate / Batch / Hybrid]
3902. Script: [Content, 150-300 words]
3913. Voice: [Tool] — [Voice ID / clone name] — Consent: [Yes / N/A]
3924. Avatar: [Tool] — [Avatar ID / custom]
3935. QA Score: [X]/100 (10 criteria)
3946. Disclosure (per region variant): [Text + placement]
3957. Publish: [Platform] — [Aspect ratio] — [Link]
396```
397
398---
399
400## Quality checklist
401
402- [ ] Information collection completed (4 questions)
403- [ ] Tier picked (Free / Pro / Enterprise) and aligns with budget + volume
404- [ ] Workflow picked (Single / Translate / Batch / Hybrid)
405- [ ] Voice clone consent recorded (if cloning a real person)
406- [ ] Avatar setup checklist completed before recording
407- [ ] Anti-detection techniques applied for the target platform
408- [ ] Region variant read and disclosure compliant
409- [ ] QA Score >= 70 before publishing
410
411---
412
413## Related skills
414
415- `25-voice-clone-podcast-global` — voice clone deep-dive + podcast pipeline
416- `04-script-video-global` — script writing for AI avatar
417- `26-thought-leadership-content-global` — content strategy for personal brand
418- `references/ai-video-disclosure-global` — full legal reference
419- `references/voice-clone-prompts-global` — voice clone training prompts
420
421---
422
423*Global Skill 24 (AI Avatar Production) | Over Powers Agency | v1.1.0*