Results for “speaker-similarity”

27 skills
github
Resemble Detect
Detect AI-generated audio, images, video, and text, trace synthesis sources, apply watermarks, verify speaker identity, and analyze media intelligence using the Resemble AI platform.
36.2k · bundle
solizardking
Sag
ElevenLabs text-to-speech with mac-style say UX.
0
azusagasaku
Brand Voice
从真实的帖子、文章、发布说明、文档或网站文案中构建基于源材料的写作风格档案,然后在内容、外展和社交工作流中重复使用该档案。当用户希望保持声音一致性而不使用通用的AI写作套路时使用。
0 · bundle
promisingcoder
Sag
ElevenLabs text-to-speech with mac-style say UX.
0
om-scogo
Sag
ElevenLabs text-to-speech with mac-style say UX.
0 · bundle
infometa
Sag
ElevenLabs text-to-speech with mac-style say UX.
228
doriangallo
Fal Audio
Text-to-speech and speech-to-text using fal.ai audio models
1
rootcastleco
Fal Audio
Text-to-speech and speech-to-text using fal.ai audio models
6
tangchunwu
Brand Voice
Build a source-derived writing style profile from real posts, essays, launch notes, docs, or site copy, then reuse that profile across content, outreach, and social workflows. Use when the user wants voice consistency without generic AI writing tropes.
1 · bundle
prime-skills
Lipsync
Lip-sync a face to a specific audio track on RunComfy via the `runcomfy` CLI. Routes across ByteDance OmniHuman (audio-driven full-body avatar from a portrait + audio), Sync Labs sync v2 / Pro (state-of-the-art mouth sync onto a video), Kling lipsync (audio-to- video and text-to-video with synced speech), and Creatify lipsync. The skill picks the right endpoint for the user's actual intent — portrait still + audio (avatar-style), source video + audio (mouth- swap on existing footage), or generate-and-sync from a script. Triggers on "lip sync", "lipsync", "make this video speak", "match audio to mouth", "dub video", "sync lips to voice", "Sync Labs", "voiceover sync", or any explicit ask to drive a face's mouth from an audio track.
33
welitonevoc
Fal Audio
Text-to-speech and speech-to-text using fal.ai audio models
1
diegojcn
Fal Audio
Text-to-speech and speech-to-text using fal.ai audio models
1
yanacuti1121
Brand Voice
Build a source-derived writing style profile from real posts, essays, launch notes, docs, or site copy, then reuse that profile across content, outreach, and social workflows. Use when the user wants voice consistency without generic AI writing tropes.
2
livelybug
Brand Voice
Build a source-derived writing style profile from real posts, essays, launch notes, docs, or site copy, then reuse that profile across content, outreach, and social workflows. Use when the user wants voice consistency without generic AI writing tropes.
0 · bundle
desesbraker
Fal Audio
Text-to-speech and speech-to-text using fal.ai audio models
2
inference-sh
Dialogue Audio
Create realistic multi-speaker dialogue audio using Dia TTS via the inference.sh CLI, with control over speaker tags, emotion, pacing, and conversation structure.
584
iamanacarolinarezende
Fal Audio
Text-to-speech and speech-to-text using fal.ai audio models
0
sinhoneyy
Behuman
Use when the user wants more human-like AI responses — less robotic, less listy, more authentic. Triggers: 'behuman', 'be real', 'like a human', 'more human', 'less AI', 'talk like a person', 'mirror mode', 'stop being so AI', or when conversations are emotionally charged (grief, job loss, relationship advice, fear). NOT for technical questions, code generation, or factual lookups.
11 · bundle
ahang1598
Seed Audio
用自然语言描述生成目标音频。把一段场景描述(人声对话、环境声、音效、背景音乐等复合音频)一次性生成成音频。当用户描述一个声音场景、要求生成/合成/制作一段音频或声音、给出形如"角色:台词"的对话脚本要转成音频、或要按参考音频的音色说话时使用。支持两种模式:纯文本描述生成(T2A)和带参考音频生成(A2A,在描述中引用参考音频指定角色音色)
9 · bundle
jackychenlu
Sag
ElevenLabs text-to-speech with mac-style say UX.
0
x402agent
Sag
ElevenLabs text-to-speech with mac-style say UX.
9
lucian55
Pujian Skill
朴树(音乐人)认知与表达框架(压缩蒸馏):内向真诚、少话多留白、生命与创作一体… 触发:平凡之路 等。不消费抑郁;非医疗建议
9 · bundle
runcomfy-com
Lipsync
Lip-sync a face to a specific audio track on RunComfy via the `runcomfy` CLI. Routes across ByteDance OmniHuman (audio-driven full-body avatar from a portrait + audio), Sync Labs sync v2 / Pro (state-of-the-art mouth sync onto a video), Kling lipsync (audio-to- video and text-to-video with synced speech), and Creatify lipsync. The skill picks the right endpoint for the user's actual intent — portrait still + audio (avatar-style), source video + audio (mouth- swap on existing footage), or generate-and-sync from a script. Triggers on "lip sync", "lipsync", "make this video speak", "match audio to mouth", "dub video", "sync lips to voice", "Sync Labs", "voiceover sync", or any explicit ask to drive a face's mouth from an audio track.
12
jrennie99-glitch
Sag
ElevenLabs text-to-speech with mac-style say UX.
0
mit-network
Fal Audio
Text-to-speech and speech-to-text using fal.ai audio models
2
danstrem2
Sag
ElevenLabs text-to-speech with mac-style say UX.
2 · bundle
inskillflow
Fal Audio
Text-to-speech and speech-to-text using fal.ai audio models
1