Packs
2 packsResults for “voice-ai”
43 skillsvoice-maestro
Use when voice AI strategy, conversational AI architecture, voice technology innovation, or voice platform leadership is needed. This agent specializes in voice AI leadership within the VoiceForge AI ecosystem.
0
azure-ai-voicelive-py
Build real-time voice AI applications using the Azure AI Voice Live SDK for bidirectional WebSocket audio communication.
2.7k · bundle
azure-ai-voicelive-ts
Build real-time voice AI applications with bidirectional WebSocket communication using the Azure AI Voice Live SDK for JavaScript/TypeScript.
2.7k · bundle
api-architect
Use when voice AI API design, integration architecture, API management, or voice service APIs is needed. This agent specializes in voice AI API architecture within the VoiceForge AI ecosystem.
0
azure-ai-voicelive-dotnet
Build real-time voice AI applications with bidirectional WebSocket communication using the Azure AI Voice Live SDK for .NET.
2.7k
azure-ai-voicelive-java
Integrate real-time bidirectional voice conversations with AI assistants using the Azure AI VoiceLive SDK for Java, including WebSocket streaming, turn detection, and voice configuration.
2.7k · bundle
More results
ai-voice-cloning
Generate natural AI voices, text-to-speech, and voice synthesis using the inference.sh CLI with models like Inworld TTS, ElevenLabs, and Kokoro TTS for voiceovers, audiobooks, podcasts, and more.
584
aivoov-automation
Automate Aivoov text-to-speech and voiceover operations through Composio's Aivoov toolkit via Rube MCP.
66.9k
voice-ai
Generates speech, transcribes audio, clones voices, and builds real-time voice agents using ElevenLabs, OpenAI TTS, Whisper, and Vapi.
10
agents
Build voice AI agents with natural conversations, multiple LLM providers, custom tools, and easy web embedding.
363 · bundle
daily
Build real-time voice and multimodal AI applications using Pipecat and Daily, covering pipeline architecture, AI service integration, and transport options.
42.4k
daily
Reference for building real-time voice and multimodal AI applications with Daily and Pipecat, covering pipeline architecture, AI service integrations, transports, and client SDKs.
3
voice-designer
Builds detailed voice profiles for characters, narrators, brands, or AI agents, covering diction, syntax, rhythm, worldview, signature patterns, and failure modes, and can reverse-engineer profiles from writing samples or audit existing ones.
1
humanizer
Humanize text: strip AI-isms and add real voice.
0 · bundle
daily
Reference for building real-time voice and multimodal AI agents with Pipecat, covering pipelines, speech services, LLMs, transports, and deployment.
2
daily
Reference for building real-time voice and multimodal AI applications with Pipecat, covering pipelines, speech services, LLMs, transports, and deployment.
5
daily
Provides a reference for building real-time voice and multimodal AI agents with Pipecat, covering pipeline architecture, speech services, LLM integration, transports, and deployment.
0 · bundle
speech
Generate spoken audio for narration, voiceovers, IVR prompts, and accessibility reads using the OpenAI Audio API with bundled CLI and built-in voices.
23.3k · bundle
text-to-speech
Generate natural speech from text using ElevenLabs voice AI, supporting 70+ languages, multiple models, and various output formats.
363 · bundle
talking-head-production
Create talking head videos with AI avatars, lipsync, and voiceover using the inference.sh CLI.
584
fal-audio
Converts text to speech and speech to text using fal.ai audio models.
5
fal-audio
Converts text to speech and speech to text using fal.ai audio models.
2
podcast-generation
Generate AI-powered podcast-style audio narratives from text using Azure OpenAI's GPT Realtime Mini model via WebSocket, with full-stack implementation from React frontend to Python FastAPI backend.
2.7k · bundle
detecting-deepfake-audio-in-vishing-attacks
Detects AI-generated deepfake audio used in voice phishing (vishing) attacks by extracting spectral features and classifying samples with machine learning models.
24.6k · bundle
ai-podcast
Generate multi-person talking head podcast videos from scratch using AI — character creation, TTS, avatar animation, and video stitching.
584
whisper
Transcribe and translate speech across 99 languages using OpenAI's Whisper model, with support for multiple model sizes, batch processing, and subtitle generation.
10.4k · bundle
daily
Reference for building real-time voice and multimodal AI applications with Pipecat, covering pipelines, speech services, LLM integration, and transports.
253
fal-audio
Convert text to speech and speech to text using fal.ai audio models.
42.4k
ai-avatar-video
Create AI avatar, talking-head, and lip-sync videos on RunComfy via the `runcomfy` CLI. Routes across ByteDance OmniHuman (audio-driven full-body avatar), Wan-AI Wan 2-7 (audio-driven mouth sync via `audio_url` on a portrait), HappyHorse 1.0 (Arena #1 t2v / i2v with in-pass audio), and Seedance v2 Pro (multi-modal cinematic with reference audio + reference subject). Picks the right model for the user's actual intent — UGC voiceover, virtual presenter, dubbed product demo, lip-synced character, dialog scene — and ships each model's documented prompting patterns plus the minimal `runcomfy run` invoke. Triggers on "talking head", "lip sync", "avatar video", "make X speak", "audio to video", "audio driven avatar", "virtual presenter", "AI spokesperson", "dubbed video", "UGC avatar", "HeyGen alternative", "Synthesia alternative", "digital human", "make this portrait talk", "video from voiceover", or any explicit ask to put words in a face.
12
ai-avatar-video
Create AI avatar, talking-head, and lip-sync videos on RunComfy via the `runcomfy` CLI. Routes across ByteDance OmniHuman (audio-driven full-body avatar), Wan-AI Wan 2-7 (audio-driven mouth sync via `audio_url` on a portrait), HappyHorse 1.0 (Arena #1 t2v / i2v with in-pass audio), and Seedance v2 Pro (multi-modal cinematic with reference audio + reference subject). Picks the right model for the user's actual intent — UGC voiceover, virtual presenter, dubbed product demo, lip-synced character, dialog scene — and ships each model's documented prompting patterns plus the minimal `runcomfy run` invoke. Triggers on "talking head", "lip sync", "avatar video", "make X speak", "audio to video", "audio driven avatar", "virtual presenter", "AI spokesperson", "dubbed video", "UGC avatar", "HeyGen alternative", "Synthesia alternative", "digital human", "make this portrait talk", "video from voiceover", or any explicit ask to put words in a face.
5
speech-engine
Add real-time voice conversations to a custom agent runtime using ElevenLabs Speech Engine, handling WebSocket servers, browser clients, and interruption-aware streaming.
363 · bundle
ruvocal
Voice/audio-enabled AI chat interface forked from HuggingFace Chat UI with RVF document store replacing MongoDB
0
ai-brand-kit
Build comprehensive AI-native brand asset systems that maintain consistency across all AI-generated content. Train AI tools on brand guidelines, create reusable prompt libraries, and manage visual/voice assets at scale. Use when ", " mentioned.
128 · bundle
fei-fei-li-expert
Embody the voice and methodology of AI researcher Fei-Fei Li to enhance technical content with data-centric and human-centered perspectives.
6
groqcloud-automation
Automate AI inference, chat completions, audio translation, and TTS voice management through GroqCloud's high-performance API via Composio.
66.9k
agents-sdk
Build AI agents on Cloudflare Workers using the Agents SDK, covering stateful agents, durable workflows, real-time WebSocket apps, scheduled tasks, MCP servers, chat applications, voice agents, and browser automation.
2.1k · bundle