Packs

1 pack

Results for “speech-to-speech”

151 skills
aniruddhaadak80
openai-whisper
Local speech-to-text with the Whisper CLI (no API key).
0
elevenlabs
voice-changer
Transform the voice in an audio recording into a different target voice while preserving emotion, timing, and delivery using the ElevenLabs Voice Changer API.
363 · bundle
phoroth
mmx-cli
Generates text, images, video, speech, and music, and performs web searches via the MiniMax AI platform using the mmx terminal CLI.
3
drnabeelkhan
voice-mode
Enables voice-driven invocation of Maxim's capabilities via hotword routing, intent classification, and decision capture, with graceful fallback when the voicemode plugin is absent.
2
joshuashepherd
fal-ai-media
Generates images, videos, and audio using fal.ai models via MCP, covering text-to-image, text/image-to-video, text-to-speech, and video-to-audio.
1
kbarbel640-del
fal
Search, explore, and run fal.ai generative AI models for image, video, audio, and 3D generation, including queue management and file uploads.
1 · bundle
mhassan0000
fal-ai-media
Generates images, videos, and audio using fal.ai models via MCP, covering text-to-image, text/image-to-video, text-to-speech, and video-to-audio.
1
q2805187159
whisper
OpenAI's general-purpose speech recognition model. Supports 99 languages, transcription, translation to English, and language identification. Six model sizes from tiny (39M params) to large (1550M params). Use for speech-to-text, podcast transcription, or multilingual audio processing. Best for robust, multilingual ASR.
3 · bundle
tianhao909
whisper
OpenAI's general-purpose speech recognition model. Supports 99 languages, transcription, translation to English, and language identification. Six model sizes from tiny (39M params) to large (1550M params). Use for speech-to-text, podcast transcription, or multilingual audio processing. Best for robust, multilingual ASR.
1 · bundle
qcmuu
whisper
OpenAI's general-purpose speech recognition model. Supports 99 languages, transcription, translation to English, and language identification. Six model sizes from tiny (39M params) to large (1550M params). Use for speech-to-text, podcast transcription, or multilingual audio processing. Best for robust, multilingual ASR.
0 · bundle
jackychenlu
whisper
OpenAI's general-purpose speech recognition model. Supports 99 languages, transcription, translation to English, and language identification. Six model sizes from tiny (39M params) to large (1550M params). Use for speech-to-text, podcast transcription, or multilingual audio processing. Best for robust, multilingual ASR.
0 · bundle
bog5d
whisper
OpenAI's general-purpose speech recognition model. Supports 99 languages, transcription, translation to English, and language identification. Six model sizes from tiny (39M params) to large (1550M params). Use for speech-to-text, podcast transcription, or multilingual audio processing. Best for robust, multilingual ASR.
0 · bundle
aniruddhaadak80
whisper
OpenAI's general-purpose speech recognition model. Supports 99 languages, transcription, translation to English, and language identification. Six model sizes from tiny (39M params) to large (1550M params). Use for speech-to-text, podcast transcription, or multilingual audio processing. Best for robust, multilingual ASR.
0 · bundle
ichichuang
whisper
OpenAI's general-purpose speech recognition model. Supports 99 languages, transcription, translation to English, and language identification. Six model sizes from tiny (39M params) to large (1550M params). Use for speech-to-text, podcast transcription, or multilingual audio processing. Best for robust, multilingual ASR.
0 · bundle
peteedoo
whisper
OpenAI's general-purpose speech recognition model. Supports 99 languages, transcription, translation to English, and language identification. Six model sizes from tiny (39M params) to large (1550M params). Use for speech-to-text, podcast transcription, or multilingual audio processing. Best for robust, multilingual ASR.
0 · bundle
sakamoto-family-smile
fal-ai-media
Generates images, videos, and audio using fal.ai models via MCP, covering text-to-image, text/image-to-video, text-to-speech, and video-to-audio.
0
whd4
voice-agents
Voice agents represent the frontier of AI interaction - humans speaking naturally with AI systems. The challenge isn't just speech recognition and synthesis, it's achieving natural conversation flow with sub-800ms latency while handling interruptions, background noise, and emotional nuance. This skill covers two architectures: speech-to-speech (OpenAI Realtime API, lowest latency, most natural) and pipeline (STT→LLM→TTS, more control, easier to debug). Key insight: latency is the constraint. Hu
0
danstrem2
voice-agents
Voice agents represent the frontier of AI interaction - humans speaking naturally with AI systems. The challenge isn't just speech recognition and synthesis, it's achieving natural conversation flow with sub-800ms latency while handling interruptions, background noise, and emotional nuance. This skill covers two architectures: speech-to-speech (OpenAI Realtime API, lowest latency, most natural) and pipeline (STT→LLM→TTS, more control, easier to debug). Key insight: latency is the constraint. Hu
2
dokhacgiakhoa
voice-agents
Voice agents represent the frontier of AI interaction - humans speaking naturally with AI systems. The challenge isn't just speech recognition and synthesis, it's achieving natural conversation flow with sub-800ms latency while handling interruptions, background noise, and emotional nuance. This skill covers two architectures: speech-to-speech (OpenAI Realtime API, lowest latency, most natural) and pipeline (STT→LLM→TTS, more control, easier to debug). Key insight: latency is the constraint. Hu
505 · bundle
microsoft
azure-communication-callautomation-java
Build server-side call automation workflows with Azure Communication Services Call Automation Java SDK, including IVR systems, call routing, recording, DTMF recognition, text-to-speech, and AI-powered call flows.
2.7k · bundle
dvcrn
fal
Search, explore, and run fal.ai generative AI models for image, video, audio, and 3D generation, including schema lookup, job submission, status polling, result retrieval, and file uploads.
32 · bundle
affaan-m
fal-ai-media
Generate images, videos, and audio using fal.ai models via MCP tools, with support for text-to-image, text/image-to-video, text-to-speech, and video-to-audio.
226k
inference-sh
ai-voice-cloning
Generate natural AI voices, text-to-speech, and voice synthesis using the inference.sh CLI with models like Inworld TTS, ElevenLabs, and Kokoro TTS for voiceovers, audiobooks, podcasts, and more.
584
om-scogo
asr
Transcribe audio files to text using local speech recognition. Triggers on: "转录", "transcribe", "语音转文字", "ASR", "识别音频", "把这段音频转成文字".
0 · bundle
artubss
speech
Use quando o usuário solicita narração em texto-para-fala, voiceovers de acessibilidade, prompts de áudio ou geração em lote via OpenAI Audio API; execute a CLI incluída (`scripts/text_to_speech.py`) com vozes integradas e requer `OPENAI_API_KEY` para chamadas diretas. Criação de vozes customizadas está fora do escopo.
10 · bundle
om-scogo
zai-tts
Text-to-speech conversion using GLM-TTS service via the `uvx zai-tts` command for generating audio from text. Use when (1) User requests audio/voice output with the "tts" trigger or keyword. (2) Content needs to be spoken rather than read (multitasking, accessibility, podcast, driving, cooking). (3) Using pre-cloned voices for speech.
0 · bundle
openai
transcribe
Transcribe audio files to text with optional speaker diarization and known-speaker hints using OpenAI models.
23.3k · bundle
inference-sh
elevenlabs-dialogue
Generate multi-speaker dialogue audio with different voices in a single file using the inference.sh CLI.
584
dvcrn
stt
Transcribes audio files to text using OpenAI Whisper, optimized for Brazilian Portuguese, with support for common audio formats and timestamped output.
32 · bundle
composiohq
groqcloud-automation
Automate AI inference, chat completions, audio translation, and TTS voice management through GroqCloud's high-performance API via Composio.
66.9k
alirezarezvani
demo-video
Create polished demo videos, product walkthroughs, and feature showcases by orchestrating browser rendering, text-to-speech, and video compositing.
20.4k · bundle
oyi77
ai-podcast
Creates fully automated AI podcasts that research, write, and narrate complete episodes, with guidance on monetization and building a podcast network.
10
seaworld008
transcribe
Transcribe audio files to text with optional diarization and known-speaker hints. Use when a user asks to transcribe speech from audio/video, extract text from recordings, or label speakers in interviews or meetings.
65 · bundle
microsoft
azure-ai-transcription-py
Transcribe audio to text using Azure AI Transcription SDK with real-time and batch support, including timestamps and diarization.
2.7k
metinduraktr-44
transcribe
Transcribe audio files to text with optional diarization and known-speaker hints. Use when a user asks to transcribe speech from audio/video, extract text from recordings, or label speakers in interviews or meetings.
0 · bundle
ranbot-ai
mmx-cli
Use mmx to generate text, images, video, speech, and music via the MiniMax AI platform. Use when the user wants to create media content, chat with MiniMax models, perform web search, or manage MiniMax
6