Packs
1 packResults for “speech-to-speech”
151 skillssag
Generates speech from text using ElevenLabs TTS with local playback, supporting voice selection, pronunciation rules, and audio tags.
61
gladia-automation
Automate Gladia speech-to-text and audio processing tasks through Composio's Gladia toolkit via Rube MCP.
66.9k
deepgram-automation
Automate Deepgram speech-to-text and audio processing tasks through Composio's Deepgram toolkit via Rube MCP.
66.9k
rev-ai-automation
Automate Rev AI speech-to-text operations through Composio's Rev AI toolkit via Rube MCP.
66.9k
speech
Generate spoken audio from text using OpenAI's API with built-in voices for narrated explainers, lecture audio, and quick voiceover tracks.
1
tts
Converts text to speech and generates MP3 audio files using Hume AI or OpenAI APIs, printing the file path for delivery.
32 · bundle
speech-to-text
Transcribe audio to text using ElevenLabs Scribe v2, supporting 90+ languages, speaker diarization, and word-level timestamps.
363 · bundle
tts
Converts text to speech and generates MP3 audio files using Hume AI or OpenAI APIs, printing the file path for delivery.
10 · bundle
asr
Transcribes audio from URLs or local files into text with speaker diarization, word timestamps, and multiple output formats via a command-line tool.
10 · bundle
sag
ElevenLabs text-to-speech with mac-style say UX.
9
sag
ElevenLabs text-to-speech with mac-style say UX.
2 · bundle
sag
ElevenLabs text-to-speech with mac-style say UX.
0
sag
ElevenLabs text-to-speech with mac-style say UX.
0
sag
ElevenLabs text-to-speech with mac-style say UX.
0
sag
ElevenLabs text-to-speech with mac-style say UX.
228
tts
Use this skill whenever the user wants to convert text into speech, generate audio from text, or produce voiceovers. Triggers include: any mention of 'TTS', 'text to speech', 'speak', 'say', 'voice', 'read aloud', 'audio narration', 'voiceover', 'dubbing', or requests to turn written content into spoken audio. Also use when converting EPUB/PDF/SRT/articles to audio, cloning voices from reference audio, controlling emotion or speed in speech, aligning speech to subtitle timelines, or producing per-segment voice-mapped audio.
0 · bundle
imam
Leads the five daily Islamic prayers and Friday khutbahs via text-to-speech, with multilingual support and step-by-step vocal guidance.
32 · bundle
tts
Converts text to speech and generates MP3 audio files using the Hume AI or OpenAI API, printing the file path for delivery.
1 · bundle
mmx-cli
Generate text, images, video, speech, and music via the MiniMax AI platform from the terminal.
12.9k
speech
Generates spoken audio clips from text for narration, voiceovers, IVR prompts, and accessibility reads, with support for single clips and batch processing.
61
geminigen-ai
Unified multimedia generation API for images, videos, and text-to-speech, replacing separate providers for a single workflow.
10
sag
ElevenLabs text-to-speech with mac-style say UX.
0 · bundle
sag
ElevenLabs text-to-speech with mac-style say UX.
0
daily
Build real-time voice and multimodal AI applications using Pipecat and Daily, covering pipeline architecture, AI service integration, and transport options.
42.4k
speech
Generate spoken audio for narration, voiceovers, IVR prompts, and accessibility reads using the OpenAI Audio API with bundled CLI and built-in voices.
23.3k · bundle
asr
Transcribes audio from URLs or local files to text using the Speech is Cheap API, with options for speaker diarization, timestamps, and multiple output formats.
1 · bundle
azure-ai-voicelive-py
Build real-time voice AI applications using the Azure AI Voice Live SDK for bidirectional WebSocket audio communication.
2.7k · bundle
elevenlabs-tts
Generate high-quality speech from text using ElevenLabs' premium voices, with support for 32 languages, multiple models, and voice tuning parameters.
584
asr
Transcribes audio from URLs or local files into text using a low-cost speech-to-text API, with options for speaker diarization, word timestamps, and multiple output formats.
1 · bundle
speech-to-text
Transcribe audio to text using ElevenLabs Scribe and Whisper models via the inference.sh CLI, supporting timestamps, speaker diarization, translation, and multi-language transcription.
584
ai-podcast-creation
Create AI-powered podcasts and audio content using text-to-speech, music generation, and audio editing via the inference.sh CLI.
584
openai-whisper
Local speech-to-text with the Whisper CLI (no API key).
9
sherpa-onnx-tts
Local text-to-speech via sherpa-onnx (offline, no cloud)
9 · bundle
openai-whisper
Local speech-to-text with the Whisper CLI (no API key).
0
openai-whisper
Local speech-to-text with the Whisper CLI (no API key).
0
sherpa-onnx-tts
Local text-to-speech via sherpa-onnx (offline, no cloud)
0 · bundle