Packs
1 packResults for “speech”
83 skillsfal-audio
Convert text to speech and speech to text using fal.ai audio models.
42.4k
fal-audio
Converts text to speech and speech to text using fal.ai audio models.
2
fal-audio
Converts text to speech and speech to text using fal.ai audio models.
5
speech-engine
Add real-time voice conversations to a custom agent runtime using ElevenLabs Speech Engine, handling WebSocket servers, browser clients, and interruption-aware streaming.
363 · bundle
azure-speech-to-text-rest-py
Transcribe short audio files (up to 60 seconds) using Azure Speech-to-Text REST API with Python, without requiring the Speech SDK.
2.7k · bundle
vox
Runs a local voice MCP server in Rust for text-to-speech and speech-to-text, with build, test, and configuration guidance.
54 · bundle
More results
sag
Generates speech from text using ElevenLabs text-to-speech with a command-line interface and local playback.
1 · bundle
text-to-speech
Generate natural speech from text using ElevenLabs voice AI, supporting 70+ languages, multiple models, and various output formats.
363 · bundle
bss-eval
Evaluates speech language models on beyond-semantic speech attributes such as dialect comprehension, multi-turn context memory, emotion perception, age-aware response generation, and non-verbal cue handling, reporting accuracy and judge-based scores.
3
daily
Reference for building real-time voice and multimodal AI applications with Pipecat, covering pipelines, speech services, LLM integration, and transports.
253
daily
Reference for building real-time voice and multimodal AI agents with Pipecat, covering pipelines, speech services, LLMs, transports, and deployment.
2
daily
Reference for building real-time voice and multimodal AI applications with Pipecat, covering pipelines, speech services, LLMs, transports, and deployment.
5
nemotron-speech
Routes NVIDIA Nemotron Speech (Riva) NIM tasks for ASR, TTS, and NMT, covering cloud-hosted inference, self-hosted Docker deployment, and custom model builds.
2.2k · bundle
azure-ai
Integrate with Azure AI services including Search, Speech, OpenAI, and Document Intelligence for search, transcription, OCR, and text-to-speech tasks.
2.7k · bundle
fal
Generate images, videos, audio, and more using fal.ai AI models, with support for text-to-image, image-to-video, text-to-speech, speech-to-text, image editing, upscaling, model search, workflow creation, and cost estimation.
567 · bundle
voice-ai
Generates speech, transcribes audio, clones voices, and builds real-time voice agents using ElevenLabs, OpenAI TTS, Whisper, and Vapi.
10
ttsds
Evaluates text-to-speech systems by measuring distributional distance between synthetic and real speech across five factors, producing a scalar score without subjective MOS ratings.
3
whisper
Transcribe and translate speech across 99 languages using OpenAI's Whisper model, with support for multiple model sizes, batch processing, and subtitle generation.
10.4k · bundle
daily
Provides a reference for building real-time voice and multimodal AI agents with Pipecat, covering pipeline architecture, speech services, LLM integration, transports, and deployment.
0 · bundle
speech
Generate spoken audio for narration, voiceovers, IVR prompts, and accessibility reads using the OpenAI Audio API with bundled CLI and built-in voices.
23.3k · bundle
tts
Converts text to speech and generates MP3 audio files using the Hume AI or OpenAI API, printing the file path for delivery.
1 · bundle
speech
Generates spoken audio clips from text for narration, voiceovers, IVR prompts, and accessibility reads, with support for single clips and batch processing.
61
ast-eval
Benchmarks automatic speech translation and recognition on English-French and English-Romanian datasets, reporting BLEU and WER on tokenized outputs.
3
elevenlabs-voice-changer
Transform any voice into a different voice while preserving speech content and emotion using the inference.sh CLI and ElevenLabs models.
584
sag
Generates speech from text using ElevenLabs TTS with local playback, supporting voice selection, pronunciation rules, and audio tags.
61
elevenlabs-automation
Automate ElevenLabs text-to-speech workflows: generate speech from text, browse and inspect voices, check subscription limits, list models, stream audio, and retrieve history via the Composio MCP integration.
66.9k
speech-to-text
Transcribe audio to text using ElevenLabs Scribe v2, supporting 90+ languages, speaker diarization, and word-level timestamps.
363 · bundle
speech-to-text
Transcribe audio to text using ElevenLabs Scribe and Whisper models via the inference.sh CLI, supporting timestamps, speaker diarization, translation, and multi-language transcription.
584
text-to-speech
Convert text to natural speech using multiple TTS models via the inference.sh CLI, with support for emotion steering, voice cloning, and multi-speaker dialogue.
584
aivoov-automation
Automate Aivoov text-to-speech and voiceover operations through Composio's Aivoov toolkit via Rube MCP.
66.9k
gladia-automation
Automate Gladia speech-to-text and audio processing tasks through Composio's Gladia toolkit via Rube MCP.
66.9k
deepgram-automation
Automate Deepgram speech-to-text and audio processing tasks through Composio's Deepgram toolkit via Rube MCP.
66.9k
rev-ai-automation
Automate Rev AI speech-to-text operations through Composio's Rev AI toolkit via Rube MCP.
66.9k
mos
Evaluates the naturalness, speaker similarity, and real-time synthesis speed of a Mandarin speech cloning system across diverse practical application scenarios.
3
mmx-cli
Generate text, images, video, speech, and music via the MiniMax AI platform using the mmx CLI.
42.4k
mmx-cli
Generate text, images, video, speech, and music via the MiniMax AI platform from the terminal.
12.9k