Speech & Audio
-
huggingface Bundle Transformers JSRun state-of-the-art machine learning models directly in JavaScript/TypeScript across browsers and server-side runtimes using Transformers.js.
Audited 10.8k Jul 6 -
microsoft Bundle Podcast GenerationGenerate AI-powered podcast-style audio narratives from text using Azure OpenAI's GPT Realtime Mini model via WebSocket, with full-stack implementation from React frontend to Python FastAPI backend.
Audit pending 2.7k Jul 6 -
microsoft Bundle Azure AIIntegrate with Azure AI services including Search, Speech, OpenAI, and Document Intelligence for search, transcription, OCR, and text-to-speech tasks.
Audited 2.7k Jul 6 -
microsoft Bundle Azure AI Voicelive JavaIntegrate real-time bidirectional voice conversations with AI assistants using the Azure AI VoiceLive SDK for Java, including WebSocket streaming, turn detection, and voice configuration.
Audited 2.7k Jul 6 -
microsoft Bundle Azure AI Voicelive PyBuild real-time voice AI applications using the Azure AI Voice Live SDK for bidirectional WebSocket audio communication.
Audited 2.7k Jul 6 -
microsoft Skill Azure AI Openai DotnetIntegrate Azure OpenAI and OpenAI services in .NET applications for chat completions, embeddings, image generation, audio transcription, and assistants.
Audit pending 2.7k Jul 6 -
microsoft Skill Azure AI Voicelive DotnetBuild real-time voice AI applications with bidirectional WebSocket communication using the Azure AI Voice Live SDK for .NET.
Audited 2.7k Jul 6 -
microsoft Skill Azure AI Transcription PyTranscribe audio to text using Azure AI Transcription SDK with real-time and batch support, including timestamps and diarization.
Audit pending 2.7k Jul 6 -
microsoft Bundle Azure AI Voicelive TSBuild real-time voice AI applications with bidirectional WebSocket communication using the Azure AI Voice Live SDK for JavaScript/TypeScript.
Audited 2.7k Jul 6 -
microsoft Bundle Azure Speech To Text REST PyTranscribe short audio files (up to 60 seconds) using Azure Speech-to-Text REST API with Python, without requiring the Speech SDK.
Audited 2.7k Jul 6 -
microsoft Bundle Azure Communication Callautomation JavaBuild server-side call automation workflows with Azure Communication Services Call Automation Java SDK, including IVR systems, call routing, recording, DTMF recognition, text-to-speech, and AI-powered call flows.
Audited 2.7k Jul 6 -
nvidia Bundle Nemotron SpeechRoutes NVIDIA Nemotron Speech (Riva) NIM tasks for ASR, TTS, and NMT, covering cloud-hosted inference, self-hosted Docker deployment, and custom model builds.
Audited 2.2k Jul 6 -
nvidia Bundle Digital Health Clinical Asr EvalScore a clinical ASR manifest against a chosen NIM, produce a five-section KER leaderboard, and route the user via a post-eval decision tree.
Audit pending 2.2k Jul 6 -
affaan-m Skill Fal AI MediaGenerate images, videos, and audio using fal.ai models via MCP tools, with support for text-to-image, text/image-to-video, text-to-speech, and video-to-audio.
Audit pending 226k Jul 6 -
antigravity Skill DailyBuild real-time voice and multimodal AI applications using Pipecat and Daily, covering pipeline architecture, AI service integration, and transport options.
Audited 42.4k Jul 6 -
antigravity Bundle VideodbIngest, index, search, and edit video and audio from files, URLs, live streams, or desktop sessions with timestamps, subtitles, overlays, and real-time alerts.
Audit pending 42.4k Jul 6 -
antigravity Skill Fal AudioConvert text to speech and speech to text using fal.ai audio models.
Audited 42.4k Jul 6 -
openai Bundle SpeechGenerate spoken audio for narration, voiceovers, IVR prompts, and accessibility reads using the OpenAI Audio API with bundled CLI and built-in voices.
Audit pending 23.3k Jul 6 -
openai Bundle TranscribeTranscribe audio files to text with optional speaker diarization and known-speaker hints using OpenAI models.
Audit pending 23.3k Jul 6 -
google-gemini Skill Gemini Live API DevBuild real-time, bidirectional streaming applications with the Gemini Live API, covering WebSocket-based audio/video/text streaming, voice activity detection, function calling, session management, and ephemeral tokens.
Audit pending 3.8k Jul 6 -
elevenlabs Bundle MusicGenerate music from text prompts using ElevenLabs Music API, supporting instrumental tracks, songs with lyrics, composition plans, video-to-music, and inpainting.
Audit pending 363 Jul 6 -
elevenlabs Bundle Sound EffectsGenerate sound effects from text descriptions using ElevenLabs, with support for looping, duration control, and prompt influence tuning.
Audit pending 363 Jul 6 -
elevenlabs Bundle Voice ChangerTransform the voice in an audio recording into a different target voice while preserving emotion, timing, and delivery using the ElevenLabs Voice Changer API.
Audit pending 363 Jul 6 -
elevenlabs Bundle Speech To TextTranscribe audio to text using ElevenLabs Scribe v2, supporting 90+ languages, speaker diarization, and word-level timestamps.
Audit pending 363 Jul 6 -
elevenlabs Bundle Text To SpeechGenerate natural speech from text using ElevenLabs voice AI, supporting 70+ languages, multiple models, and various output formats.
Audit pending 363 Jul 6 -
minimax-ai Bundle Minimax Music PlaylistAnalyzes music listening data from Apple Music or Spotify exports to build a taste profile, then generates personalized playlists with AI-generated songs and album covers.
Audit pending 12.9k Jul 6 -
minimax-ai Skill Mmx CLIGenerate text, images, video, speech, and music via the MiniMax AI platform from the terminal.
Audit pending 12.9k Jul 6 -
composiohq Skill Aivoov AutomationAutomate Aivoov text-to-speech and voiceover operations through Composio's Aivoov toolkit via Rube MCP.
Audit pending 66.9k Jul 6 -
composiohq Skill Gladia AutomationAutomate Gladia speech-to-text and audio processing tasks through Composio's Gladia toolkit via Rube MCP.
Audit pending 66.9k Jul 6 -
composiohq Skill Rev AI AutomationAutomate Rev AI speech-to-text operations through Composio's Rev AI toolkit via Rube MCP.
Audit pending 66.9k Jul 6 -
composiohq Skill Deepgram AutomationAutomate Deepgram speech-to-text and audio processing tasks through Composio's Deepgram toolkit via Rube MCP.
Audit pending 66.9k Jul 6 -
composiohq Skill Groqcloud AutomationAutomate AI inference, chat completions, audio translation, and TTS voice management through GroqCloud's high-performance API via Composio.
Audit pending 66.9k Jul 6 -
composiohq Skill Elevenlabs AutomationAutomate ElevenLabs text-to-speech workflows: generate speech from text, browse and inspect voices, check subscription limits, list models, stream audio, and retrieve history via the Composio MCP integration.
Audit pending 66.9k Jul 6 -
composiohq Skill Castingwords AutomationAutomate Castingwords transcription and captioning tasks through Composio's toolkit via Rube MCP.
Audit pending 66.9k Jul 6 -
browser-act Bundle Youtube Transcript Extractor API SkillExtracts YouTube video transcripts and metadata (title, publisher, likes) via the BrowserAct API without CAPTCHA or IP restrictions.
Audit pending 3.7k Jul 6 -
inference-sh Bundle Infsh CLIRun 250+ AI apps from the command line: generate images and videos, call LLMs, search the web, create 3D models, and automate Twitter posts.
Audit pending 584 Jul 6