Speech & Audio Agent Skills

Speech & Audio

107 skills
huggingface
transformers-js
Run state-of-the-art machine learning models directly in JavaScript/TypeScript across browsers and server-side runtimes using Transformers.js.
10.8k · bundle
microsoft
podcast-generation
Generate AI-powered podcast-style audio narratives from text using Azure OpenAI's GPT Realtime Mini model via WebSocket, with full-stack implementation from React frontend to Python FastAPI backend.
2.7k · bundle
microsoft
azure-ai
Integrate with Azure AI services including Search, Speech, OpenAI, and Document Intelligence for search, transcription, OCR, and text-to-speech tasks.
2.7k · bundle
microsoft
azure-ai-voicelive-java
Integrate real-time bidirectional voice conversations with AI assistants using the Azure AI VoiceLive SDK for Java, including WebSocket streaming, turn detection, and voice configuration.
2.7k · bundle
microsoft
azure-ai-voicelive-py
Build real-time voice AI applications using the Azure AI Voice Live SDK for bidirectional WebSocket audio communication.
2.7k · bundle
microsoft
azure-ai-openai-dotnet
Integrate Azure OpenAI and OpenAI services in .NET applications for chat completions, embeddings, image generation, audio transcription, and assistants.
2.7k
microsoft
azure-ai-voicelive-dotnet
Build real-time voice AI applications with bidirectional WebSocket communication using the Azure AI Voice Live SDK for .NET.
2.7k
microsoft
azure-ai-transcription-py
Transcribe audio to text using Azure AI Transcription SDK with real-time and batch support, including timestamps and diarization.
2.7k
microsoft
azure-ai-voicelive-ts
Build real-time voice AI applications with bidirectional WebSocket communication using the Azure AI Voice Live SDK for JavaScript/TypeScript.
2.7k · bundle
microsoft
azure-speech-to-text-rest-py
Transcribe short audio files (up to 60 seconds) using Azure Speech-to-Text REST API with Python, without requiring the Speech SDK.
2.7k · bundle
microsoft
azure-communication-callautomation-java
Build server-side call automation workflows with Azure Communication Services Call Automation Java SDK, including IVR systems, call routing, recording, DTMF recognition, text-to-speech, and AI-powered call flows.
2.7k · bundle
nvidia
nemotron-speech
Routes NVIDIA Nemotron Speech (Riva) NIM tasks for ASR, TTS, and NMT, covering cloud-hosted inference, self-hosted Docker deployment, and custom model builds.
2.2k · bundle
nvidia
digital-health-clinical-asr-eval
Score a clinical ASR manifest against a chosen NIM, produce a five-section KER leaderboard, and route the user via a post-eval decision tree.
2.2k · bundle
affaan-m
fal-ai-media
Generate images, videos, and audio using fal.ai models via MCP tools, with support for text-to-image, text/image-to-video, text-to-speech, and video-to-audio.
226k
antigravity
daily
Build real-time voice and multimodal AI applications using Pipecat and Daily, covering pipeline architecture, AI service integration, and transport options.
42.4k
antigravity
videodb
Ingest, index, search, and edit video and audio from files, URLs, live streams, or desktop sessions with timestamps, subtitles, overlays, and real-time alerts.
42.4k · bundle
antigravity
fal-audio
Convert text to speech and speech to text using fal.ai audio models.
42.4k
openai
speech
Generate spoken audio for narration, voiceovers, IVR prompts, and accessibility reads using the OpenAI Audio API with bundled CLI and built-in voices.
23.3k · bundle
openai
transcribe
Transcribe audio files to text with optional speaker diarization and known-speaker hints using OpenAI models.
23.3k · bundle
google-gemini
gemini-live-api-dev
Build real-time, bidirectional streaming applications with the Gemini Live API, covering WebSocket-based audio/video/text streaming, voice activity detection, function calling, session management, and ephemeral tokens.
3.8k
elevenlabs
music
Generate music from text prompts using ElevenLabs Music API, supporting instrumental tracks, songs with lyrics, composition plans, video-to-music, and inpainting.
363 · bundle
elevenlabs
sound-effects
Generate sound effects from text descriptions using ElevenLabs, with support for looping, duration control, and prompt influence tuning.
363 · bundle
elevenlabs
voice-changer
Transform the voice in an audio recording into a different target voice while preserving emotion, timing, and delivery using the ElevenLabs Voice Changer API.
363 · bundle
elevenlabs
speech-to-text
Transcribe audio to text using ElevenLabs Scribe v2, supporting 90+ languages, speaker diarization, and word-level timestamps.
363 · bundle
elevenlabs
text-to-speech
Generate natural speech from text using ElevenLabs voice AI, supporting 70+ languages, multiple models, and various output formats.
363 · bundle
minimax-ai
minimax-music-playlist
Analyzes music listening data from Apple Music or Spotify exports to build a taste profile, then generates personalized playlists with AI-generated songs and album covers.
12.9k · bundle
minimax-ai
mmx-cli
Generate text, images, video, speech, and music via the MiniMax AI platform from the terminal.
12.9k
composiohq
aivoov-automation
Automate Aivoov text-to-speech and voiceover operations through Composio's Aivoov toolkit via Rube MCP.
66.9k
composiohq
gladia-automation
Automate Gladia speech-to-text and audio processing tasks through Composio's Gladia toolkit via Rube MCP.
66.9k
composiohq
rev-ai-automation
Automate Rev AI speech-to-text operations through Composio's Rev AI toolkit via Rube MCP.
66.9k
composiohq
deepgram-automation
Automate Deepgram speech-to-text and audio processing tasks through Composio's Deepgram toolkit via Rube MCP.
66.9k
composiohq
groqcloud-automation
Automate AI inference, chat completions, audio translation, and TTS voice management through GroqCloud's high-performance API via Composio.
66.9k
composiohq
elevenlabs-automation
Automate ElevenLabs text-to-speech workflows: generate speech from text, browse and inspect voices, check subscription limits, list models, stream audio, and retrieve history via the Composio MCP integration.
66.9k
composiohq
castingwords-automation
Automate Castingwords transcription and captioning tasks through Composio's toolkit via Rube MCP.
66.9k
browser-act
youtube-transcript-extractor-api-skill
Extracts YouTube video transcripts and metadata (title, publisher, likes) via the BrowserAct API without CAPTCHA or IP restrictions.
3.7k · bundle
inference-sh
infsh-cli
Run 250+ AI apps from the command line: generate images and videos, call LLMs, search the web, create 3D models, and automate Twitter posts.
584 · bundle