Speech & Audio Agent Skills
Speech & Audio
107 skillsai-podcast
Creates fully automated AI podcasts that research, write, and narrate complete episodes, with guidance on monetization and building a podcast network.
10
geminigen-ai
Unified multimedia generation API for images, videos, and text-to-speech, replacing separate providers for a single workflow.
10
172-rvc-7a57af2e
Guides downloading and configuring RVC voice conversion models, including HuBERT and index files, and running voice conversion scripts.
7 · bundle
tts
Converts text to speech and generates MP3 audio files using the Hume AI or OpenAI API, printing the file path for delivery.
1 · bundle
asr
Transcribes audio from URLs or local files to text using the Speech is Cheap API, with options for speaker diarization, timestamps, and multiple output formats.
1 · bundle
videodb
Ingests video and audio from files, URLs, and live streams, builds searchable visual and spoken indexes, edits timelines with subtitles and overlays, and generates real-time alerts.
3 · bundle
daily
Reference for building real-time voice and multimodal AI applications with Pipecat, covering pipelines, speech services, LLMs, transports, and deployment.
5
mos
Evaluates the naturalness, speaker similarity, and real-time synthesis speed of a Mandarin speech cloning system across diverse practical application scenarios.
3
atxp
Access ATXP's paid API tools for web search, AI image generation, music creation, video generation, and X/Twitter search via CLI or programmatic client.
2 · bundle
digital-health-clinical-asr-setup
Bootstraps a clinical ASR evaluation environment by verifying NVIDIA_API_KEY, installing Python dependencies, and running a smoke test against hosted TTS/ASR services.
2.2k · bundle
mmx-cli
Generate text, images, video, speech, and music via the MiniMax AI platform using the mmx CLI.
42.4k
runapi-cli
Generate AI images, videos, and music/audio from agents using the RunAPI CLI.
42.4k
speech-engine
Add real-time voice conversations to a custom agent runtime using ElevenLabs Speech Engine, handling WebSocket servers, browser clients, and interruption-aware streaming.
363 · bundle
minimax-music-gen
Generate songs, instrumental tracks, and covers using the MiniMax Music API with basic or advanced control modes.
12.9k · bundle
text-to-speech
Convert text to natural speech using multiple TTS models via the inference.sh CLI, with support for emotion steering, voice cloning, and multi-speaker dialogue.
584
elevenlabs-dialogue
Generate multi-speaker dialogue audio with different voices in a single file using the inference.sh CLI.
584
ai-podcast-creation
Create AI-powered podcasts and audio content using text-to-speech, music generation, and audio editing via the inference.sh CLI.
584
audiocraft-audio-generation
Generate music and sound effects from text descriptions using Meta's AudioCraft library, with support for melody conditioning, stereo output, and style transfer.
10.4k · bundle
asr
Transcribe audio files to text using the z-ai-web-dev-sdk, with CLI and SDK examples for single files, batches, and directories.
567 · bundle
whisper
Transcribe and translate audio across 99 languages using OpenAI's Whisper model, with options for model size, language detection, timestamps, and batch processing.
2
daily
Provides a reference for building real-time voice and multimodal AI agents with Pipecat, covering pipeline architecture, speech services, LLM integration, transports, and deployment.
0 · bundle
gemini-live-api-dev
Builds real-time, bidirectional streaming applications with the Gemini Live API, covering WebSocket audio/video/text streaming, VAD, function calling, session management, ephemeral tokens, and live translation across Python and JavaScript SDKs.
0
mmx-cli
Generates text, images, video, speech, and music, and performs web searches via the MiniMax AI platform using the mmx terminal CLI.
3
ast-eval
Benchmarks automatic speech translation and recognition on English-French and English-Romanian datasets, reporting BLEU and WER on tokenized outputs.
3
digital-health-clinical-asr-build
Curates clinical-specialty term lists, generates IPA-tagged synthetic audio via TTS, and produces NeMo-format manifests for ASR benchmark evaluation.
2.2k · bundle
songsee
Generates spectrograms and multi-panel audio feature visualizations from audio files via a command-line tool.
61
videocut
Generates and burns subtitles into videos: extracts audio, transcribes via Volcano Engine, corrects errors, reviews, and burns subtitles with ffmpeg.
61
stt
Transcribes audio files to text using OpenAI Whisper, optimized for Brazilian Portuguese, with support for common audio formats and timestamped output.
32 · bundle
fal-audio
Converts text to speech and speech to text using fal.ai audio models.
2
asr
Transcribes audio from URLs or local files into text using a low-cost speech-to-text API, with options for speaker diarization, word timestamps, and multiple output formats.
1 · bundle
sag
Generates speech from text using ElevenLabs text-to-speech with a command-line interface and local playback.
1 · bundle
asr
Transcribes audio from URLs or local files into text with speaker diarization, word timestamps, and multiple output formats via a command-line tool.
10 · bundle
fal-api
Generates images, videos, and audio transcripts using fal.ai's API, supporting models like FLUX, Stable Diffusion, and Whisper.
1 · bundle
daily
Reference for building real-time voice and multimodal AI applications with Daily and Pipecat, covering pipeline architecture, AI service integrations, transports, and client SDKs.
3
fal-audio
Converts text to speech and speech to text using fal.ai audio models.
5