Speech & Audio Agent Skills

Speech & Audio

107 skills
inference-sh
ai-podcast
Generate multi-person talking head podcast videos from scratch using AI — character creation, TTS, avatar animation, and video stitching.
584
inference-sh
dialogue-audio
Create realistic multi-speaker dialogue audio using Dia TTS via the inference.sh CLI, with control over speaker tags, emotion, pacing, and conversation structure.
584
inference-sh
elevenlabs-stt
Transcribe audio with high accuracy using ElevenLabs Scribe models, supporting speaker diarization, audio event tagging, forced alignment, and subtitle generation via the inference.sh CLI.
584
inference-sh
elevenlabs-tts
Generate high-quality speech from text using ElevenLabs' premium voices, with support for 32 languages, multiple models, and voice tuning parameters.
584
inference-sh
speech-to-text
Transcribe audio to text using ElevenLabs Scribe and Whisper models via the inference.sh CLI, supporting timestamps, speaker diarization, translation, and multi-language transcription.
584
inference-sh
ai-voice-cloning
Generate natural AI voices, text-to-speech, and voice synthesis using the inference.sh CLI with models like Inworld TTS, ElevenLabs, and Kokoro TTS for voiceovers, audiobooks, podcasts, and more.
584
inference-sh
elevenlabs-music
Generate original music from text prompts using ElevenLabs AI, with control over genre, mood, instruments, and duration up to 10 minutes.
584
inference-sh
elevenlabs-dubbing
Translate and dub audio/video into 29 languages while preserving speaker voice using the inference.sh CLI.
584
inference-sh
ai-music-generation
Generate music and songs using ElevenLabs, Diffrythm, and Tencent Song Generation models via the inference.sh CLI.
584
inference-sh
ai-content-pipeline
Build multi-step AI content creation pipelines combining image, video, audio, and text using the inference.sh CLI.
584
inference-sh
talking-head-production
Create talking head videos with AI avatars, lipsync, and voiceover using the inference.sh CLI.
584
inference-sh
elevenlabs-sound-effects
Generate AI sound effects from text descriptions using the inference.sh CLI, with control over duration and prompt influence.
584
inference-sh
elevenlabs-voice-changer
Transform any voice into a different voice while preserving speech content and emotion using the inference.sh CLI and ElevenLabs models.
584
inference-sh
elevenlabs-voice-isolator
Remove background noise and isolate vocals from audio files using the inference.sh CLI and ElevenLabs voice isolator.
584
orchestra-research
whisper
Transcribe and translate speech across 99 languages using OpenAI's Whisper model, with support for multiple model sizes, batch processing, and subtitle generation.
10.4k · bundle
higgsfield-ai
higgsfield-generate
Generate images, videos, 3D assets, and audio via the Higgsfield AI CLI, including Marketing Studio ads and Virality Predictor analysis.
518 · bundle
samuraigpt
muapi-media-generation
Generate AI images, videos, music, and audio from the terminal via muapi.ai — supports 100+ models including Flux, Midjourney v7, Kling 3.0, Veo3, and Suno V5.
3.7k · bundle
samuraigpt
muapi-music-video
Generates a short music video from a song theme by creating keyframes, animating them, and producing a matching soundtrack.
3.7k
heygen
media-use
Resolves, generates, and operates on media assets (audio, images, icons, logos, voice, color grades, LUTs) for HyperFrames projects, using a local cache and the HeyGen CLI for free-usage catalog search and TTS.
· bundle
majiayu000
fal
Generate images, videos, audio, and more using fal.ai AI models, with support for text-to-image, image-to-video, text-to-speech, speech-to-text, image editing, upscaling, model search, workflow creation, and cost estimation.
567 · bundle
comeonoliver
sag
Generates speech from text using ElevenLabs TTS with local playback, supporting voice selection, pronunciation rules, and audio tags.
61
comeonoliver
speech
Generates spoken audio clips from text for narration, voiceovers, IVR prompts, and accessibility reads, with support for single clips and batch processing.
61
dvcrn
fal
Search, explore, and run fal.ai generative AI models for image, video, audio, and 3D generation, including schema lookup, job submission, status polling, result retrieval, and file uploads.
32 · bundle
lingxling
daily
Reference for building real-time voice and multimodal AI applications with Pipecat, covering pipelines, speech services, LLM integration, and transports.
253
nimoqup046-collab
daily
Reference for building real-time voice and multimodal AI agents with Pipecat, covering pipelines, speech services, LLMs, transports, and deployment.
2
sakamoto-family-smile
fal-ai-media
Generates images, videos, and audio using fal.ai models via MCP, covering text-to-image, text/image-to-video, text-to-speech, and video-to-audio.
0
lord1egypt
songsee
Generates spectrograms and multi-panel audio feature visualizations (mel, chroma, MFCC) from audio files via a Go CLI.
2
lord1egypt
heartmula
Generates full songs from lyrics and tags using the open-source HeartMuLa music models, with multilingual support and local GPU or CPU inference.
2
kbarbel640-del
fal
Search, explore, and run fal.ai generative AI models for image, video, audio, and 3D generation, including queue management and file uploads.
1 · bundle
joshuashepherd
audio-scrape
Discovers podcasts via the iTunes Search API, parses RSS feeds, downloads audio, transcribes with OpenAI Whisper, chunks transcripts, and upserts results into a database table.
1
joshuashepherd
fal-ai-media
Generates images, videos, and audio using fal.ai models via MCP, covering text-to-image, text/image-to-video, text-to-speech, and video-to-audio.
1
gabrielmoreira
vox-explainer
Produces a complete narrated, subtitled, scored explainer video from a single topic prompt using a six-stage pipeline with script, voiceover, keyframes, animation, music, and local assembly.
17 · bundle
gabrielmoreira
ai-media-generator
Generates high-quality prompts for AI image, video, and music generation platforms, with optional browser automation to submit them to target sites.
17 · bundle
mhassan0000
fal-ai-media
Generates images, videos, and audio using fal.ai models via MCP, covering text-to-image, text/image-to-video, text-to-speech, and video-to-audio.
1
diegosouzapw
vox
Runs a local voice MCP server in Rust for text-to-speech and speech-to-text, with build, test, and configuration guidance.
54 · bundle
oyi77
voice-ai
Generates speech, transcribes audio, clones voices, and builds real-time voice agents using ElevenLabs, OpenAI TTS, Whisper, and Vapi.
10