Plugins
2 pluginscurated
Text-to-Speech Podcast
Convert a script into a multi-speaker podcast audio with TTS, dialogue, and optional music.
4 skills · plugin
curated
Build Gemini Live API App
Build real-time, bidirectional streaming applications with the Gemini Live API, covering WebSocket-based audio/video/text streaming and function calling.
4 skills · plugin
Results for “audio”
158 skillsSongsee
Generates spectrograms and multi-panel audio feature visualizations from audio files via a command-line tool.
61
Agent Audiocraft V2
Expert en AudioCraft avancé (MusicGen text-to-music, AudioGen text-to-sound, fine-tuning)
6
Audiocraft Audio Generation
Generate music and sound effects from text descriptions using Meta's AudioCraft library, with support for melody conditioning, stereo output, and style transfer.
10.4k · bundle
Fal Audio
Convert text to speech and speech to text using fal.ai audio models.
42.4k
Fal Audio
Converts text to speech and speech to text using fal.ai audio models.
2
Fal Audio
Converts text to speech and speech to text using fal.ai audio models.
5
More results
Songsee
Generates spectrograms and multi-panel audio feature visualizations (mel, chroma, MFCC) from audio files via a Go CLI.
2
Elevenlabs Dialogue
Generate multi-speaker dialogue audio with different voices in a single file using the inference.sh CLI.
584
Stt
Transcribes audio files to text using OpenAI Whisper, optimized for Brazilian Portuguese, with support for common audio formats and timestamped output.
32 · bundle
Speech
Generate spoken audio for narration, voiceovers, IVR prompts, and accessibility reads using the OpenAI Audio API with bundled CLI and built-in voices.
23.3k · bundle
Dialogue Audio
Create realistic multi-speaker dialogue audio using Dia TTS via the inference.sh CLI, with control over speaker tags, emotion, pacing, and conversation structure.
584
Ascii Video
Converts video, audio, or images into colored ASCII art videos (MP4/GIF) with generative effects, audio-reactive visuals, and text overlays.
2 · bundle
Asr
Transcribe audio files to text using the z-ai-web-dev-sdk, with CLI and SDK examples for single files, batches, and directories.
567 · bundle
Audio Scrape
Discovers podcasts via the iTunes Search API, parses RSS feeds, downloads audio, transcribes with OpenAI Whisper, chunks transcripts, and upserts results into a database table.
1
AI Podcast Creation
Create AI-powered podcasts and audio content using text-to-speech, music generation, and audio editing via the inference.sh CLI.
584
Elevenlabs Stt
Transcribe audio with high accuracy using ElevenLabs Scribe models, supporting speaker diarization, audio event tagging, forced alignment, and subtitle generation via the inference.sh CLI.
584
Detecting Deepfake Audio In Vishing Attacks
Detects AI-generated deepfake audio used in voice phishing (vishing) attacks by extracting spectral features and classifying samples with machine learning models.
24.6k · bundle
Transcribe
Transcribe audio files to text with optional speaker diarization and known-speaker hints using OpenAI models.
23.3k · bundle
Caa Eval
Benchmarks large audio-language models against adversarial audio attacks using the CAA dataset, computing WER, ROUGE-L, cosine similarity, and coherence scores to assess robustness in conversational settings.
3
Sound Effects
Generate sound effects from text descriptions using ElevenLabs, with support for looping, duration control, and prompt influence tuning.
363 · bundle
Sag
Generates speech from text using ElevenLabs TTS with local playback, supporting voice selection, pronunciation rules, and audio tags.
61
Tone
Game audio generation agent. Produces code (Python/JS/TS/Shell) for SFX, BGM, Voice, Ambient, and UI sounds using ElevenLabs/Stable Audio/MusicGen/Suno/OpenAI TTS/JSFXR. Handles LUFS normalization and middleware integration.
65 · bundle
Tts
Converts text to speech and generates MP3 audio files using the Hume AI or OpenAI API, printing the file path for delivery.
1 · bundle
Songsee
Audio spectrograms/features (mel, chroma, MFCC) via CLI.
0
Voice Changer
Transform the voice in an audio recording into a different target voice while preserving emotion, timing, and delivery using the ElevenLabs Voice Changer API.
363 · bundle
Fal AI Media
Generates images, videos, and audio using fal.ai models via MCP, covering text-to-image, text/image-to-video, text-to-speech, and video-to-audio.
0
Runapi CLI
Generate AI images, videos, and music/audio from agents using the RunAPI CLI.
42.4k
Videodb
Ingest, index, search, and edit video and audio content with timestamps, subtitles, overlays, and live-stream alerts.
5 · bundle
Fal AI Media
Generate images, videos, and audio using fal.ai models via MCP tools, with support for text-to-image, text/image-to-video, text-to-speech, and video-to-audio.
226k
Fal AI Media
Unified media generation via fal.ai MCP — image, video, and audio. Covers text-to-image (Nano Banana), text/image-to-video (Seedance, Kling, Veo 3), text-to-speech (CSM-1B), and video-to-audio (ThinkSound). Use when the user wants to generate images, videos, or audio with AI.
1
Fal AI Media
Unified media generation via fal.ai MCP — image, video, and audio. Covers text-to-image (Nano Banana), text/image-to-video (Seedance, Kling, Veo 3), text-to-speech (CSM-1B), and video-to-audio (ThinkSound). Use when the user wants to generate images, videos, or audio with AI.
0
Fal AI Media
Unified media generation via fal.ai MCP — image, video, and audio. Covers text-to-image (Nano Banana), text/image-to-video (Seedance, Kling, Veo 3), text-to-speech (CSM-1B), and video-to-audio (ThinkSound). Use when the user wants to generate images, videos, or audio with AI.
0
Gladia Automation
Automate Gladia speech-to-text and audio processing tasks through Composio's Gladia toolkit via Rube MCP.
66.9k
Deepgram Automation
Automate Deepgram speech-to-text and audio processing tasks through Composio's Deepgram toolkit via Rube MCP.
66.9k
Fal AI Media
Unified media generation via fal.ai MCP — image, video, and audio. Covers text-to-image (Nano Banana), text/image-to-video (Seedance, Kling, Veo 3), text-to-speech (CSM-1B), and video-to-audio (ThinkSound). Use when the user wants to generate images, videos, or audio with AI.
0
Fal AI Media
Unified media generation via fal.ai MCP — image, video, and audio. Covers text-to-image (Nano Banana), text/image-to-video (Seedance, Kling, Veo 3), text-to-speech (CSM-1B), and video-to-audio (ThinkSound). Use when the user wants to generate images, videos, or audio with AI.
1