Packs
2 packscurated
Text-to-Speech Podcast
Convert a script into a multi-speaker podcast audio with TTS, dialogue, and optional music.
4 skills · pack
curated
Build Gemini Live API App
Build real-time, bidirectional streaming applications with the Gemini Live API, covering WebSocket-based audio/video/text streaming and function calling.
4 skills · pack
Results for “audio”
41 skillsgame-audio
Expert game audio designer and implementer specializing in interactive sound design, adaptive music systems, spatial audio, and audio middleware integration. Brings deep knowledge of FMOD, Wwise, and native engine audio systems to create immersive sonic experiences that respond dynamically to gameplay. Use when "game audio, sound design, game music, FMOD, Wwise, spatial audio, 3D sound, audio middleware, game sfx, adaptive music, interactive audio, audio bus, game mixing, audio occlusion, reverb zones, audio pooling, sound manager, audio, sound, music, game-audio, fmod, wwise, spatial-audio, middleware, mixing, sound-design" mentioned.
128 · bundle
game-audio
Game audio principles. Sound design, music integration, adaptive audio systems.
3
game-development-game-audio-engineer
Interactive audio specialist - Masters FMOD/Wwise integration, adaptive music systems, spatial audio, and audio performance budgeting across all game engines
2
openai-whisper-api
Transcribe audio via OpenAI Audio Transcriptions API (Whisper).
9 · bundle
sound-engineer
Expert in spatial audio, procedural sound design, game audio middleware, and app UX sound design. Specializes in HRTF/Ambisonics, Wwise/FMOD integration, UI sound design, and adaptive music systems. Activate on 'spatial audio', 'HRTF', 'binaural', 'Wwise', 'FMOD', 'procedural sound', 'footstep system', 'adaptive music', 'UI sounds', 'notification audio', 'sonic branding'. NOT for music composition/production (use DAW), audio post-production for film (linear media), voice cloning/TTS (use voice-audio-engineer), podcast editing (use standard audio editors), or hardware design.
10 · bundle
speech
Generate spoken audio from text using OpenAI's API with built-in voices for narrated explainers, lecture audio, and quick voiceover tracks.
1
More results
voice-isolator
Remove background noise and isolate vocals or speech from audio files using the ElevenLabs Voice Isolator API.
363 · bundle
audio-scrape
Discovers podcasts via the iTunes Search API, parses RSS feeds, downloads audio, transcribes with OpenAI Whisper, chunks transcripts, and upserts results into a database table.
1
app-sound-design
Design decisions for when and what sounds to use in web apps, mobile apps, and desktop native apps — plus how to create them with AI. Covers UI sound vocabulary (confirms, alerts, transitions, errors), platform conventions (iOS, macOS, Material), creating sounds with AI tools (ElevenLabs, AudioCraft, Stable Audio), Web Audio API for procedural sounds, haptic-audio pairing, and the psychology of satisfying vs. annoying UI sounds. Activate on 'app sounds', 'UI sounds', 'notification sound', 'sound design app', 'earcons', 'sonic branding', 'interaction sounds', 'feedback sounds', 'app audio', 'satisfying sounds'. NOT for professional audio engineering (use sound-engineer), Windows 3.1 retro sounds specifically (use win31-audio-design), music composition (use sound-engineer).
10
tts
Converts text to speech and generates MP3 audio files using Hume AI or OpenAI APIs.
1 · bundle
gladia-automation
Automate Gladia speech-to-text and audio processing tasks through Composio's Gladia toolkit via Rube MCP.
66.9k
runapi-cli
Generate AI images, videos, and music/audio from agents using the RunAPI CLI.
42.4k
deepgram-automation
Automate Deepgram speech-to-text and audio processing tasks through Composio's Deepgram toolkit via Rube MCP.
66.9k
tts
Converts text to speech and generates MP3 audio files using Hume AI or OpenAI APIs, printing the file path for delivery.
32 · bundle
groqcloud-automation
Automate AI inference, chat completions, audio translation, and TTS voice management through GroqCloud's high-performance API via Composio.
66.9k
tts
Converts text to speech and generates MP3 audio files using Hume AI or OpenAI APIs, printing the file path for delivery.
10 · bundle
tts
Converts text to speech and generates MP3 audio files using the Hume AI or OpenAI API, printing the file path for delivery.
1 · bundle
azure-ai-openai-dotnet
Integrate Azure OpenAI and OpenAI services in .NET applications for chat completions, embeddings, image generation, audio transcription, and assistants.
2.7k
videodb
Ingest, index, search, and edit video and audio from files, URLs, live streams, or desktop sessions with timestamps, subtitles, overlays, and real-time alerts.
42.4k · bundle
openai-whisper-api
OpenAI Audio Transcriptions API via curl; gpt-4o-transcribe, mini, diarize, or whisper-1.
0 · bundle
voice-changer
Transform the voice in an audio recording into a different target voice while preserving emotion, timing, and delivery using the ElevenLabs Voice Changer API.
363 · bundle
elevenlabs-automation
Automate ElevenLabs text-to-speech workflows: generate speech from text, browse and inspect voices, check subscription limits, list models, stream audio, and retrieve history via the Composio MCP integration.
66.9k
comfyui
Generate images, video, and audio with ComfyUI — install, launch, manage nodes/models, run workflows with parameter injection. Uses the official comfy-cli for lifecycle and direct REST/WebSocket API for execution.
2
transloadit-media-processing
Encode video to HLS/MP4, generate thumbnails, resize or watermark images, extract audio, concatenate clips, add subtitles, OCR documents, and build multi-step media processing pipelines using Transloadit's cloud infrastructure.
36.2k
openai-api
OpenAI API skill. Use when working with OpenAI for assistants, audio, batches. Covers 148 endpoints.
6 · bundle
gemini-live-api-dev
Build real-time, bidirectional streaming applications with the Gemini Live API, covering WebSocket-based audio/video/text streaming, voice activity detection, function calling, session management, and ephemeral tokens.
3.8k
gemini-live-api-dev
Builds real-time, bidirectional streaming applications with the Gemini Live API, covering WebSocket audio/video/text streaming, VAD, function calling, session management, ephemeral tokens, and live translation across Python and JavaScript SDKs.
0
speech
Use quando o usuário solicita narração em texto-para-fala, voiceovers de acessibilidade, prompts de áudio ou geração em lote via OpenAI Audio API; execute a CLI incluída (`scripts/text_to_speech.py`) com vozes integradas e requer `OPENAI_API_KEY` para chamadas diretas. Criação de vozes customizadas está fora do escopo.
10 · bundle
notebooklm
Provides programmatic access to Google NotebookLM, enabling creation of notebooks, adding sources, generating artifacts, and downloading results in multiple formats.
2 · bundle
text-to-speech
Generate natural speech from text using ElevenLabs voice AI, supporting 70+ languages, multiple models, and various output formats.
363 · bundle
comfyui
Generate images, video, and audio with ComfyUI — install, launch, manage nodes/models, run workflows with parameter injection. Uses the official comfy-cli for lifecycle and direct REST/WebSocket API for execution.
0 · bundle
ima-sdk-basics
Integrate client-side video and audio ads using the IMA SDK across web, Android, iOS, and TV platforms with VAST/VMAP support.
14.4k · bundle
videodb
Ingest, index, search, edit, and generate video and audio content from files, URLs, live streams, or desktop capture.
226k · bundle
notebooklm-integration
Wraps the notebooklm-py CLI to ingest sources and generate synthesized artifacts like podcasts, slide decks, and quizzes from Google NotebookLM.
2
voice-ai
Generates speech, transcribes audio, clones voices, and builds real-time voice agents using ElevenLabs, OpenAI TTS, Whisper, and Vapi.
10
fal-api
Generates images, videos, and audio transcripts using fal.ai's API, supporting models like FLUX, Stable Diffusion, and Whisper.
1 · bundle