Plugins
2 pluginscurated
Text-to-Speech Podcast
Convert a script into a multi-speaker podcast audio with TTS, dialogue, and optional music.
4 skills · plugin
curated
Build Gemini Live API App
Build real-time, bidirectional streaming applications with the Gemini Live API, covering WebSocket-based audio/video/text streaming and function calling.
4 skills · plugin
Results for “audio”
335 skillsAscii Video
Production pipeline for ASCII art video — any format. Converts video/audio/images/generative input into colored ASCII character video output (MP4, GIF, image sequence). Covers: video-to-ASCII conversion, audio-reactive music visualizers, generative ASCII art animations, hybrid video+audio reactive, text/lyrics overlays, real-time terminal rendering. Use when users request: ASCII video, text art video, terminal-style video, character art animation, retro text visualization, audio visualizer in ASCII, converting video to ASCII art, matrix-style effects, or any animated ASCII output.
0 · bundle
Notebooklm
Browser-automates Google's NotebookLM to read notebooks, add sources, generate Studio outputs (Audio/Video Overviews, Mind Maps, Reports), and create new notebooks.
20.4k · bundle
Elevenlabs Music
Generate original music from text prompts using ElevenLabs AI, with control over genre, mood, instruments, and duration up to 10 minutes.
584
Elevenlabs Sound Effects
Generate AI sound effects from text descriptions using the inference.sh CLI, with control over duration and prompt influence.
584
Whisper
Transcribe and translate audio across 99 languages using OpenAI's Whisper model, with options for model size, language detection, timestamps, and batch processing.
2
Remotion Best Practices
Loads Remotion best practices for React video creation, covering animations, composition setup, assets, audio, captions, sequencing, transitions, media metadata, Three.js, Tailwind, and rendering patterns.
7 · bundle
Asr
Transcribes audio from URLs or local files to text using the Speech is Cheap API, with options for speaker diarization, timestamps, and multiple output formats.
1 · bundle
Transcribe
Transcribe audio files to text with optional diarization and known-speaker hints. Use when a user asks to transcribe speech from audio/video, extract text from recordings, or label speakers in interviews or meetings.
0 · bundle
Higgsfield Generate
Generate images, videos, 3D assets, and audio via the Higgsfield AI CLI, including Marketing Studio ads and Virality Predictor analysis.
518 · bundle
Hyperframes Creative
Provides creative direction for HyperFrames videos, handling design specs, palettes, typography, narration, beat planning, audio-reactive visuals, composition patterns, and brand/style decisions.
· bundle
Videodb
Ingests video and audio from files, URLs, and live streams, builds searchable visual and spoken indexes, edits timelines with subtitles and overlays, and generates real-time alerts.
3 · bundle
Sdr
Quantifies audio source separation quality by computing the signal-to-distortion ratio (SDR) between ground-truth and estimated stems, with per-stem and record-level averaging.
3
Transcribe
Transcrever arquivos de áudio para texto com diarização opcional e dicas de falantes conhecidos. Use quando um usuário pedir para transcrever fala de áudio/vídeo, extrair texto de gravações ou identificar falantes em entrevistas ou reuniões.
10 · bundle
Specialized
Specialized hardware interfaces and system reliability for Zephyr RTOS. Covers LVGL GUI development, Audio I2S/Codecs, Watchdog timers, and Fault Injection. Trigger when building human-machine interfaces (HMI), audio devices, or high-reliability mission-critical systems.
60 · bundle
Ascii Video
ASCII video: convert video/audio to colored ASCII MP4/GIF.
0 · bundle
Ascii Video
ASCII video: convert video/audio to colored ASCII MP4/GIF.
1 · bundle
Ascii Video
ASCII video: convert video/audio to colored ASCII MP4/GIF.
0 · bundle
Ascii Video
ASCII video: convert video/audio to colored ASCII MP4/GIF.
0 · bundle
Videodb
Ingest, index, search, and edit video and audio from files, URLs, live streams, or desktop sessions with timestamps, subtitles, overlays, and real-time alerts.
42.4k · bundle
Voice Changer
Transform the voice in an audio recording into a different target voice while preserving emotion, timing, and delivery using the ElevenLabs Voice Changer API.
363 · bundle
Asr
Transcribes audio from URLs or local files into text using a low-cost speech-to-text API, with options for speaker diarization, word timestamps, and multiple output formats.
1 · bundle
Videodb
Ingest, index, search, edit, and generate video and audio assets from files, URLs, RTSP feeds, or desktop capture, with real-time alerts and stream links.
1 · bundle
Azure Speech To Text REST Py
Transcribe short audio files (up to 60 seconds) using Azure Speech-to-Text REST API with Python, without requiring the Speech SDK.
2.7k · bundle
Speech To Text
Transcribe audio to text using ElevenLabs Scribe and Whisper models via the inference.sh CLI, supporting timestamps, speaker diarization, translation, and multi-language transcription.
584
Meeting Transcription
Transcribe meeting audio with speaker diarization, generate structured summaries with action items, decisions, and follow-ups, and support multiple audio formats and languages. Use when the user requests meeting transcription or provides relevant inputs for this workflow.
159
Songsee
Generate spectrograms and feature-panel visualizations from audio with the songsee CLI.
9
Songsee
Generate spectrograms and feature-panel visualizations from audio with the songsee CLI.
2 · bundle
Songsee
Generate spectrograms and feature-panel visualizations from audio with the songsee CLI.
0
Replicate Run
Run any Replicate model (image gen, audio, video) by version ID
118 · bundle
Songsee
Generate spectrograms and feature-panel visualizations from audio with the songsee CLI.
0
Songsee
Generate spectrograms and feature-panel visualizations from audio with the songsee CLI.
0
Songsee
Generate spectrograms and feature-panel visualizations from audio with the songsee CLI.
228
Markitdown
Convert various file formats (PDF, Office documents, images, audio, web content, structured data) to Markdown optimized for LLM processing. Use when converting documents to markdown, extracting text from PDFs/Office files, transcribing audio, performing OCR on images, extracting YouTube transcripts, or processing batches of files. Supports 20+ formats including DOCX, XLSX, PPTX, PDF, HTML, EPUB, CSV, JSON, images with OCR, and audio with transcription.
1k · bundle
Videodb
Ingests video and audio from files, URLs, live feeds, or desktop capture; indexes and searches moments with timestamps; transcodes, edits timelines, generates media assets, and emits real-time alerts.
1 · bundle
Sound Cues
Create and modify SoundCue assets — add/connect nodes (mixer, random, delay, attenuation, modulator, etc.) and set audio properties (SoundCueService). Use when the user asks to create a Sound Cue, build a SoundCue node graph, or wire audio playback logic.
605 · bundle
Summarize
Summarize URLs or files with the summarize CLI (web, PDFs, images, audio, YouTube).
0 · bundle