Results for “audio-processing”
30 skillsgladia-automation
Automate Gladia speech-to-text and audio processing tasks through Composio's Gladia toolkit via Rube MCP.
66.9k
deepgram-automation
Automate Deepgram speech-to-text and audio processing tasks through Composio's Deepgram toolkit via Rube MCP.
66.9k
whisper
Transcribe and translate audio across 99 languages using OpenAI's Whisper model, with options for model size, language detection, timestamps, and batch processing.
2
videodb
Ingest, index, search, and edit video and audio from files, URLs, live streams, or desktop sessions with timestamps, subtitles, overlays, and real-time alerts.
42.4k · bundle
ascii-video
Converts video, audio, or images into colored ASCII art videos (MP4/GIF) with generative effects, audio-reactive visuals, and text overlays.
2 · bundle
elevenlabs-stt
Transcribe audio with high accuracy using ElevenLabs Scribe models, supporting speaker diarization, audio event tagging, forced alignment, and subtitle generation via the inference.sh CLI.
584
More results
speech-to-text
Transcribe audio to text using ElevenLabs Scribe v2, supporting 90+ languages, speaker diarization, and word-level timestamps.
363 · bundle
speech
Generates spoken audio clips from text for narration, voiceovers, IVR prompts, and accessibility reads, with support for single clips and batch processing.
61
whisper
Transcribe and translate speech across 99 languages using OpenAI's Whisper model, with support for multiple model sizes, batch processing, and subtitle generation.
10.4k · bundle
voice-changer
Transform the voice in an audio recording into a different target voice while preserving emotion, timing, and delivery using the ElevenLabs Voice Changer API.
363 · bundle
songsee
Generates spectrograms and multi-panel audio feature visualizations from audio files via a command-line tool.
61
sound-effects
Generate sound effects from text descriptions using ElevenLabs, with support for looping, duration control, and prompt influence tuning.
363 · bundle
transcribe
Transcribe audio files to text with optional speaker diarization and known-speaker hints using OpenAI models.
23.3k · bundle
songsee
Generates spectrograms and multi-panel audio feature visualizations (mel, chroma, MFCC) from audio files via a Go CLI.
2
detecting-deepfake-audio-in-vishing-attacks
Detects AI-generated deepfake audio used in voice phishing (vishing) attacks by extracting spectral features and classifying samples with machine learning models.
24.6k · bundle
tone
Game audio generation agent. Produces code (Python/JS/TS/Shell) for SFX, BGM, Voice, Ambient, and UI sounds using ElevenLabs/Stable Audio/MusicGen/Suno/OpenAI TTS/JSFXR. Handles LUFS normalization and middleware integration.
65 · bundle
muapi-media-generation
Generate AI images, videos, music, and audio from the terminal via muapi.ai — supports 100+ models including Flux, Midjourney v7, Kling 3.0, Veo3, and Suno V5.
3.7k · bundle
agent-songsee-v2
Expert en analyse audio avancé (spectrograms, mel, chroma, MFCC, feature extraction, CLI)
6
videodb
Ingest, index, search, edit, and generate video and audio content from files, URLs, live streams, or desktop capture.
226k · bundle
187-step-459c2d7b
Guides analysis of Neuropixels recordings from raw data to curated units, covering preprocessing, motion correction, spike sorting, quality metrics, and export.
7 · bundle
fal-audio
Converts text to speech and speech to text using fal.ai audio models.
5
songsee
Audio spectrograms/features (mel, chroma, MFCC) via CLI.
0
whisper
OpenAI's general-purpose speech recognition model. Supports 99 languages, transcription, translation to English, and language identification. Six model sizes from tiny (39M params) to large (1550M params). Use for speech-to-text, podcast transcription, or multilingual audio processing. Best for robust, multilingual ASR.
3 · bundle
whisper
OpenAI's general-purpose speech recognition model. Supports 99 languages, transcription, translation to English, and language identification. Six model sizes from tiny (39M params) to large (1550M params). Use for speech-to-text, podcast transcription, or multilingual audio processing. Best for robust, multilingual ASR.
1 · bundle
whisper
OpenAI's general-purpose speech recognition model. Supports 99 languages, transcription, translation to English, and language identification. Six model sizes from tiny (39M params) to large (1550M params). Use for speech-to-text, podcast transcription, or multilingual audio processing. Best for robust, multilingual ASR.
0 · bundle
whisper
OpenAI's general-purpose speech recognition model. Supports 99 languages, transcription, translation to English, and language identification. Six model sizes from tiny (39M params) to large (1550M params). Use for speech-to-text, podcast transcription, or multilingual audio processing. Best for robust, multilingual ASR.
0 · bundle
whisper
OpenAI's general-purpose speech recognition model. Supports 99 languages, transcription, translation to English, and language identification. Six model sizes from tiny (39M params) to large (1550M params). Use for speech-to-text, podcast transcription, or multilingual audio processing. Best for robust, multilingual ASR.
0 · bundle
whisper
OpenAI's general-purpose speech recognition model. Supports 99 languages, transcription, translation to English, and language identification. Six model sizes from tiny (39M params) to large (1550M params). Use for speech-to-text, podcast transcription, or multilingual audio processing. Best for robust, multilingual ASR.
0 · bundle
whisper
OpenAI's general-purpose speech recognition model. Supports 99 languages, transcription, translation to English, and language identification. Six model sizes from tiny (39M params) to large (1550M params). Use for speech-to-text, podcast transcription, or multilingual audio processing. Best for robust, multilingual ASR.
0 · bundle
whisper
OpenAI's general-purpose speech recognition model. Supports 99 languages, transcription, translation to English, and language identification. Six model sizes from tiny (39M params) to large (1550M params). Use for speech-to-text, podcast transcription, or multilingual audio processing. Best for robust, multilingual ASR.
0 · bundle