Results for “audio-transcription”

31 skills
More results
majiayu000
Asr
Transcribe audio files to text using the z-ai-web-dev-sdk, with CLI and SDK examples for single files, batches, and directories.
567 · bundle
microsoft
Azure AI Contentunderstanding Py
Extract semantic content from documents, images, audio, and video using Azure AI Content Understanding SDK for Python.
2.7k
elevenlabs
Speech To Text
Transcribe audio to text using ElevenLabs Scribe v2, supporting 90+ languages, speaker diarization, and word-level timestamps.
363 · bundle
inference-sh
Elevenlabs Stt
Transcribe audio with high accuracy using ElevenLabs Scribe models, supporting speaker diarization, audio event tagging, forced alignment, and subtitle generation via the inference.sh CLI.
584
demerzels-lab
Asr
Transcribes audio from URLs or local files into text with speaker diarization, word timestamps, and multiple output formats via a command-line tool.
10 · bundle
lord1egypt
Whisper
Transcribe and translate audio across 99 languages using OpenAI's Whisper model, with options for model size, language detection, timestamps, and batch processing.
2
johnalbertini14-glitch
Asr
Transcribes audio from URLs or local files to text using the Speech is Cheap API, with options for speaker diarization, timestamps, and multiple output formats.
1 · bundle
antigravity
Videodb
Ingest, index, search, and edit video and audio from files, URLs, live streams, or desktop sessions with timestamps, subtitles, overlays, and real-time alerts.
42.4k · bundle
orchestra-research
Whisper
Transcribe and translate speech across 99 languages using OpenAI's Whisper model, with support for multiple model sizes, batch processing, and subtitle generation.
10.4k · bundle
lord1egypt
Songsee
Generates spectrograms and multi-panel audio feature visualizations (mel, chroma, MFCC) from audio files via a Go CLI.
2
aniruddhaadak80
Whisper
OpenAI's general-purpose speech recognition model. Supports 99 languages, transcription, translation to English, and language identification. Six model sizes from tiny (39M params) to large (1550M params). Use for speech-to-text, podcast transcription, or multilingual audio processing. Best for robust, multilingual ASR.
0 · bundle
tianhao909
Whisper
OpenAI's general-purpose speech recognition model. Supports 99 languages, transcription, translation to English, and language identification. Six model sizes from tiny (39M params) to large (1550M params). Use for speech-to-text, podcast transcription, or multilingual audio processing. Best for robust, multilingual ASR.
1 · bundle
qcmuu
Whisper
OpenAI's general-purpose speech recognition model. Supports 99 languages, transcription, translation to English, and language identification. Six model sizes from tiny (39M params) to large (1550M params). Use for speech-to-text, podcast transcription, or multilingual audio processing. Best for robust, multilingual ASR.
0 · bundle
kbarbel640-del
Asr
Transcribes audio from URLs or local files into text using a low-cost speech-to-text API, with options for speaker diarization, word timestamps, and multiple output formats.
1 · bundle
aniruddhaadak80
Songsee
Audio spectrograms/features (mel, chroma, MFCC) via CLI.
0
peteedoo
Whisper
OpenAI's general-purpose speech recognition model. Supports 99 languages, transcription, translation to English, and language identification. Six model sizes from tiny (39M params) to large (1550M params). Use for speech-to-text, podcast transcription, or multilingual audio processing. Best for robust, multilingual ASR.
0 · bundle
ichichuang
Whisper
OpenAI's general-purpose speech recognition model. Supports 99 languages, transcription, translation to English, and language identification. Six model sizes from tiny (39M params) to large (1550M params). Use for speech-to-text, podcast transcription, or multilingual audio processing. Best for robust, multilingual ASR.
0 · bundle
bog5d
Whisper
OpenAI's general-purpose speech recognition model. Supports 99 languages, transcription, translation to English, and language identification. Six model sizes from tiny (39M params) to large (1550M params). Use for speech-to-text, podcast transcription, or multilingual audio processing. Best for robust, multilingual ASR.
0 · bundle
q2805187159
Whisper
OpenAI's general-purpose speech recognition model. Supports 99 languages, transcription, translation to English, and language identification. Six model sizes from tiny (39M params) to large (1550M params). Use for speech-to-text, podcast transcription, or multilingual audio processing. Best for robust, multilingual ASR.
3 · bundle
qhjqhj00
Ast Eval
Benchmarks automatic speech translation and recognition on English-French and English-Romanian datasets, reporting BLEU and WER on tokenized outputs.
3
jackychenlu
Whisper
OpenAI's general-purpose speech recognition model. Supports 99 languages, transcription, translation to English, and language identification. Six model sizes from tiny (39M params) to large (1550M params). Use for speech-to-text, podcast transcription, or multilingual audio processing. Best for robust, multilingual ASR.
0 · bundle
comeonoliver
Songsee
Generates spectrograms and multi-panel audio feature visualizations from audio files via a command-line tool.
61
microsoft
Azure Speech To Text REST Py
Transcribe short audio files (up to 60 seconds) using Azure Speech-to-Text REST API with Python, without requiring the Speech SDK.
2.7k · bundle
comeonoliver
Videocut
Generates and burns subtitles into videos: extracts audio, transcribes via Volcano Engine, corrects errors, reviews, and burns subtitles with ffmpeg.
61
affaan-m
Videodb
Ingest, index, search, edit, and generate video and audio content from files, URLs, live streams, or desktop capture.
226k · bundle