Transcribe Anything

Transcribes audio and video files to text using pluggable ASR backends. Default backend is local whisper/MLX Whisper for ASR. Supports pyannote.audio direct diarization, whisperX, insanely-fast-whisper, faster-whisper, whisper.cpp, OpenAI Whisper API, Groq Whisper API, Deepgram, AssemblyAI, Gemini, and Hugging Face models. Handles very long files (1-8+ hours) by preprocessing with ffmpeg: extracts audio from video, converts to optimal ASR format, detects and skips silence, and chunks for API size limits. Supports speaker diarization, word-level timestamps, custom vocabulary, and multiple output formats. Use this skill when someone says "transcribe this", "convert to text", "speech to text", "get the transcript", "transcribe this video/audio/podcast/recording", or provides a media file and wants text output.

swyxio Updated

File contents

swyxio/skills/tree/main/transcribe-anything commit b5bb3a3a2a

Frequently asked questions

npx skillmds@latest add swyxio/transcribe-anything