Transcribir

Transcribe one or more audio or video files to a structured Markdown document with timestamps and speaker diarization, using Google Gemini 3.5 Flash. Use this skill whenever the user provides audio/video files and wants them transcribed, or when the user says "transcribe esto", "transcribe estos audios", "pasame esto a texto", "que dice este audio", "transcripción", "voice note", "audio nota", "WhatsApp PTT", or similar. Supports .ogg, .mp3, .m4a, .wav, .opus, .flac, .aac, .webm, .mp4, .mov, .mpeg, .mpga. Auto-detects language (Spanish, English, etc.), identifies multiple speakers with timestamps, and outputs a clean MD next to each source file. Processes multiple files in parallel. Defaults to Gemini 3.5 Flash for the best speed/quality/cost balance (handles Spanish, Latin American accents, regionalisms, and multi-speaker diarization). Use the --pro flag to upgrade to gemini-pro-latest for maximum quality on critical audio, or --model X to specify any other model. Use when this capability is needed.

tomevault-io c216b88 2 files · 7.0 KB Updated

File contents

tomevault-io/skills-registry/tree/main/josuebustosn--gemini-transcribe--transcribir commit c216b88aa9

Frequently asked questions

npx skillmds@latest add tomevault-io/transcribir