WhisperX Speech Recognition with Word-Level Timestamps and Diarization

WhisperX extends OpenAI Whisper with batched inference for 70x realtime transcription, phoneme-based word-level timestamp alignment via wav2vec2, voice activity detection, and speaker diarization. It produces accurate per-word timestamps and speaker labels from audio files.

agentskillexchange Updated 28 repo stars

File contents

agentskillexchange/skills/tree/main/skills/whisperx-speech-recognition-timestamps-diarization commit be1e7f6562

Frequently asked questions

npx skillmds@latest add agentskillexchange/whisperx-speech-recognition-with-word-level-timestamps-and-d