Transcribe

Transcribe audio and video files using the configured speech-to-text provider

vellum-ai 33ba830 5 files · 19.5 KB Updated

File contents

Transcribe audio and video files using the configured speech-to-text provider. Supports multiple STT providers including OpenAI Whisper, Deepgram, and Google Gemini — the active provider is selected in Settings under Speech-to-Text (services.stt).

Usage Notes

  • The tool accepts a file_path (absolute path to a local audio or video file) to transcribe.
  • Supported formats: any video (mp4, mov, etc.) or audio (mp3, wav, m4a, etc.) file.
  • For video files, audio is automatically extracted via ffmpeg before transcription.
  • Large files are automatically split into chunks for processing.
  • If no STT provider credentials are configured, the tool will return an error with setup instructions.
  • The STT provider (services.stt) is shared between transcription and telephony call paths.

Maintenance

When adding or modifying an STT provider, follow the onboarding checklist at assistant/docs/stt-provider-onboarding.md. That document covers the daemon catalog, config schema, adapter wiring, client catalog parity, and required tests.

vellum-ai/vellum-assistant/tree/main/assistant/src/config/bundled-skills/transcribe commit 33ba8302cb

Frequently asked questions

npx skillmds@latest add vellum-ai/transcribe