Openai Whisper API

Transcribe audio files via OpenAI Audio Transcriptions API (Whisper). Supports multiple formats with language detection and custom prompts.

yunseo-kim Updated

File contents

OpenAI Whisper API (curl)

Transcribe an audio file via OpenAI’s /v1/audio/transcriptions endpoint.

Quick start

{skillDir}/scripts/transcribe.sh /path/to/audio.m4a

Defaults:

  • Model: whisper-1
  • Output: <input>.txt

Useful flags

{skillDir}/scripts/transcribe.sh /path/to/audio.ogg --model whisper-1 --out /tmp/transcript.txt
{skillDir}/scripts/transcribe.sh /path/to/audio.m4a --language en
{skillDir}/scripts/transcribe.sh /path/to/audio.m4a --prompt "Speaker names: Peter, Daniel"
{skillDir}/scripts/transcribe.sh /path/to/audio.m4a --json --out /tmp/transcript.json

API key

Set OPENAI_API_KEY, or configure it in ~/.the assistant/the assistant.json:

{
  skills: {
    "openai-whisper-api": {
      apiKey: "OPENAI_KEY_HERE",
    },
  },
}

yunseo-kim/agent-toolbox/tree/main/catalog/skills/openai-whisper-api commit 4d1b1fdcca

Frequently asked questions

npx skillmds@latest add yunseo-kim/openai-whisper-api