Audio Transcription with Sber Salute Speech
Transcribe audio/video files to text with timestamps via Salute Speech async REST API.
Requirements
- API Key: Environment variable
SALUTE_AUTH_DATA must be set (Base64-encoded client_id:client_secret or raw authorization key from https://developers.sber.ru/studio/).
- SSL note: The script disables SSL verification by default (
verify_ssl=False) because Sber's certificate chain is non-standard. This is expected.
Supported formats & encodings
| Audio encoding |
Content-Type |
Typical extensions |
MP3 |
audio/mpeg |
.mp3 |
PCM_S16LE |
audio/wav |
.wav |
OPUS |
audio/ogg |
.ogg, .opus |
FLAC |
audio/flac |
.flac |
ALAW |
audio/alaw |
.alaw |
MULAW |
audio/mulaw |
.mulaw |
Supported languages
ru-RU, en-US, kk-KZ (Kazakh), ky-KG (Kyrgyz), uz-UZ (Uzbek).
Workflow
- Identify input files — from user request.
- Read API key from host environment.
- Run transcription — execute
salute_transcribe.py with uv and appropriate arguments.
- Deliver results — present to user human-readable transcript with timestamps to the user and give a direct link to files.
Usage
uv run --with requests {baseDir}/salute_transcribe.py \
--file /path/to/audio.mp3 \
--output_dir ~/.openclaw/workspace/transcriptions \
--lang ru-RU
Arguments
| Argument |
Required |
Default |
Description |
--file |
Yes |
— |
Path to audio/video file |
--output_dir |
No |
~/.openclaw/workspace/transcribations |
Output directory for results |
--lang |
No |
ru-RU |
Language code: ru-RU, en-US, kk-KZ, ky-KG, uz-UZ |
--audio-encoding |
No |
MP3 |
Codec: MP3, PCM_S16LE, OPUS, FLAC, ALAW, MULAW |
--model |
No |
general |
Recognition model: general or callcenter |
--hyp-count |
No |
1 |
Number of alternative hypotheses: 1 or 2 |
--max-wait-time |
No |
300 |
Max seconds to wait for async result |
--print |
No |
off |
Also print transcription to stdout |
Content-Type mapping
When the file extension doesn't match audio/mpeg, adjust content_type in the script or add logic. Current default is audio/mpeg (MP3). For .wav files use audio/wav, etc.
Output files
For input file meetingABC.mp3 the script produces:
| File |
Description |
meetingABC_recognition_orig.json |
Raw API response (full JSON with all hypotheses, timing, confidence) |
meetingABC_pretty.txt |
Formatted human-readable transcript with timestamps |
Output text format
[00:01 - 00:20]:
Ну, даже если сосредоточиться на идее узкой щели.
[00:20 - 00:45]:
Следующий фрагмент текста здесь.
Notes
- Token is valid for ~30 minutes; the script fetches a new one each run.
- Large files (>1 hour) may need
--max-wait-time increased beyond 300s.
- The
callcenter model is optimized for telephony audio (8kHz, mono).
- Profanity filter is disabled by default (
enable_profanity_filter=False).
- The script uses normalized text by default (numbers as digits, abbreviations expanded). Raw text is also available in the JSON output.
1---2name: salute-speech3description: Transcribe audio files using Sber Salute Speech async API. Russian-first STT with support for ru-RU, en-US, kk-KZ, ky-KG, uz-UZ.4---5
6# Audio Transcription with Sber Salute Speech
7
8Transcribe audio/video files to text with timestamps via Salute Speech async REST API.
9
10## Requirements
11
12- **API Key**: Environment variable `SALUTE_AUTH_DATA` must be set (Base64-encoded `client_id:client_secret` or raw authorization key from https://developers.sber.ru/studio/).
13- **SSL note**: The script disables SSL verification by default (`verify_ssl=False`) because Sber's certificate chain is non-standard. This is expected.
14
15## Supported formats & encodings
16
17| Audio encoding | Content-Type | Typical extensions |
18|---------------|-------------|--------------------|
19| `MP3` | `audio/mpeg` | `.mp3` |
20| `PCM_S16LE` | `audio/wav` | `.wav` |
21| `OPUS` | `audio/ogg` | `.ogg`, `.opus` |
22| `FLAC` | `audio/flac` | `.flac` |
23| `ALAW` | `audio/alaw` | `.alaw` |
24| `MULAW` | `audio/mulaw` | `.mulaw` |
25
26## Supported languages
27
28`ru-RU`, `en-US`, `kk-KZ` (Kazakh), `ky-KG` (Kyrgyz), `uz-UZ` (Uzbek).
29
30## Workflow
31
321. **Identify input files** — from user request.
332. **Read API key** from host environment.
343. **Run transcription** — execute `salute_transcribe.py` with `uv` and appropriate arguments.
354. **Deliver results** — present to user human-readable transcript with timestamps to the user and give a direct link to files.
36
37## Usage
38
39```bash
40uv run --with requests {baseDir}/salute_transcribe.py \
41 --file /path/to/audio.mp3 \
42 --output_dir ~/.openclaw/workspace/transcriptions \
43 --lang ru-RU
44```
45
46### Arguments
47
48| Argument | Required | Default | Description |
49|----------|----------|---------|-------------|
50| `--file` | **Yes** | — | Path to audio/video file |
51| `--output_dir` | No | `~/.openclaw/workspace/transcribations` | Output directory for results |
52| `--lang` | No | `ru-RU` | Language code: `ru-RU`, `en-US`, `kk-KZ`, `ky-KG`, `uz-UZ` |
53| `--audio-encoding` | No | `MP3` | Codec: `MP3`, `PCM_S16LE`, `OPUS`, `FLAC`, `ALAW`, `MULAW` |
54| `--model` | No | `general` | Recognition model: `general` or `callcenter` |
55| `--hyp-count` | No | `1` | Number of alternative hypotheses: `1` or `2` |
56| `--max-wait-time` | No | `300` | Max seconds to wait for async result |
57| `--print` | No | off | Also print transcription to stdout |
58
59### Content-Type mapping
60
61When the file extension doesn't match `audio/mpeg`, adjust `content_type` in the script or add logic. Current default is `audio/mpeg` (MP3). For `.wav` files use `audio/wav`, etc.
62
63## Output files
64
65For input file `meetingABC.mp3` the script produces:
66
67| File | Description |
68|------|-------------|
69| `meetingABC_recognition_orig.json` | Raw API response (full JSON with all hypotheses, timing, confidence) |
70| `meetingABC_pretty.txt` | Formatted human-readable transcript with timestamps |
71
72### Output text format
73
74```
75[00:01 - 00:20]:
76Ну, даже если сосредоточиться на идее узкой щели.
77
78[00:20 - 00:45]:
79Следующий фрагмент текста здесь.
80```
81
82## Notes
83
84- Token is valid for ~30 minutes; the script fetches a new one each run.
85- Large files (>1 hour) may need `--max-wait-time` increased beyond 300s.
86- The `callcenter` model is optimized for telephony audio (8kHz, mono).
87- Profanity filter is disabled by default (`enable_profanity_filter=False`).
88- The script uses **normalized text** by default (numbers as digits, abbreviations expanded). Raw text is also available in the JSON output.