# azure-ai-transcription-py

> Transcribe audio to text using Azure AI Transcription SDK with real-time and batch support, including timestamps and diarization.

- Skill: `microsoft/azure-ai-transcription-py` (Agent Skill)
- Install (CLI): `npx skillmds add microsoft/azure-ai-transcription-py`
- Raw SKILL.md: https://api.skillmd.com/api/skills/microsoft/azure-ai-transcription-py/raw
- Safety review: CAUTION (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: AI & ML, DevOps & Infra, Speech & Audio
- Tags: Azure, Azure Sdk, Batch Transcription, Python, Real Time Transcription, Speech To Text, Transcription
- License: MIT
- Author: Microsoft (https://skillmd.com/u/microsoft), verified publisher
- Updated: 2026-07-06
- Page: https://skillmd.com/skills/microsoft/azure-ai-transcription-py

---


# Azure AI Transcription SDK for Python

Client library for Azure AI Transcription (speech-to-text) with real-time and batch transcription.

## Installation

```bash
pip install azure-ai-transcription
```

## Environment Variables

```bash
TRANSCRIPTION_ENDPOINT=https://<resource>.cognitiveservices.azure.com
TRANSCRIPTION_KEY=<your-key>
```

## Authentication

Use subscription key authentication (DefaultAzureCredential is not supported for this client):

```python
import os
from azure.ai.transcription import TranscriptionClient

with TranscriptionClient(
    endpoint=os.environ["TRANSCRIPTION_ENDPOINT"],
    credential=os.environ["TRANSCRIPTION_KEY"],
) as client:
    transcriptions = list(client.list_transcriptions())
```

## Transcription (Batch)

```python
import os
from azure.ai.transcription import TranscriptionClient

with TranscriptionClient(
    endpoint=os.environ["TRANSCRIPTION_ENDPOINT"],
    credential=os.environ["TRANSCRIPTION_KEY"],
) as client:
    job = client.begin_transcription(
        name="meeting-transcription",
        locale="en-US",
        content_urls=["https://<storage>/audio.wav"],
        diarization_enabled=True,
    )
    result = job.result()
    print(result.status)
```

## Transcription (Real-time)

```python
import os
from azure.ai.transcription import TranscriptionClient

with TranscriptionClient(
    endpoint=os.environ["TRANSCRIPTION_ENDPOINT"],
    credential=os.environ["TRANSCRIPTION_KEY"],
) as client:
    stream = client.begin_stream_transcription(locale="en-US")
    stream.send_audio_file("audio.wav")
    for event in stream:
        print(event.text)
```

## Best Practices

1. **Pick sync OR async and stay consistent.** Do not mix `azure.xxx` sync clients with `azure.xxx.aio` async clients in the same call path. Choose one mode per module.
2. **Always use context managers for clients and async credentials.** Wrap every client in `with Client(...) as client:` (sync) or `async with Client(...) as client:` (async). For async `DefaultAzureCredential` from `azure.identity.aio`, also use `async with credential:` so tokens and transports are cleaned up.
3. **Enable diarization** when multiple speakers are present
4. **Use batch transcription** for long files stored in blob storage
5. **Capture timestamps** for subtitle generation
6. **Specify language** to improve recognition accuracy
7. **Handle streaming backpressure** for real-time transcription
8. **Close transcription sessions** when complete

