Deepgram Production Checklist
Prerequisites
- A named service/data owner, approved model/provider/data-routing decision, and production rollback plan.
- Passing staging evidence for credentials, audio retention/consent, observability, rate limits, quality, and incident escalation.
Examples
Run the final production gate with a non-sensitive canary, verify alert routing and the redacted request receipt, then observe against the stated quality/latency/error threshold. If any gate fails, halt rollout and restore the previous approved configuration rather than extending the canary or disabling controls.
Overview
Comprehensive go-live checklist for Deepgram integrations. Covers singleton client, health checks, Prometheus metrics, alert rules, error handling, and a phased go-live timeline.
Production Readiness Matrix
| Category |
Item |
Status |
| Auth |
Production API key with scoped permissions |
[ ] |
| Auth |
Key stored in secret manager (not env file) |
[ ] |
| Auth |
Key rotation schedule (90-day) configured |
[ ] |
| Auth |
Fallback key provisioned and tested |
[ ] |
| Resilience |
Retry with exponential backoff on 429/5xx |
[ ] |
| Resilience |
Circuit breaker for cascade failure prevention |
[ ] |
| Resilience |
Request timeout set (30s pre-recorded, 10s TTS) |
[ ] |
| Resilience |
Graceful degradation when API unavailable |
[ ] |
| Performance |
Singleton client (not creating per-request) |
[ ] |
| Performance |
Concurrency limited (50-80% of plan limit) |
[ ] |
| Performance |
Audio preprocessed (16kHz mono for best results) |
[ ] |
| Performance |
Large files use callback URL (async) |
[ ] |
| Monitoring |
Health check endpoint testing Deepgram API |
[ ] |
| Monitoring |
Prometheus metrics: latency, error rate, usage |
[ ] |
| Monitoring |
Alerts: error rate >5%, latency >10s, circuit open |
[ ] |
| Security |
PII redaction enabled if handling sensitive audio |
[ ] |
| Security |
Audio URLs validated (HTTPS, no private IPs) |
[ ] |
| Security |
Audit logging on all operations |
[ ] |
Instructions
Step 1: Production Singleton Client
import { createClient, DeepgramClient } from '@deepgram/sdk';
class ProductionDeepgram {
private static client: DeepgramClient | null = null;
static getClient(): DeepgramClient {
if (!this.client) {
const key = process.env.DEEPGRAM_API_KEY;
if (!key) throw new Error('DEEPGRAM_API_KEY required for production');
this.client = createClient(key);
}
return this.client;
}
// Force re-init (for key rotation)
static reset() { this.client = null; }
}
Step 2: Health Check Endpoint
import express from 'express';
import { createClient } from '@deepgram/sdk';
const app = express();
const deepgram = createClient(process.env.DEEPGRAM_API_KEY!);
app.get('/health', async (req, res) => {
const start = Date.now();
try {
// Test API connectivity by listing projects
const { error } = await deepgram.manage.getProjects();
const latency = Date.now() - start;
if (error) {
return res.status(503).json({
status: 'unhealthy',
deepgram: 'error',
error: error.message,
latency_ms: latency,
});
}
res.json({
status: 'healthy',
deepgram: 'connected',
latency_ms: latency,
timestamp: new Date().toISOString(),
});
} catch (err: any) {
res.status(503).json({
status: 'unhealthy',
deepgram: 'unreachable',
error: err.message,
latency_ms: Date.now() - start,
});
}
});
Step 3: Prometheus Metrics
import { Counter, Histogram, Gauge, Registry } from 'prom-client';
const registry = new Registry();
const transcriptionRequests = new Counter({
name: 'deepgram_requests_total',
help: 'Total Deepgram API requests',
labelNames: ['method', 'model', 'status'],
registers: [registry],
});
const transcriptionLatency = new Histogram({
name: 'deepgram_latency_seconds',
help: 'Deepgram API request latency',
labelNames: ['method', 'model'],
buckets: [0.5, 1, 2, 5, 10, 30],
registers: [registry],
});
const audioProcessed = new Counter({
name: 'deepgram_audio_seconds_total',
help: 'Total audio seconds processed',
labelNames: ['model'],
registers: [registry],
});
const activeConnections = new Gauge({
name: 'deepgram_active_connections',
help: 'Active WebSocket connections',
registers: [registry],
});
// Instrumented transcription
async function instrumentedTranscribe(url: string, model = 'nova-3') {
const timer = transcriptionLatency.startTimer({ method: 'prerecorded', model });
try {
const { result, error } = await deepgram.listen.prerecorded.transcribeUrl(
{ url }, { model, smart_format: true }
);
timer();
transcriptionRequests.inc({ method: 'prerecorded', model, status: error ? 'error' : 'ok' });
if (result?.metadata?.duration) {
audioProcessed.inc({ model }, result.metadata.duration);
}
if (error) throw error;
return result;
} catch (err) {
timer();
transcriptionRequests.inc({ method: 'prerecorded', model, status: 'error' });
throw err;
}
}
// Expose metrics endpoint
app.get('/metrics', async (req, res) => {
res.set('Content-Type', registry.contentType);
res.send(await registry.metrics());
});
Step 4: Alert Rules (Prometheus/AlertManager)
groups:
- name: deepgram
rules:
- alert: DeepgramHighErrorRate
expr: rate(deepgram_requests_total{status="error"}[5m]) / rate(deepgram_requests_total[5m]) > 0.05
for: 5m
labels:
severity: critical
annotations:
summary: "Deepgram error rate > 5%"
- alert: DeepgramHighLatency
expr: histogram_quantile(0.95, rate(deepgram_latency_seconds_bucket[5m])) > 10
for: 5m
labels:
severity: warning
annotations:
summary: "Deepgram P95 latency > 10s"
- alert: DeepgramHealthCheckFailed
expr: up{job="deepgram-service"} == 0
for: 2m
labels:
severity: critical
annotations:
summary: "Deepgram health check failed for 2+ minutes"
Step 5: Error Handling Wrapper
async function safeTranscribe(url: string, options: Record<string, any> = {}) {
const timeout = options.timeout ?? 30000;
const controller = new AbortController();
const timeoutId = setTimeout(() => controller.abort(), timeout);
try {
const result = await Promise.race([
instrumentedTranscribe(url, options.model ?? 'nova-3'),
new Promise((_, reject) =>
setTimeout(() => reject(new Error('Transcription timeout')), timeout)
),
]);
clearTimeout(timeoutId);
return result;
} catch (err: any) {
clearTimeout(timeoutId);
// Log structured error
console.error(JSON.stringify({
level: 'error',
service: 'deepgram',
message: err.message,
url: url.substring(0, 100),
timestamp: new Date().toISOString(),
}));
throw err;
}
}
Step 6: Go-Live Timeline
| Phase |
When |
Actions |
| D-7 |
1 week before |
Load test at 2x expected volume, security review |
| D-3 |
3 days before |
Smoke test with production key, verify all alerts fire |
| D-1 |
Day before |
Confirm on-call rotation, validate dashboards |
| D-0 |
Launch |
Shadow mode (10% traffic), monitoring open |
| D+1 |
Day after |
Review error rate, latency, verify no anomalies |
| D+7 |
1 week after |
Full traffic, tune alert thresholds based on baselines |
Output
- Singleton client with reset capability
- Health check endpoint with latency reporting
- Prometheus metrics (requests, latency, audio, connections)
- AlertManager rules for error rate, latency, availability
- Timeout-safe transcription wrapper
- Phased go-live timeline
Error Handling
| Issue |
Cause |
Solution |
| Health check 503 |
API key expired |
Rotate key, check secret manager |
| Metrics not scraped |
Wrong port/path |
Verify Prometheus target config |
| Alert storms |
Thresholds too tight |
Add for: duration, tune values |
| Timeout on large files |
Sync mode too slow |
Switch to callback URL pattern |
Resources
1---2name: deepgram-prod-checklist3description: Execute Deepgram production deployment checklist. Use when preparing for production launch, auditing production readiness, or verifying deployment configurations. Trigger: "deepgram production", "deploy deepgram", "deepgram prod checklist", "deepgram go-live", "production ready deepgram".4license: MIT5---6# Deepgram Production Checklist
7
8## Prerequisites
9
10- A named service/data owner, approved model/provider/data-routing decision, and production rollback plan.
11- Passing staging evidence for credentials, audio retention/consent, observability, rate limits, quality, and incident escalation.
12
13## Examples
14
15Run the final production gate with a non-sensitive canary, verify alert routing and the redacted request receipt, then observe against the stated quality/latency/error threshold. If any gate fails, halt rollout and restore the previous approved configuration rather than extending the canary or disabling controls.
16
17## Overview
18
19Comprehensive go-live checklist for Deepgram integrations. Covers singleton client, health checks, Prometheus metrics, alert rules, error handling, and a phased go-live timeline.
20
21## Production Readiness Matrix
22
23| Category | Item | Status |
24|----------|------|--------|
25| **Auth** | Production API key with scoped permissions | [ ] |
26| **Auth** | Key stored in secret manager (not env file) | [ ] |
27| **Auth** | Key rotation schedule (90-day) configured | [ ] |
28| **Auth** | Fallback key provisioned and tested | [ ] |
29| **Resilience** | Retry with exponential backoff on 429/5xx | [ ] |
30| **Resilience** | Circuit breaker for cascade failure prevention | [ ] |
31| **Resilience** | Request timeout set (30s pre-recorded, 10s TTS) | [ ] |
32| **Resilience** | Graceful degradation when API unavailable | [ ] |
33| **Performance** | Singleton client (not creating per-request) | [ ] |
34| **Performance** | Concurrency limited (50-80% of plan limit) | [ ] |
35| **Performance** | Audio preprocessed (16kHz mono for best results) | [ ] |
36| **Performance** | Large files use callback URL (async) | [ ] |
37| **Monitoring** | Health check endpoint testing Deepgram API | [ ] |
38| **Monitoring** | Prometheus metrics: latency, error rate, usage | [ ] |
39| **Monitoring** | Alerts: error rate >5%, latency >10s, circuit open | [ ] |
40| **Security** | PII redaction enabled if handling sensitive audio | [ ] |
41| **Security** | Audio URLs validated (HTTPS, no private IPs) | [ ] |
42| **Security** | Audit logging on all operations | [ ] |
43
44## Instructions
45
46### Step 1: Production Singleton Client
47
48```typescript
49import { createClient, DeepgramClient } from '@deepgram/sdk';
50
51class ProductionDeepgram {
52 private static client: DeepgramClient | null = null;
53
54 static getClient(): DeepgramClient {
55 if (!this.client) {
56 const key = process.env.DEEPGRAM_API_KEY;
57 if (!key) throw new Error('DEEPGRAM_API_KEY required for production');
58 this.client = createClient(key);
59 }
60 return this.client;
61 }
62
63 // Force re-init (for key rotation)
64 static reset() { this.client = null; }
65}
66```
67
68### Step 2: Health Check Endpoint
69
70```typescript
71import express from 'express';
72import { createClient } from '@deepgram/sdk';
73
74const app = express();
75const deepgram = createClient(process.env.DEEPGRAM_API_KEY!);
76
77app.get('/health', async (req, res) => {
78 const start = Date.now();
79 try {
80 // Test API connectivity by listing projects
81 const { error } = await deepgram.manage.getProjects();
82 const latency = Date.now() - start;
83
84 if (error) {
85 return res.status(503).json({
86 status: 'unhealthy',
87 deepgram: 'error',
88 error: error.message,
89 latency_ms: latency,
90 });
91 }
92
93 res.json({
94 status: 'healthy',
95 deepgram: 'connected',
96 latency_ms: latency,
97 timestamp: new Date().toISOString(),
98 });
99 } catch (err: any) {
100 res.status(503).json({
101 status: 'unhealthy',
102 deepgram: 'unreachable',
103 error: err.message,
104 latency_ms: Date.now() - start,
105 });
106 }
107});
108```
109
110### Step 3: Prometheus Metrics
111
112```typescript
113import { Counter, Histogram, Gauge, Registry } from 'prom-client';
114
115const registry = new Registry();
116
117const transcriptionRequests = new Counter({
118 name: 'deepgram_requests_total',
119 help: 'Total Deepgram API requests',
120 labelNames: ['method', 'model', 'status'],
121 registers: [registry],
122});
123
124const transcriptionLatency = new Histogram({
125 name: 'deepgram_latency_seconds',
126 help: 'Deepgram API request latency',
127 labelNames: ['method', 'model'],
128 buckets: [0.5, 1, 2, 5, 10, 30],
129 registers: [registry],
130});
131
132const audioProcessed = new Counter({
133 name: 'deepgram_audio_seconds_total',
134 help: 'Total audio seconds processed',
135 labelNames: ['model'],
136 registers: [registry],
137});
138
139const activeConnections = new Gauge({
140 name: 'deepgram_active_connections',
141 help: 'Active WebSocket connections',
142 registers: [registry],
143});
144
145// Instrumented transcription
146async function instrumentedTranscribe(url: string, model = 'nova-3') {
147 const timer = transcriptionLatency.startTimer({ method: 'prerecorded', model });
148 try {
149 const { result, error } = await deepgram.listen.prerecorded.transcribeUrl(
150 { url }, { model, smart_format: true }
151 );
152 timer();
153 transcriptionRequests.inc({ method: 'prerecorded', model, status: error ? 'error' : 'ok' });
154 if (result?.metadata?.duration) {
155 audioProcessed.inc({ model }, result.metadata.duration);
156 }
157 if (error) throw error;
158 return result;
159 } catch (err) {
160 timer();
161 transcriptionRequests.inc({ method: 'prerecorded', model, status: 'error' });
162 throw err;
163 }
164}
165
166// Expose metrics endpoint
167app.get('/metrics', async (req, res) => {
168 res.set('Content-Type', registry.contentType);
169 res.send(await registry.metrics());
170});
171```
172
173### Step 4: Alert Rules (Prometheus/AlertManager)
174
175```yaml
176groups:
177 - name: deepgram
178 rules:
179 - alert: DeepgramHighErrorRate
180 expr: rate(deepgram_requests_total{status="error"}[5m]) / rate(deepgram_requests_total[5m]) > 0.05
181 for: 5m
182 labels:
183 severity: critical
184 annotations:
185 summary: "Deepgram error rate > 5%"
186
187 - alert: DeepgramHighLatency
188 expr: histogram_quantile(0.95, rate(deepgram_latency_seconds_bucket[5m])) > 10
189 for: 5m
190 labels:
191 severity: warning
192 annotations:
193 summary: "Deepgram P95 latency > 10s"
194
195 - alert: DeepgramHealthCheckFailed
196 expr: up{job="deepgram-service"} == 0
197 for: 2m
198 labels:
199 severity: critical
200 annotations:
201 summary: "Deepgram health check failed for 2+ minutes"
202```
203
204### Step 5: Error Handling Wrapper
205
206```typescript
207async function safeTranscribe(url: string, options: Record<string, any> = {}) {
208 const timeout = options.timeout ?? 30000;
209
210 const controller = new AbortController();
211 const timeoutId = setTimeout(() => controller.abort(), timeout);
212
213 try {
214 const result = await Promise.race([
215 instrumentedTranscribe(url, options.model ?? 'nova-3'),
216 new Promise((_, reject) =>
217 setTimeout(() => reject(new Error('Transcription timeout')), timeout)
218 ),
219 ]);
220 clearTimeout(timeoutId);
221 return result;
222 } catch (err: any) {
223 clearTimeout(timeoutId);
224 // Log structured error
225 console.error(JSON.stringify({
226 level: 'error',
227 service: 'deepgram',
228 message: err.message,
229 url: url.substring(0, 100),
230 timestamp: new Date().toISOString(),
231 }));
232 throw err;
233 }
234}
235```
236
237### Step 6: Go-Live Timeline
238
239| Phase | When | Actions |
240|-------|------|---------|
241| D-7 | 1 week before | Load test at 2x expected volume, security review |
242| D-3 | 3 days before | Smoke test with production key, verify all alerts fire |
243| D-1 | Day before | Confirm on-call rotation, validate dashboards |
244| D-0 | Launch | Shadow mode (10% traffic), monitoring open |
245| D+1 | Day after | Review error rate, latency, verify no anomalies |
246| D+7 | 1 week after | Full traffic, tune alert thresholds based on baselines |
247
248## Output
249
250- Singleton client with reset capability
251- Health check endpoint with latency reporting
252- Prometheus metrics (requests, latency, audio, connections)
253- AlertManager rules for error rate, latency, availability
254- Timeout-safe transcription wrapper
255- Phased go-live timeline
256
257## Error Handling
258
259| Issue | Cause | Solution |
260|-------|-------|----------|
261| Health check 503 | API key expired | Rotate key, check secret manager |
262| Metrics not scraped | Wrong port/path | Verify Prometheus target config |
263| Alert storms | Thresholds too tight | Add `for:` duration, tune values |
264| Timeout on large files | Sync mode too slow | Switch to `callback` URL pattern |
265
266## Resources
267
268- Deepgram Production Guide
269- [Prometheus Best Practices](https://prometheus.io/docs/practices/)
270- Deepgram SLA