Logging & Observability
Patterns for building observable systems across the three pillars: logs, metrics, and traces.
Three Pillars
| Pillar |
Purpose |
Question It Answers |
Example |
| Logs |
What happened |
Why did this request fail? |
{"level":"error","msg":"payment declined","user_id":"u_82"} |
| Metrics |
How much / how fast |
Is latency increasing? |
http_request_duration_seconds{route="/api/orders"} 0.342 |
| Traces |
Request flow |
Where is the bottleneck? |
Span: api-gateway → auth → order-service → db |
Each pillar is strongest when correlated. Embed trace_id in every log line to jump from a log entry to the full distributed trace.
Structured Logging
Always emit logs as structured JSON — never free-text strings.
Required Fields
| Field |
Purpose |
Required |
timestamp |
ISO-8601 with milliseconds |
Yes |
level |
Severity (DEBUG … FATAL) |
Yes |
service |
Originating service name |
Yes |
message |
Human-readable description |
Yes |
trace_id |
Distributed trace correlation |
Yes |
span_id |
Current span within trace |
Yes |
correlation_id |
Business-level correlation (order ID) |
When applicable |
error |
Structured error object |
On errors |
context |
Request-specific metadata |
Recommended |
Context Enrichment
Attach context at the middleware level so downstream logs inherit automatically:
app.use((req, res, next) => {
const ctx = {
trace_id: req.headers['x-trace-id'] || crypto.randomUUID(),
request_id: crypto.randomUUID(),
user_id: req.user?.id,
method: req.method,
path: req.path,
};
asyncLocalStorage.run(ctx, () => next());
});
Library Recommendations
| Library |
Language |
Strengths |
Perf |
| Pino |
Node.js |
Fastest Node logger, low overhead |
Excellent |
| structlog |
Python |
Composable processors, context binding |
Good |
| zerolog |
Go |
Zero-allocation JSON logging |
Excellent |
| zap |
Go |
High performance, typed fields |
Excellent |
| tracing |
Rust |
Spans + events, async-aware |
Excellent |
Choose a logger that outputs structured JSON natively. Avoid loggers requiring post-processing.
Log Levels
| Level |
When to Use |
Example |
| FATAL |
App cannot continue, process will exit |
Database connection pool exhausted |
| ERROR |
Operation failed, needs attention |
Payment charge failed: CARD_DECLINED |
| WARN |
Unexpected but recoverable |
Retry 2/3 for upstream timeout |
| INFO |
Normal business events |
Order ORD-1234 placed successfully |
| DEBUG |
Developer troubleshooting |
Cache miss for key user:82:preferences |
| TRACE |
Very fine-grained (rarely in prod) |
Entering validateAddress with payload |
Rules: Production default = INFO and above. If you log an ERROR, someone should act on it. Every FATAL should trigger an alert.
Distributed Tracing
OpenTelemetry Setup
Always prefer OpenTelemetry over vendor-specific SDKs:
import { NodeSDK } from '@opentelemetry/sdk-node';
import { OTLPTraceExporter } from '@opentelemetry/exporter-trace-otlp-http';
import { getNodeAutoInstrumentations } from '@opentelemetry/auto-instrumentations-node';
const sdk = new NodeSDK({
serviceName: 'order-service',
traceExporter: new OTLPTraceExporter({
url: 'http://otel-collector:4318/v1/traces',
}),
instrumentations: [getNodeAutoInstrumentations()],
});
sdk.start();
Span Creation
const tracer = trace.getTracer('order-service');
async function processOrder(order: Order) {
return tracer.startActiveSpan('processOrder', async (span) => {
try {
span.setAttribute('order.id', order.id);
span.setAttribute('order.total_cents', order.totalCents);
await validateInventory(order);
await chargePayment(order);
span.setStatus({ code: SpanStatusCode.OK });
} catch (err) {
span.setStatus({ code: SpanStatusCode.ERROR, message: err.message });
span.recordException(err);
throw err;
} finally {
span.end();
}
});
}
Context Propagation
- Use W3C Trace Context (
traceparent header) — default in OTel
- Propagate across HTTP, gRPC, and message queues
- For async workers: serialise
traceparent into the job payload
Trace Sampling
| Strategy |
Use When |
| Always On |
Low-traffic services, debugging |
| Probabilistic (N%) |
General production use |
| Rate-limited (N/sec) |
High-throughput services |
| Tail-based |
When you need all error traces |
Always sample 100% of error traces regardless of strategy.
Metrics Collection
RED Method (Request-Driven)
Monitor these three for every service endpoint:
| Metric |
What It Measures |
Prometheus Example |
| Rate |
Requests/sec |
rate(http_requests_total[5m]) |
| Errors |
Failed request ratio |
rate(http_requests_total{status=~"5.."}[5m]) |
| Duration |
Response time |
histogram_quantile(0.99, http_request_duration_seconds) |
USE Method (Resource-Driven)
For infrastructure components (CPU, memory, disk, network):
| Metric |
What It Measures |
Example |
| Utilization |
% resource busy |
CPU usage at 78% |
| Saturation |
Work queued/waiting |
12 requests queued in thread pool |
| Errors |
Error events on resource |
3 disk I/O errors in last minute |
Monitoring Stack
| Tool |
Category |
Best For |
| Prometheus |
Metrics |
Pull-based metrics, alerting rules |
| Grafana |
Visualisation |
Dashboards for metrics, logs, traces |
| Jaeger |
Tracing |
Distributed trace visualisation |
| Loki |
Logs |
Log aggregation (pairs with Grafana) |
| OpenTelemetry |
Collection |
Vendor-neutral telemetry collection |
Recommendation: Start with OTel Collector → Prometheus + Grafana + Loki + Jaeger. Migrate to SaaS only when operational overhead justifies cost.
Alert Design
Severity Levels
| Severity |
Response Time |
Example |
| P1 |
Immediate |
Service fully down, data loss |
| P2 |
< 30 min |
Error rate > 5%, latency p99 > 5s |
| P3 |
Business hours |
Disk > 80%, cert expiring in 7 days |
| P4 |
Best effort |
Non-critical deprecation warning |
Alert Fatigue Prevention
- Alert on symptoms, not causes — "error rate > 5%" not "pod restarted"
- Multi-window, multi-burn-rate — catch both sudden spikes and slow burns
- Require runbook links — every alert must link to diagnosis and remediation
- Review monthly — delete or tune alerts that never fire or always fire
- Group related alerts — use inhibition rules to suppress child alerts
- Set appropriate thresholds — if alert fires daily and is ignored, raise threshold or delete
Dashboard Patterns
Overview Dashboard ("War Room")
- Total requests/sec across all services
- Global error rate (%) with trendline
- p50 / p95 / p99 latency
- Active alerts count by severity
- Deployment markers overlaid on graphs
Service Dashboard (Per-Service)
- RED metrics for each endpoint
- Dependency health (upstream/downstream success rates)
- Resource utilisation (CPU, memory, connections)
- Top errors table with count and last seen
Observability Checklist
Every service must have:
Anti-Patterns
| Anti-Pattern |
Problem |
Fix |
| Logging PII |
Privacy/compliance violation |
Mask or exclude PII; use token references |
| Excessive logging |
Storage costs balloon, signal drowns |
Log business events, not data flow |
| Unstructured logs |
Cannot query or alert on fields |
Use structured JSON with consistent schema |
| String interpolation |
Breaks structured fields, injection risk |
Pass fields as metadata, not in message |
| Missing correlation IDs |
Cannot trace across services |
Generate and propagate trace_id everywhere |
| Alert storms |
On-call fatigue, real issues buried |
Use grouping, inhibition, deduplication |
| Metrics with high cardinality |
Prometheus OOM, dashboard timeouts |
Never use user ID or request ID as label |
NEVER Do
- NEVER log passwords, tokens, API keys, or secrets — even at DEBUG level
- NEVER use console.log / print in production — use a structured logger
- NEVER use user IDs, emails, or request IDs as metric labels — cardinality will explode
- NEVER create alerts without a runbook link — unactionable alerts erode trust
- NEVER rely on logs alone — you need metrics and traces for full observability
- NEVER log request/response bodies by default — opt-in only, with PII redaction
- NEVER ignore log volume — set budgets and alert when a service exceeds daily quota
- NEVER skip context propagation in async flows — broken traces are worse than no traces
1---2name: logging-observability3description: Structured logging, distributed tracing, and metrics collection patterns for building observable systems. Use when implementing logging infrastructure, setting up distributed tracing with OpenTelemetry, designing metrics collection (RED/USE methods), configuring alerting and dashboards, or reviewing observability practices. Covers structured JSON logging, context propagation, trace sampling, Prometheus/Grafana stack, alert design, and PII/secret scrubbing.4---5
6# Logging & Observability
7
8Patterns for building observable systems across the three pillars: logs, metrics, and traces.
9
10## Three Pillars
11
12| Pillar | Purpose | Question It Answers | Example |
13|--------|---------|---------------------|---------|
14| **Logs** | What happened | Why did this request fail? | `{"level":"error","msg":"payment declined","user_id":"u_82"}` |
15| **Metrics** | How much / how fast | Is latency increasing? | `http_request_duration_seconds{route="/api/orders"} 0.342` |
16| **Traces** | Request flow | Where is the bottleneck? | Span: `api-gateway → auth → order-service → db` |
17
18Each pillar is strongest when correlated. Embed `trace_id` in every log line to jump from a log entry to the full distributed trace.
19
20---
21
22## Structured Logging
23
24Always emit logs as structured JSON — never free-text strings.
25
26### Required Fields
27
28| Field | Purpose | Required |
29|-------|---------|----------|
30| `timestamp` | ISO-8601 with milliseconds | Yes |
31| `level` | Severity (DEBUG … FATAL) | Yes |
32| `service` | Originating service name | Yes |
33| `message` | Human-readable description | Yes |
34| `trace_id` | Distributed trace correlation | Yes |
35| `span_id` | Current span within trace | Yes |
36| `correlation_id` | Business-level correlation (order ID) | When applicable |
37| `error` | Structured error object | On errors |
38| `context` | Request-specific metadata | Recommended |
39
40### Context Enrichment
41
42Attach context at the middleware level so downstream logs inherit automatically:
43
44```typescript
45app.use((req, res, next) => {
46 const ctx = {
47 trace_id: req.headers['x-trace-id'] || crypto.randomUUID(),
48 request_id: crypto.randomUUID(),
49 user_id: req.user?.id,
50 method: req.method,
51 path: req.path,
52 };
53 asyncLocalStorage.run(ctx, () => next());
54});
55```
56
57### Library Recommendations
58
59| Library | Language | Strengths | Perf |
60|---------|----------|-----------|------|
61| **Pino** | Node.js | Fastest Node logger, low overhead | Excellent |
62| **structlog** | Python | Composable processors, context binding | Good |
63| **zerolog** | Go | Zero-allocation JSON logging | Excellent |
64| **zap** | Go | High performance, typed fields | Excellent |
65| **tracing** | Rust | Spans + events, async-aware | Excellent |
66
67Choose a logger that outputs structured JSON natively. Avoid loggers requiring post-processing.
68
69---
70
71## Log Levels
72
73| Level | When to Use | Example |
74|-------|-------------|---------|
75| **FATAL** | App cannot continue, process will exit | Database connection pool exhausted |
76| **ERROR** | Operation failed, needs attention | Payment charge failed: CARD_DECLINED |
77| **WARN** | Unexpected but recoverable | Retry 2/3 for upstream timeout |
78| **INFO** | Normal business events | Order ORD-1234 placed successfully |
79| **DEBUG** | Developer troubleshooting | Cache miss for key user:82:preferences |
80| **TRACE** | Very fine-grained (rarely in prod) | Entering validateAddress with payload |
81
82**Rules:** Production default = INFO and above. If you log an ERROR, someone should act on it. Every FATAL should trigger an alert.
83
84---
85
86## Distributed Tracing
87
88### OpenTelemetry Setup
89
90Always prefer OpenTelemetry over vendor-specific SDKs:
91
92```typescript
93import { NodeSDK } from '@opentelemetry/sdk-node';
94import { OTLPTraceExporter } from '@opentelemetry/exporter-trace-otlp-http';
95import { getNodeAutoInstrumentations } from '@opentelemetry/auto-instrumentations-node';
96
97const sdk = new NodeSDK({
98 serviceName: 'order-service',
99 traceExporter: new OTLPTraceExporter({
100 url: 'http://otel-collector:4318/v1/traces',
101 }),
102 instrumentations: [getNodeAutoInstrumentations()],
103});
104sdk.start();
105```
106
107### Span Creation
108
109```typescript
110const tracer = trace.getTracer('order-service');
111
112async function processOrder(order: Order) {
113 return tracer.startActiveSpan('processOrder', async (span) => {
114 try {
115 span.setAttribute('order.id', order.id);
116 span.setAttribute('order.total_cents', order.totalCents);
117 await validateInventory(order);
118 await chargePayment(order);
119 span.setStatus({ code: SpanStatusCode.OK });
120 } catch (err) {
121 span.setStatus({ code: SpanStatusCode.ERROR, message: err.message });
122 span.recordException(err);
123 throw err;
124 } finally {
125 span.end();
126 }
127 });
128}
129```
130
131### Context Propagation
132
133- Use W3C Trace Context (`traceparent` header) — default in OTel
134- Propagate across HTTP, gRPC, and message queues
135- For async workers: serialise `traceparent` into the job payload
136
137### Trace Sampling
138
139| Strategy | Use When |
140|----------|----------|
141| **Always On** | Low-traffic services, debugging |
142| **Probabilistic** (N%) | General production use |
143| **Rate-limited** (N/sec) | High-throughput services |
144| **Tail-based** | When you need all error traces |
145
146Always sample 100% of error traces regardless of strategy.
147
148---
149
150## Metrics Collection
151
152### RED Method (Request-Driven)
153
154Monitor these three for every service endpoint:
155
156| Metric | What It Measures | Prometheus Example |
157|--------|-----------------|-------------------|
158| **Rate** | Requests/sec | `rate(http_requests_total[5m])` |
159| **Errors** | Failed request ratio | `rate(http_requests_total{status=~"5.."}[5m])` |
160| **Duration** | Response time | `histogram_quantile(0.99, http_request_duration_seconds)` |
161
162### USE Method (Resource-Driven)
163
164For infrastructure components (CPU, memory, disk, network):
165
166| Metric | What It Measures | Example |
167|--------|-----------------|---------|
168| **Utilization** | % resource busy | CPU usage at 78% |
169| **Saturation** | Work queued/waiting | 12 requests queued in thread pool |
170| **Errors** | Error events on resource | 3 disk I/O errors in last minute |
171
172---
173
174## Monitoring Stack
175
176| Tool | Category | Best For |
177|------|----------|----------|
178| **Prometheus** | Metrics | Pull-based metrics, alerting rules |
179| **Grafana** | Visualisation | Dashboards for metrics, logs, traces |
180| **Jaeger** | Tracing | Distributed trace visualisation |
181| **Loki** | Logs | Log aggregation (pairs with Grafana) |
182| **OpenTelemetry** | Collection | Vendor-neutral telemetry collection |
183
184**Recommendation:** Start with OTel Collector → Prometheus + Grafana + Loki + Jaeger. Migrate to SaaS only when operational overhead justifies cost.
185
186---
187
188## Alert Design
189
190### Severity Levels
191
192| Severity | Response Time | Example |
193|----------|---------------|---------|
194| **P1** | Immediate | Service fully down, data loss |
195| **P2** | < 30 min | Error rate > 5%, latency p99 > 5s |
196| **P3** | Business hours | Disk > 80%, cert expiring in 7 days |
197| **P4** | Best effort | Non-critical deprecation warning |
198
199### Alert Fatigue Prevention
200
201- **Alert on symptoms, not causes** — "error rate > 5%" not "pod restarted"
202- **Multi-window, multi-burn-rate** — catch both sudden spikes and slow burns
203- **Require runbook links** — every alert must link to diagnosis and remediation
204- **Review monthly** — delete or tune alerts that never fire or always fire
205- **Group related alerts** — use inhibition rules to suppress child alerts
206- **Set appropriate thresholds** — if alert fires daily and is ignored, raise threshold or delete
207
208---
209
210## Dashboard Patterns
211
212### Overview Dashboard ("War Room")
213- Total requests/sec across all services
214- Global error rate (%) with trendline
215- p50 / p95 / p99 latency
216- Active alerts count by severity
217- Deployment markers overlaid on graphs
218
219### Service Dashboard (Per-Service)
220- RED metrics for each endpoint
221- Dependency health (upstream/downstream success rates)
222- Resource utilisation (CPU, memory, connections)
223- Top errors table with count and last seen
224
225---
226
227## Observability Checklist
228
229Every service must have:
230
231- [ ] Structured JSON logging with consistent schema
232- [ ] Correlation / trace IDs propagated on all requests
233- [ ] RED metrics exposed for every external endpoint
234- [ ] Health check endpoints (`/healthz` and `/readyz`)
235- [ ] Distributed tracing with OpenTelemetry
236- [ ] Dashboards for RED metrics and resource utilisation
237- [ ] Alerts for error rate, latency, and saturation with runbook links
238- [ ] Log level configurable at runtime without redeployment
239- [ ] PII scrubbing verified and tested
240- [ ] Retention policies defined for logs, metrics, and traces
241
242## Anti-Patterns
243
244| Anti-Pattern | Problem | Fix |
245|-------------|---------|-----|
246| Logging PII | Privacy/compliance violation | Mask or exclude PII; use token references |
247| Excessive logging | Storage costs balloon, signal drowns | Log business events, not data flow |
248| Unstructured logs | Cannot query or alert on fields | Use structured JSON with consistent schema |
249| String interpolation | Breaks structured fields, injection risk | Pass fields as metadata, not in message |
250| Missing correlation IDs | Cannot trace across services | Generate and propagate trace_id everywhere |
251| Alert storms | On-call fatigue, real issues buried | Use grouping, inhibition, deduplication |
252| Metrics with high cardinality | Prometheus OOM, dashboard timeouts | Never use user ID or request ID as label |
253
254## NEVER Do
255
2561. **NEVER log passwords, tokens, API keys, or secrets** — even at DEBUG level
2572. **NEVER use console.log / print in production** — use a structured logger
2583. **NEVER use user IDs, emails, or request IDs as metric labels** — cardinality will explode
2594. **NEVER create alerts without a runbook link** — unactionable alerts erode trust
2605. **NEVER rely on logs alone** — you need metrics and traces for full observability
2616. **NEVER log request/response bodies by default** — opt-in only, with PII redaction
2627. **NEVER ignore log volume** — set budgets and alert when a service exceeds daily quota
2638. **NEVER skip context propagation in async flows** — broken traces are worse than no traces