Logging & Observability
Patterns for building observable systems across the three pillars: logs, metrics, and traces.
Three Pillars
| Pillar |
Purpose |
Question It Answers |
Example |
| Logs |
What happened |
Why did this request fail? |
{"level":"error","msg":"payment declined","user_id":"u_82"} |
| Metrics |
How much / how fast |
Is latency increasing? |
http_request_duration_seconds{route="/api/orders"} 0.342 |
| Traces |
Request flow |
Where is the bottleneck? |
Span: api-gateway → auth → order-service → db |
Each pillar is strongest when correlated. Embed trace_id in every log line to jump from a log entry to the full distributed trace.
Structured Logging
Always emit logs as structured JSON — never free-text strings.
Required Fields
| Field |
Purpose |
Required |
timestamp |
ISO-8601 with milliseconds |
Yes |
level |
Severity (DEBUG … FATAL) |
Yes |
service |
Originating service name |
Yes |
message |
Human-readable description |
Yes |
trace_id |
Distributed trace correlation |
Yes |
span_id |
Current span within trace |
Yes |
correlation_id |
Business-level correlation (order ID) |
When applicable |
error |
Structured error object |
On errors |
context |
Request-specific metadata |
Recommended |
Context Enrichment
Attach context at the middleware level so downstream logs inherit automatically:
app.use((req, res, next) => {
const ctx = {
trace_id: req.headers['x-trace-id'] || crypto.randomUUID(),
request_id: crypto.randomUUID(),
user_id: req.user?.id,
method: req.method,
path: req.path,
};
asyncLocalStorage.run(ctx, () => next());
});
Library Recommendations
| Library |
Language |
Strengths |
Perf |
| Pino |
Node.js |
Fastest Node logger, low overhead |
Excellent |
| structlog |
Python |
Composable processors, context binding |
Good |
| zerolog |
Go |
Zero-allocation JSON logging |
Excellent |
| zap |
Go |
High performance, typed fields |
Excellent |
| tracing |
Rust |
Spans + events, async-aware |
Excellent |
Choose a logger that outputs structured JSON natively. Avoid loggers requiring post-processing.
Log Levels
| Level |
When to Use |
Example |
| FATAL |
App cannot continue, process will exit |
Database connection pool exhausted |
| ERROR |
Operation failed, needs attention |
Payment charge failed: CARD_DECLINED |
| WARN |
Unexpected but recoverable |
Retry 2/3 for upstream timeout |
| INFO |
Normal business events |
Order ORD-1234 placed successfully |
| DEBUG |
Developer troubleshooting |
Cache miss for key user:82:preferences |
| TRACE |
Very fine-grained (rarely in prod) |
Entering validateAddress with payload |
Rules: Production default = INFO and above. If you log an ERROR, someone should act on it. Every FATAL should trigger an alert.
Distributed Tracing
OpenTelemetry Setup
Always prefer OpenTelemetry over vendor-specific SDKs:
import { NodeSDK } from '@opentelemetry/sdk-node';
import { OTLPTraceExporter } from '@opentelemetry/exporter-trace-otlp-http';
import { getNodeAutoInstrumentations } from '@opentelemetry/auto-instrumentations-node';
const sdk = new NodeSDK({
serviceName: 'order-service',
traceExporter: new OTLPTraceExporter({
url: 'http://otel-collector:4318/v1/traces',
}),
instrumentations: [getNodeAutoInstrumentations()],
});
sdk.start();
Span Creation
const tracer = trace.getTracer('order-service');
async function processOrder(order: Order) {
return tracer.startActiveSpan('processOrder', async (span) => {
try {
span.setAttribute('order.id', order.id);
span.setAttribute('order.total_cents', order.totalCents);
await validateInventory(order);
await chargePayment(order);
span.setStatus({ code: SpanStatusCode.OK });
} catch (err) {
span.setStatus({ code: SpanStatusCode.ERROR, message: err.message });
span.recordException(err);
throw err;
} finally {
span.end();
}
});
}
Context Propagation
- Use W3C Trace Context (
traceparent header) — default in OTel
- Propagate across HTTP, gRPC, and message queues
- For async workers: serialise
traceparent into the job payload
Trace Sampling
| Strategy |
Use When |
| Always On |
Low-traffic services, debugging |
| Probabilistic (N%) |
General production use |
| Rate-limited (N/sec) |
High-throughput services |
| Tail-based |
When you need all error traces |
Always sample 100% of error traces regardless of strategy.
Metrics Collection
RED Method (Request-Driven)
Monitor these three for every service endpoint:
| Metric |
What It Measures |
Prometheus Example |
| Rate |
Requests/sec |
rate(http_requests_total[5m]) |
| Errors |
Failed request ratio |
rate(http_requests_total{status=~"5.."}[5m]) |
| Duration |
Response time |
histogram_quantile(0.99, http_request_duration_seconds) |
USE Method (Resource-Driven)
For infrastructure components (CPU, memory, disk, network):
| Metric |
What It Measures |
Example |
| Utilization |
% resource busy |
CPU usage at 78% |
| Saturation |
Work queued/waiting |
12 requests queued in thread pool |
| Errors |
Error events on resource |
3 disk I/O errors in last minute |
Monitoring Stack
| Tool |
Category |
Best For |
| Prometheus |
Metrics |
Pull-based metrics, alerting rules |
| Grafana |
Visualisation |
Dashboards for metrics, logs, traces |
| Jaeger |
Tracing |
Distributed trace visualisation |
| Loki |
Logs |
Log aggregation (pairs with Grafana) |
| OpenTelemetry |
Collection |
Vendor-neutral telemetry collection |
Recommendation: Start with OTel Collector → Prometheus + Grafana + Loki + Jaeger. Migrate to SaaS only when operational overhead justifies cost.
Alert Design
Severity Levels
| Severity |
Response Time |
Example |
| P1 |
Immediate |
Service fully down, data loss |
| P2 |
< 30 min |
Error rate > 5%, latency p99 > 5s |
| P3 |
Business hours |
Disk > 80%, cert expiring in 7 days |
| P4 |
Best effort |
Non-critical deprecation warning |
Alert Fatigue Prevention
- Alert on symptoms, not causes — "error rate > 5%" not "pod restarted"
- Multi-window, multi-burn-rate — catch both sudden spikes and slow burns
- Require runbook links — every alert must link to diagnosis and remediation
- Review monthly — delete or tune alerts that never fire or always fire
- Group related alerts — use inhibition rules to suppress child alerts
- Set appropriate thresholds — if alert fires daily and is ignored, raise threshold or delete
Dashboard Patterns
Overview Dashboard ("War Room")
- Total requests/sec across all services
- Global error rate (%) with trendline
- p50 / p95 / p99 latency
- Active alerts count by severity
- Deployment markers overlaid on graphs
Service Dashboard (Per-Service)
- RED metrics for each endpoint
- Dependency health (upstream/downstream success rates)
- Resource utilisation (CPU, memory, connections)
- Top errors table with count and last seen
Observability Checklist
Every service must have:
Anti-Patterns
| Anti-Pattern |
Problem |
Fix |
| Logging PII |
Privacy/compliance violation |
Mask or exclude PII; use token references |
| Excessive logging |
Storage costs balloon, signal drowns |
Log business events, not data flow |
| Unstructured logs |
Cannot query or alert on fields |
Use structured JSON with consistent schema |
| String interpolation |
Breaks structured fields, injection risk |
Pass fields as metadata, not in message |
| Missing correlation IDs |
Cannot trace across services |
Generate and propagate trace_id everywhere |
| Alert storms |
On-call fatigue, real issues buried |
Use grouping, inhibition, deduplication |
| Metrics with high cardinality |
Prometheus OOM, dashboard timeouts |
Never use user ID or request ID as label |
NEVER Do
- NEVER log passwords, tokens, API keys, or secrets — even at DEBUG level
- NEVER use console.log / print in production — use a structured logger
- NEVER use user IDs, emails, or request IDs as metric labels — cardinality will explode
- NEVER create alerts without a runbook link — unactionable alerts erode trust
- NEVER rely on logs alone — you need metrics and traces for full observability
- NEVER log request/response bodies by default — opt-in only, with PII redaction
- NEVER ignore log volume — set budgets and alert when a service exceeds daily quota
- NEVER skip context propagation in async flows — broken traces are worse than no traces
1---2name: logging-observability3description: Structured logging, distributed tracing, and metrics collection patterns for building observable systems. Use when implementing logging infrastructure, setting up distributed tracing with OpenTelemetry, designing metrics collection (RED/USE methods), configuring alerting and dashboards, or reviewing observability practices. Covers structured JSON logging, context propagation, trace sampling, Prometheus/Grafana stack, alert design, and PII/secret scrubbing.4---56# Logging & Observability78Patterns for building observable systems across the three pillars: logs, metrics, and traces.910## Three Pillars1112| Pillar | Purpose | Question It Answers | Example |13|--------|---------|---------------------|---------|14| **Logs** | What happened | Why did this request fail? | `{"level":"error","msg":"payment declined","user_id":"u_82"}` |15| **Metrics** | How much / how fast | Is latency increasing? | `http_request_duration_seconds{route="/api/orders"} 0.342` |16| **Traces** | Request flow | Where is the bottleneck? | Span: `api-gateway → auth → order-service → db` |1718Each pillar is strongest when correlated. Embed `trace_id` in every log line to jump from a log entry to the full distributed trace.1920---2122## Structured Logging2324Always emit logs as structured JSON — never free-text strings.2526### Required Fields2728| Field | Purpose | Required |29|-------|---------|----------|30| `timestamp` | ISO-8601 with milliseconds | Yes |31| `level` | Severity (DEBUG … FATAL) | Yes |32| `service` | Originating service name | Yes |33| `message` | Human-readable description | Yes |34| `trace_id` | Distributed trace correlation | Yes |35| `span_id` | Current span within trace | Yes |36| `correlation_id` | Business-level correlation (order ID) | When applicable |37| `error` | Structured error object | On errors |38| `context` | Request-specific metadata | Recommended |3940### Context Enrichment4142Attach context at the middleware level so downstream logs inherit automatically:4344```typescript45app.use((req, res, next) => {46 const ctx = {47 trace_id: req.headers['x-trace-id'] || crypto.randomUUID(),48 request_id: crypto.randomUUID(),49 user_id: req.user?.id,50 method: req.method,51 path: req.path,52 };53 asyncLocalStorage.run(ctx, () => next());54});55```5657### Library Recommendations5859| Library | Language | Strengths | Perf |60|---------|----------|-----------|------|61| **Pino** | Node.js | Fastest Node logger, low overhead | Excellent |62| **structlog** | Python | Composable processors, context binding | Good |63| **zerolog** | Go | Zero-allocation JSON logging | Excellent |64| **zap** | Go | High performance, typed fields | Excellent |65| **tracing** | Rust | Spans + events, async-aware | Excellent |6667Choose a logger that outputs structured JSON natively. Avoid loggers requiring post-processing.6869---7071## Log Levels7273| Level | When to Use | Example |74|-------|-------------|---------|75| **FATAL** | App cannot continue, process will exit | Database connection pool exhausted |76| **ERROR** | Operation failed, needs attention | Payment charge failed: CARD_DECLINED |77| **WARN** | Unexpected but recoverable | Retry 2/3 for upstream timeout |78| **INFO** | Normal business events | Order ORD-1234 placed successfully |79| **DEBUG** | Developer troubleshooting | Cache miss for key user:82:preferences |80| **TRACE** | Very fine-grained (rarely in prod) | Entering validateAddress with payload |8182**Rules:** Production default = INFO and above. If you log an ERROR, someone should act on it. Every FATAL should trigger an alert.8384---8586## Distributed Tracing8788### OpenTelemetry Setup8990Always prefer OpenTelemetry over vendor-specific SDKs:9192```typescript93import { NodeSDK } from '@opentelemetry/sdk-node';94import { OTLPTraceExporter } from '@opentelemetry/exporter-trace-otlp-http';95import { getNodeAutoInstrumentations } from '@opentelemetry/auto-instrumentations-node';9697const sdk = new NodeSDK({98 serviceName: 'order-service',99 traceExporter: new OTLPTraceExporter({100 url: 'http://otel-collector:4318/v1/traces',101 }),102 instrumentations: [getNodeAutoInstrumentations()],103});104sdk.start();105```106107### Span Creation108109```typescript110const tracer = trace.getTracer('order-service');111112async function processOrder(order: Order) {113 return tracer.startActiveSpan('processOrder', async (span) => {114 try {115 span.setAttribute('order.id', order.id);116 span.setAttribute('order.total_cents', order.totalCents);117 await validateInventory(order);118 await chargePayment(order);119 span.setStatus({ code: SpanStatusCode.OK });120 } catch (err) {121 span.setStatus({ code: SpanStatusCode.ERROR, message: err.message });122 span.recordException(err);123 throw err;124 } finally {125 span.end();126 }127 });128}129```130131### Context Propagation132133- Use W3C Trace Context (`traceparent` header) — default in OTel134- Propagate across HTTP, gRPC, and message queues135- For async workers: serialise `traceparent` into the job payload136137### Trace Sampling138139| Strategy | Use When |140|----------|----------|141| **Always On** | Low-traffic services, debugging |142| **Probabilistic** (N%) | General production use |143| **Rate-limited** (N/sec) | High-throughput services |144| **Tail-based** | When you need all error traces |145146Always sample 100% of error traces regardless of strategy.147148---149150## Metrics Collection151152### RED Method (Request-Driven)153154Monitor these three for every service endpoint:155156| Metric | What It Measures | Prometheus Example |157|--------|-----------------|-------------------|158| **Rate** | Requests/sec | `rate(http_requests_total[5m])` |159| **Errors** | Failed request ratio | `rate(http_requests_total{status=~"5.."}[5m])` |160| **Duration** | Response time | `histogram_quantile(0.99, http_request_duration_seconds)` |161162### USE Method (Resource-Driven)163164For infrastructure components (CPU, memory, disk, network):165166| Metric | What It Measures | Example |167|--------|-----------------|---------|168| **Utilization** | % resource busy | CPU usage at 78% |169| **Saturation** | Work queued/waiting | 12 requests queued in thread pool |170| **Errors** | Error events on resource | 3 disk I/O errors in last minute |171172---173174## Monitoring Stack175176| Tool | Category | Best For |177|------|----------|----------|178| **Prometheus** | Metrics | Pull-based metrics, alerting rules |179| **Grafana** | Visualisation | Dashboards for metrics, logs, traces |180| **Jaeger** | Tracing | Distributed trace visualisation |181| **Loki** | Logs | Log aggregation (pairs with Grafana) |182| **OpenTelemetry** | Collection | Vendor-neutral telemetry collection |183184**Recommendation:** Start with OTel Collector → Prometheus + Grafana + Loki + Jaeger. Migrate to SaaS only when operational overhead justifies cost.185186---187188## Alert Design189190### Severity Levels191192| Severity | Response Time | Example |193|----------|---------------|---------|194| **P1** | Immediate | Service fully down, data loss |195| **P2** | < 30 min | Error rate > 5%, latency p99 > 5s |196| **P3** | Business hours | Disk > 80%, cert expiring in 7 days |197| **P4** | Best effort | Non-critical deprecation warning |198199### Alert Fatigue Prevention200201- **Alert on symptoms, not causes** — "error rate > 5%" not "pod restarted"202- **Multi-window, multi-burn-rate** — catch both sudden spikes and slow burns203- **Require runbook links** — every alert must link to diagnosis and remediation204- **Review monthly** — delete or tune alerts that never fire or always fire205- **Group related alerts** — use inhibition rules to suppress child alerts206- **Set appropriate thresholds** — if alert fires daily and is ignored, raise threshold or delete207208---209210## Dashboard Patterns211212### Overview Dashboard ("War Room")213- Total requests/sec across all services214- Global error rate (%) with trendline215- p50 / p95 / p99 latency216- Active alerts count by severity217- Deployment markers overlaid on graphs218219### Service Dashboard (Per-Service)220- RED metrics for each endpoint221- Dependency health (upstream/downstream success rates)222- Resource utilisation (CPU, memory, connections)223- Top errors table with count and last seen224225---226227## Observability Checklist228229Every service must have:230231- [ ] Structured JSON logging with consistent schema232- [ ] Correlation / trace IDs propagated on all requests233- [ ] RED metrics exposed for every external endpoint234- [ ] Health check endpoints (`/healthz` and `/readyz`)235- [ ] Distributed tracing with OpenTelemetry236- [ ] Dashboards for RED metrics and resource utilisation237- [ ] Alerts for error rate, latency, and saturation with runbook links238- [ ] Log level configurable at runtime without redeployment239- [ ] PII scrubbing verified and tested240- [ ] Retention policies defined for logs, metrics, and traces241242## Anti-Patterns243244| Anti-Pattern | Problem | Fix |245|-------------|---------|-----|246| Logging PII | Privacy/compliance violation | Mask or exclude PII; use token references |247| Excessive logging | Storage costs balloon, signal drowns | Log business events, not data flow |248| Unstructured logs | Cannot query or alert on fields | Use structured JSON with consistent schema |249| String interpolation | Breaks structured fields, injection risk | Pass fields as metadata, not in message |250| Missing correlation IDs | Cannot trace across services | Generate and propagate trace_id everywhere |251| Alert storms | On-call fatigue, real issues buried | Use grouping, inhibition, deduplication |252| Metrics with high cardinality | Prometheus OOM, dashboard timeouts | Never use user ID or request ID as label |253254## NEVER Do2552561. **NEVER log passwords, tokens, API keys, or secrets** — even at DEBUG level2572. **NEVER use console.log / print in production** — use a structured logger2583. **NEVER use user IDs, emails, or request IDs as metric labels** — cardinality will explode2594. **NEVER create alerts without a runbook link** — unactionable alerts erode trust2605. **NEVER rely on logs alone** — you need metrics and traces for full observability2616. **NEVER log request/response bodies by default** — opt-in only, with PII redaction2627. **NEVER ignore log volume** — set budgets and alert when a service exceeds daily quota2638. **NEVER skip context propagation in async flows** — broken traces are worse than no traces