Log Aggregation Architect
Expert in designing centralized log pipelines with structured logging, efficient collection, and cost-effective retention.
Activation Triggers
Activate on: "log aggregation", "structured logging", "Fluentd config", "Vector pipeline", "Loki setup", "ELK stack", "log pipeline", "log retention policy", "centralized logging", "log shipping"
NOT for: Metrics/dashboards → monitoring-stack-deployer | Distributed tracing → logging-observability | Alerting → site-reliability-engineer
Quick Start
- Standardize log format — JSON structured logs with consistent fields across all services
- Deploy collection agents — Vector or Fluentd as DaemonSet on every node
- Choose storage backend — Grafana Loki (cost-effective), Elasticsearch (full-text search), or ClickHouse (analytics)
- Define retention tiers — hot (7d searchable), warm (30d compressed), cold (1y archived)
- Build correlation — trace ID propagation so logs link to traces and metrics
Core Capabilities
| Domain |
Technologies |
| Collection |
Vector 0.43, Fluentd 1.17, Fluent Bit 3.2, OTEL Collector |
| Storage |
Grafana Loki 3.x, Elasticsearch 8.x, ClickHouse, S3 archive |
| Structured Logging |
JSON, logfmt, OpenTelemetry Logs, pino, winston, slog (Go) |
| Pipeline |
Transform, filter, route, sample, deduplicate, redact PII |
| Visualization |
Grafana (Loki), Kibana (Elastic), Grafana Explore |
Architecture Patterns
Vector Pipeline (Recommended 2026)
# vector.toml — collect, transform, route
[sources.kubernetes]
type = "kubernetes_logs"
auto_partial_merge = true
[transforms.structured]
type = "remap"
inputs = ["kubernetes"]
source = '''
# Parse JSON logs, fallback to raw message
. = parse_json(.message) ?? {"message": .message}
.timestamp = now()
.service = .kubernetes.pod_labels."app.kubernetes.io/name" ?? "unknown"
.environment = .kubernetes.pod_namespace
# Redact PII
.message = redact(.message, filters: ["pattern"],
patterns: [r'\b[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Z|a-z]{2,}\b'])
'''
[transforms.sampler]
type = "sample"
inputs = ["structured"]
rate = 10 # Keep 1 in 10 debug logs
exclude."level" = ["error", "warn", "info"] # Always keep non-debug
[sinks.loki]
type = "loki"
inputs = ["sampler"]
endpoint = "http://loki-gateway:3100"
labels.service = "{{ service }}"
labels.level = "{{ level }}"
encoding.codec = "json"
Structured Log Schema (Cross-Language Standard)
{
"timestamp": "2026-03-20T14:30:00.000Z",
"level": "info",
"message": "Order processed successfully",
"service": "order-api",
"trace_id": "abc123def456",
"span_id": "789ghi",
"user_id": "usr_masked",
"order_id": "ord_12345",
"duration_ms": 142,
"http": {
"method": "POST",
"path": "/api/v1/orders",
"status": 201
}
}
Retention Tier Strategy
HOT (0-7 days):
├─ Full-text searchable in Loki/Elasticsearch
├─ Instant query response (<1s)
└─ Cost: $$$ (SSD, indexed)
WARM (7-30 days):
├─ Compressed, queryable with delay
├─ Query response 5-30s
└─ Cost: $$ (HDD, partial index)
COLD (30-365 days):
├─ S3/GCS archive, queryable via Athena/BigQuery
├─ Query response: minutes
└─ Cost: $ (object storage, no index)
DELETED (365+ days):
└─ Lifecycle policy auto-deletes (compliance permitting)
Anti-Patterns
- Unstructured string logs —
console.log("User " + id + " did thing") is unsearchable. Use structured JSON with consistent field names.
- Logging sensitive data — PII, tokens, passwords in logs. Redact at the pipeline level (Vector remap, Fluentd filter) before storage.
- No log levels — everything at INFO. Use DEBUG for development, INFO for business events, WARN for recoverable issues, ERROR for failures requiring attention.
- Unbounded retention — keeping all logs forever. Define retention tiers with automatic lifecycle policies. Most logs lose value after 30 days.
- Missing trace correlation — logs without trace IDs cannot be correlated with distributed traces. Propagate OpenTelemetry trace context into every log line.
Quality Checklist
[ ] All services emit JSON structured logs
[ ] Consistent field schema across services (timestamp, level, service, trace_id)
[ ] Log collection agents deployed as DaemonSet (Vector or Fluent Bit)
[ ] PII redaction applied in pipeline before storage
[ ] Debug logs sampled (not all collected in production)
[ ] Retention tiers defined: hot, warm, cold with lifecycle policies
[ ] Trace IDs propagated into log context
[ ] Log-based alerts configured for error rate spikes
[ ] Grafana Explore or Kibana connected for log search
[ ] Storage costs monitored and budget-capped
[ ] Log pipeline has backpressure handling (no data loss under load)
[ ] Compliance requirements met (GDPR right-to-erasure for logs with PII)
1---2name: log-aggregation-architect3description: Centralized log pipeline architect with structured logging, Fluentd/Vector, and retention policies. Activate on: log aggregation, structured logging, Fluentd, Vector, Loki, ELK stack, log pipeline, log retention, centralized logging. NOT for: metrics and dashboards (use monitoring-stack-deployer), distributed tracing (use logging-observability), alerting rules (use site-reliability-engineer).4license: Apache-2.05---6
7# Log Aggregation Architect
8
9Expert in designing centralized log pipelines with structured logging, efficient collection, and cost-effective retention.
10
11## Activation Triggers
12
13**Activate on:** "log aggregation", "structured logging", "Fluentd config", "Vector pipeline", "Loki setup", "ELK stack", "log pipeline", "log retention policy", "centralized logging", "log shipping"
14
15**NOT for:** Metrics/dashboards → `monitoring-stack-deployer` | Distributed tracing → `logging-observability` | Alerting → `site-reliability-engineer`
16
17## Quick Start
18
191. **Standardize log format** — JSON structured logs with consistent fields across all services
202. **Deploy collection agents** — Vector or Fluentd as DaemonSet on every node
213. **Choose storage backend** — Grafana Loki (cost-effective), Elasticsearch (full-text search), or ClickHouse (analytics)
224. **Define retention tiers** — hot (7d searchable), warm (30d compressed), cold (1y archived)
235. **Build correlation** — trace ID propagation so logs link to traces and metrics
24
25## Core Capabilities
26
27| Domain | Technologies |
28|--------|-------------|
29| **Collection** | Vector 0.43, Fluentd 1.17, Fluent Bit 3.2, OTEL Collector |
30| **Storage** | Grafana Loki 3.x, Elasticsearch 8.x, ClickHouse, S3 archive |
31| **Structured Logging** | JSON, logfmt, OpenTelemetry Logs, pino, winston, slog (Go) |
32| **Pipeline** | Transform, filter, route, sample, deduplicate, redact PII |
33| **Visualization** | Grafana (Loki), Kibana (Elastic), Grafana Explore |
34
35## Architecture Patterns
36
37### Vector Pipeline (Recommended 2026)
38
39```toml
40# vector.toml — collect, transform, route
41[sources.kubernetes]
42type = "kubernetes_logs"
43auto_partial_merge = true
44
45[transforms.structured]
46type = "remap"
47inputs = ["kubernetes"]
48source = '''
49 # Parse JSON logs, fallback to raw message
50 . = parse_json(.message) ?? {"message": .message}
51 .timestamp = now()
52 .service = .kubernetes.pod_labels."app.kubernetes.io/name" ?? "unknown"
53 .environment = .kubernetes.pod_namespace
54 # Redact PII
55 .message = redact(.message, filters: ["pattern"],
56 patterns: [r'\b[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Z|a-z]{2,}\b'])
57'''
58
59[transforms.sampler]
60type = "sample"
61inputs = ["structured"]
62rate = 10 # Keep 1 in 10 debug logs
63exclude."level" = ["error", "warn", "info"] # Always keep non-debug
64
65[sinks.loki]
66type = "loki"
67inputs = ["sampler"]
68endpoint = "http://loki-gateway:3100"
69labels.service = "{{ service }}"
70labels.level = "{{ level }}"
71encoding.codec = "json"
72```
73
74### Structured Log Schema (Cross-Language Standard)
75
76```json
77{
78 "timestamp": "2026-03-20T14:30:00.000Z",
79 "level": "info",
80 "message": "Order processed successfully",
81 "service": "order-api",
82 "trace_id": "abc123def456",
83 "span_id": "789ghi",
84 "user_id": "usr_masked",
85 "order_id": "ord_12345",
86 "duration_ms": 142,
87 "http": {
88 "method": "POST",
89 "path": "/api/v1/orders",
90 "status": 201
91 }
92}
93```
94
95### Retention Tier Strategy
96
97```
98HOT (0-7 days):
99 ├─ Full-text searchable in Loki/Elasticsearch
100 ├─ Instant query response (<1s)
101 └─ Cost: $$$ (SSD, indexed)
102
103WARM (7-30 days):
104 ├─ Compressed, queryable with delay
105 ├─ Query response 5-30s
106 └─ Cost: $$ (HDD, partial index)
107
108COLD (30-365 days):
109 ├─ S3/GCS archive, queryable via Athena/BigQuery
110 ├─ Query response: minutes
111 └─ Cost: $ (object storage, no index)
112
113DELETED (365+ days):
114 └─ Lifecycle policy auto-deletes (compliance permitting)
115```
116
117## Anti-Patterns
118
1191. **Unstructured string logs** — `console.log("User " + id + " did thing")` is unsearchable. Use structured JSON with consistent field names.
1202. **Logging sensitive data** — PII, tokens, passwords in logs. Redact at the pipeline level (Vector remap, Fluentd filter) before storage.
1213. **No log levels** — everything at INFO. Use DEBUG for development, INFO for business events, WARN for recoverable issues, ERROR for failures requiring attention.
1224. **Unbounded retention** — keeping all logs forever. Define retention tiers with automatic lifecycle policies. Most logs lose value after 30 days.
1235. **Missing trace correlation** — logs without trace IDs cannot be correlated with distributed traces. Propagate OpenTelemetry trace context into every log line.
124
125## Quality Checklist
126
127```
128[ ] All services emit JSON structured logs
129[ ] Consistent field schema across services (timestamp, level, service, trace_id)
130[ ] Log collection agents deployed as DaemonSet (Vector or Fluent Bit)
131[ ] PII redaction applied in pipeline before storage
132[ ] Debug logs sampled (not all collected in production)
133[ ] Retention tiers defined: hot, warm, cold with lifecycle policies
134[ ] Trace IDs propagated into log context
135[ ] Log-based alerts configured for error rate spikes
136[ ] Grafana Explore or Kibana connected for log search
137[ ] Storage costs monitored and budget-capped
138[ ] Log pipeline has backpressure handling (no data loss under load)
139[ ] Compliance requirements met (GDPR right-to-erasure for logs with PII)
140```