Observability Compliance
When to use
Use this skill when:
- reviewing a PR that adds new service operations, endpoints, or background jobs
- generating instrumentation code for new features
- auditing a service's observability coverage
- validating that log statements, metrics, and traces are correctly implemented
1. Structured Logging
- OC-01: ALL log statements MUST use structured logging (key/value pairs) — never concatenated strings.
- OC-02: Every backend log entry MUST include:
service, environment, correlationId.
- OC-03:
traceId and spanId MUST be included when distributed tracing context is available.
- OC-04:
eventName or operation MUST be included to identify the log's business context.
- OC-05: Secrets, tokens, API keys, passwords, and PII MUST NEVER appear in log statements.
- OC-06: Payloads MUST be redacted by default — only safe, non-sensitive fields are logged.
- OC-07: Log levels MUST be used correctly:
Debug — high-volume, local diagnostics only; disabled in production by default
Info — business lifecycle milestones (order placed, payment confirmed)
Warning — recoverable anomalies (retry attempted, fallback activated)
Error — failed operations requiring attention
Critical — immediate action required (data corruption risk, auth outage)
- OC-08: Errors MUST be logged at
Error or Critical — never swallowed silently.
2. Correlation & Trace Context
- OC-09:
correlationId MUST be propagated from every inbound request/message to ALL outbound calls (HTTP, broker, background jobs).
- OC-10:
causationId MUST be included on all domain and integration events (references the command or event that caused it).
- OC-11: HTTP services MUST accept an incoming correlation header (
X-Correlation-Id or equivalent), generate one if missing, and return it in the response.
- OC-12: Message consumers MUST extract and propagate correlation context from the message envelope.
- OC-13: Background jobs MUST generate or inherit a correlation context at start.
3. Metrics
3.1 Required per HTTP endpoint
- OC-14: Request rate (requests/second, by endpoint and method)
- OC-15: Error rate (4xx and 5xx, by endpoint)
- OC-16: Latency distribution (p50, p95, p99, by endpoint)
3.2 Required per use case (Command/Query)
- OC-17: Commands executed (count, by command type)
- OC-18: Command failure rate (count, by command type and error category)
- OC-19: Queries executed (count, by query type)
- OC-20: Handler duration (p50/p95/p99, by handler)
3.3 Required per messaging operation
- OC-21: Events published (count, by event type)
- OC-22: Events consumed (count, by event type and consumer)
- OC-23: Consumer processing failures (count, by event type)
- OC-24: Outbox pending count
- OC-25: DLQ message count (by queue)
- OC-26: Consumer lag (time between event occurrence and processing)
3.4 Required per external dependency call
- OC-27: Dependency error rate (by dependency name)
- OC-28: Dependency latency (p95/p99, by dependency)
- OC-29: Circuit breaker state (CLOSED/OPEN/HALF-OPEN, by dependency)
4. Distributed Tracing
- OC-30: Spans MUST be created for:
- inbound HTTP requests and gRPC calls
- outbound HTTP and gRPC calls
- database queries (at minimum for slow query threshold)
- message broker publish and consume operations
- significant background job phases
- OC-31: Span attributes MUST include:
- use case or operation name
- aggregate id (when safe — not PII)
- event type (for broker operations)
- outcome (success/failure)
- OC-32: Span names MUST be descriptive:
order.place, payment.charge, outbox.publish — not generic HTTP POST.
- OC-33: Trace context MUST be propagated via W3C
traceparent header or OpenTelemetry conventions.
5. Alerting Baseline
New services MUST configure baseline alerts:
- OC-34: p95 latency exceeds SLA threshold for > 5 minutes
- OC-35: Error rate exceeds threshold (default 1%) for > 2 minutes
- OC-36: DLQ count above threshold (default > 0) for > 5 minutes
- OC-37: Outbox backlog above threshold for > 5 minutes
- OC-38: Circuit breaker OPEN for > 60 seconds
- OC-39: Authentication error spike (> 10x normal rate)
6. Client Observability (Flutter / React Native)
- OC-40: Crashes MUST be captured and reported to crash reporting service.
- OC-41: Network failures MUST be captured (timeout, connection error, 5xx responses).
- OC-42: Slow screens/frames MUST be captured (render time > 500ms).
- OC-43:
correlationId MUST be propagated in all outbound API calls when provided by backend.
- OC-44: User-sensitive data MUST NOT be captured in crash reports or performance traces.
Examples
✅ Good:
_logger.LogInformation(
"Order placed successfully. OrderId={OrderId} CustomerId={CustomerId} CorrelationId={CorrelationId}",
order.Id, order.CustomerId, correlationId);
❌ Bad:
_logger.LogInformation($"Order placed: {order.Id} for user {user.Email} token={accessToken}");
// String concat + PII (email) + secret (token)
✅ Good: New endpoint registers counter, histogram, and propagates correlationId.
❌ Bad: New background job runs silently with no metrics and no correlation context.
Quick Mode
For low-context activation, load .enterprise/governance/agent-skills/observability-compliance/SKILL-QUICK.md or QUICK.md first. Load this full skill for deep analysis, violation fixing, or formal review gates.
1---2name: observability-compliance3description: Use when performing a deep observability audit or remediating missing structured logging, metrics, traces, or alerting rules4license: Apache-2.05---6
7# Observability Compliance
8
9## When to use
10Use this skill when:
11- reviewing a PR that adds new service operations, endpoints, or background jobs
12- generating instrumentation code for new features
13- auditing a service's observability coverage
14- validating that log statements, metrics, and traces are correctly implemented
15
16---
17
18## 1. Structured Logging
19
20- OC-01: ALL log statements MUST use **structured logging** (key/value pairs) — never concatenated strings.
21- OC-02: Every backend log entry MUST include: `service`, `environment`, `correlationId`.
22- OC-03: `traceId` and `spanId` MUST be included when distributed tracing context is available.
23- OC-04: `eventName` or `operation` MUST be included to identify the log's business context.
24- OC-05: Secrets, tokens, API keys, passwords, and PII MUST NEVER appear in log statements.
25- OC-06: Payloads MUST be redacted by default — only safe, non-sensitive fields are logged.
26- OC-07: Log levels MUST be used correctly:
27 - `Debug` — high-volume, local diagnostics only; disabled in production by default
28 - `Info` — business lifecycle milestones (order placed, payment confirmed)
29 - `Warning` — recoverable anomalies (retry attempted, fallback activated)
30 - `Error` — failed operations requiring attention
31 - `Critical` — immediate action required (data corruption risk, auth outage)
32- OC-08: Errors MUST be logged at `Error` or `Critical` — never swallowed silently.
33
34---
35
36## 2. Correlation & Trace Context
37
38- OC-09: `correlationId` MUST be propagated from every inbound request/message to ALL outbound calls (HTTP, broker, background jobs).
39- OC-10: `causationId` MUST be included on all domain and integration events (references the command or event that caused it).
40- OC-11: HTTP services MUST accept an incoming correlation header (`X-Correlation-Id` or equivalent), generate one if missing, and return it in the response.
41- OC-12: Message consumers MUST extract and propagate correlation context from the message envelope.
42- OC-13: Background jobs MUST generate or inherit a correlation context at start.
43
44---
45
46## 3. Metrics
47
48### 3.1 Required per HTTP endpoint
49- OC-14: Request rate (requests/second, by endpoint and method)
50- OC-15: Error rate (4xx and 5xx, by endpoint)
51- OC-16: Latency distribution (p50, p95, p99, by endpoint)
52
53### 3.2 Required per use case (Command/Query)
54- OC-17: Commands executed (count, by command type)
55- OC-18: Command failure rate (count, by command type and error category)
56- OC-19: Queries executed (count, by query type)
57- OC-20: Handler duration (p50/p95/p99, by handler)
58
59### 3.3 Required per messaging operation
60- OC-21: Events published (count, by event type)
61- OC-22: Events consumed (count, by event type and consumer)
62- OC-23: Consumer processing failures (count, by event type)
63- OC-24: Outbox pending count
64- OC-25: DLQ message count (by queue)
65- OC-26: Consumer lag (time between event occurrence and processing)
66
67### 3.4 Required per external dependency call
68- OC-27: Dependency error rate (by dependency name)
69- OC-28: Dependency latency (p95/p99, by dependency)
70- OC-29: Circuit breaker state (CLOSED/OPEN/HALF-OPEN, by dependency)
71
72---
73
74## 4. Distributed Tracing
75
76- OC-30: Spans MUST be created for:
77 - inbound HTTP requests and gRPC calls
78 - outbound HTTP and gRPC calls
79 - database queries (at minimum for slow query threshold)
80 - message broker publish and consume operations
81 - significant background job phases
82- OC-31: Span attributes MUST include:
83 - use case or operation name
84 - aggregate id (when safe — not PII)
85 - event type (for broker operations)
86 - outcome (success/failure)
87- OC-32: Span names MUST be descriptive: `order.place`, `payment.charge`, `outbox.publish` — not generic `HTTP POST`.
88- OC-33: Trace context MUST be propagated via W3C `traceparent` header or OpenTelemetry conventions.
89
90---
91
92## 5. Alerting Baseline
93
94New services MUST configure baseline alerts:
95
96- OC-34: p95 latency exceeds SLA threshold for > 5 minutes
97- OC-35: Error rate exceeds threshold (default 1%) for > 2 minutes
98- OC-36: DLQ count above threshold (default > 0) for > 5 minutes
99- OC-37: Outbox backlog above threshold for > 5 minutes
100- OC-38: Circuit breaker OPEN for > 60 seconds
101- OC-39: Authentication error spike (> 10x normal rate)
102
103---
104
105## 6. Client Observability (Flutter / React Native)
106
107- OC-40: Crashes MUST be captured and reported to crash reporting service.
108- OC-41: Network failures MUST be captured (timeout, connection error, 5xx responses).
109- OC-42: Slow screens/frames MUST be captured (render time > 500ms).
110- OC-43: `correlationId` MUST be propagated in all outbound API calls when provided by backend.
111- OC-44: User-sensitive data MUST NOT be captured in crash reports or performance traces.
112
113---
114
115## Examples
116
117✅ Good:
118```csharp
119_logger.LogInformation(
120 "Order placed successfully. OrderId={OrderId} CustomerId={CustomerId} CorrelationId={CorrelationId}",
121 order.Id, order.CustomerId, correlationId);
122```
123
124❌ Bad:
125```csharp
126_logger.LogInformation($"Order placed: {order.Id} for user {user.Email} token={accessToken}");
127// String concat + PII (email) + secret (token)
128```
129
130✅ Good: New endpoint registers counter, histogram, and propagates correlationId.
131❌ Bad: New background job runs silently with no metrics and no correlation context.
132
133
134## Quick Mode
135
136For low-context activation, load `.enterprise/governance/agent-skills/observability-compliance/SKILL-QUICK.md` or `QUICK.md` first. Load this full skill for deep analysis, violation fixing, or formal review gates.
137