Monitoring and Observability
If the system can fail in a way users notice, you should be able to see it before they tell you.
Context
Monitoring and observability translate system behavior into signals responders can trust.
See context and anti-pattern notes.
Inputs
I/O contract notes define required inputs and authority.
Process
Step 1: Map User-Critical and Boundary-Critical Signals
Start from the reviewed architecture, API contract, and delivery slice:
- which flows must succeed
- which flows must fail closed
- which queues, bridges, or background workers can amplify failures
- which rollback or coexistence indicators matter after release
Do not start from "what metrics are easy to collect."
Step 2: Define a Small Set of Actionable Signals
At minimum, cover:
- availability and latency for user-facing routes
- error-rate segmentation for supported vs unsupported paths where that distinction matters
- queue depth, retry rate, or backfill lag for async seams
- release or deploy markers so responders can correlate changes quickly
- rollback or fallback health signals if the current slice depends on them
Step 3: Build Dashboards That Support Triage
Dashboards should answer:
- what is broken
- who is affected
- whether the system is degrading or recovering
- whether rollback, fail-closed behavior, or fallback mode is working
Prefer a small number of responder dashboards over a wall of charts.
Step 4: Design Alerts for Actionability
Each alert should have:
- a clear trigger
- user or business impact context
- owner or escalation target
- immediate next step or linked runbook
Alerts that cannot change behavior are noise.
Step 5: Validate with Release and Incident Scenarios
Before relying on the setup:
- verify alerts fire for the risky boundary you care about
- verify dashboards show release markers and recovery clearly
- verify the signal is strong enough to support incident-response and rollback decisions
Outputs
Produce only declared outputs at their documented quality boundary.
Quality Gate
- User-facing and boundary-critical flows are explicitly instrumented
- Alerts are actionable and mapped to an owner or runbook
- Release markers and rollback or fallback health are visible
- Async seams or coexistence boundaries are observable where applicable
- Responder dashboard supports triage without requiring ad hoc queries first