SITUATION: Monitoring & Observability Setup
Skill gerado a partir do pack templates-claude-code. Arquivo de origem: situacao/14-monitoramento-observabilidade.md. Use como baseline e adapte ao projeto antes de mudancas grandes.
Conteudo do template
CONTEXT
Use this configuration when setting up or improving observability for a live application — adding metrics, structured logging, distributed tracing, alerting, SLO definitions, or runbook creation. Observability is not an add-on — an unmonitored production service is a liability you can't reason about. This configuration ensures you can answer: "Is the service healthy right now? How do I know? What do I do when it's not?"
OBJECTIVES
- Instrument all three pillars of observability: metrics, logs, and traces
- Define SLOs/SLIs before incidents reveal you needed them
- Alert on symptoms (what users experience) not just causes (CPU usage)
- Create actionable runbooks linked from every alert
- Enable a developer to debug a production incident using only the observability stack
APPROACH RULES
Three pillars: metrics, logs, traces. Metrics tell you something is wrong. Logs tell you what happened. Traces tell you where in the system it happened. You need all three for full observability.
Alert on symptoms, not causes. "Error rate > 1%" is a symptom — users are affected. "CPU > 80%" is a cause — but doesn't necessarily mean users are impacted. Start with symptom-based alerts.
Every alert must have a runbook. If an alert fires, the on-call engineer must know what to do. An alert without a runbook is just noise that trains people to ignore alerts. Runbook link must be in the alert body.
SLO > SLA. Define Service Level Objectives (SLOs) and track your error budget. If you're burning through error budget, slow down feature development and fix reliability. If you have budget headroom, safe to move fast.
Structured logging everywhere. Never console.log("user logged in"). Always logger.info({ event: "user_login", userId: user.id, ip: req.ip }). Structured logs are queryable. Strings are not.
Correlation IDs across services. Every request gets a unique trace/correlation ID at the entry point. Every log line, every downstream call includes that ID. Without this, distributed debugging is a nightmare.
Monitor what users do, not just what systems do. Business metrics (conversion rate, checkout completions, active users) belong on the same dashboard as technical metrics. When technical metrics are fine but business metrics drop, you have a bug.
ROUTING TABLE
| If you encounter |
Then |
| No logging in the application |
Add structured logger first (pino, winston, structlog). All existing console.log → logger.info/error. |
| console.log in production code |
Replace with structured logger. console.log is not queryable, not leveled, not structured. |
| Alert without runbook |
Write the runbook before the alert goes live. Template: What is this alert? What is the impact? What are the steps? |
| P99 latency not tracked |
Add P95 and P99 latency histograms. Averages hide worst-case user experience. |
| Errors logged without stack trace |
Fix immediately. Logs without stack traces are unusable for debugging. |
| Alert firing with no action taken |
Either fix the condition or remove the alert. Alert fatigue kills incident response. |
| No distributed tracing |
Add OpenTelemetry SDK. Instrument HTTP calls, DB queries, queue operations. |
| SLO not defined |
Define Error Rate SLO (e.g., 99.9% of requests succeed) and Latency SLO (P95 < 500ms). |
| Logs growing without limit |
Add log rotation. Set retention policies. Production logs: 30 days hot, 90 days cold. |
| Multiple services with no centralized logging |
Implement log aggregation: ELK stack, Loki + Grafana, or Datadog/New Relic. |
Three Pillars Implementation Guide
Metrics (What is happening)
// Using prom-client (Node.js)
import { Counter, Histogram, register } from 'prom-client'
const httpRequestDuration = new Histogram({
name: 'http_request_duration_seconds',
help: 'HTTP request duration in seconds',
labelNames: ['method', 'route', 'status_code'],
buckets: [0.01, 0.05, 0.1, 0.25, 0.5, 1, 2.5, 5]
})
const httpRequestTotal = new Counter({
name: 'http_requests_total',
help: 'Total HTTP requests',
labelNames: ['method', 'route', 'status_code']
})
// Expose at /metrics endpoint
app.get('/metrics', async (req, res) => {
res.set('Content-Type', register.contentType)
res.send(await register.metrics())
})
Logs (What happened)
// Using pino (Node.js)
import pino from 'pino'
const logger = pino({
level: process.env.LOG_LEVEL || 'info',
base: { service: 'api', version: process.env.APP_VERSION },
timestamp: pino.stdTimeFunctions.isoTime
})
// ✅ Good: structured, queryable
logger.info({ event: 'payment_processed', userId, amount, currency }, 'Payment processed')
logger.error({ err: error, userId, requestId }, 'Payment failed')
// ❌ Bad: string only, not queryable
console.log(`Payment failed for user ${userId}: ${error.message}`)
Traces (Where it happened)
// Using OpenTelemetry
import { trace, context, propagation } from '@opentelemetry/api'
const tracer = trace.getTracer('my-service')
async function processOrder(orderId: string) {
const span = tracer.startSpan('processOrder', {
attributes: { 'order.id': orderId }
})
try {
await span.setAttribute('order.status', 'processing')
const result = await doWork(orderId)
span.setStatus({ code: SpanStatusCode.OK })
return result
} catch (err) {
span.recordException(err)
span.setStatus({ code: SpanStatusCode.ERROR })
throw err
} finally {
span.end()
}
}
SLO Definition Template
service: checkout-api
slos:
- name: Checkout Success Rate
description: Percentage of checkout requests that succeed
sli: (sum of successful checkout requests) / (sum of all checkout requests)
target: 99.5%
measurement_window: 30 days
error_budget: 0.5% → 216 minutes/month allowable downtime
- name: Checkout Latency
description: P95 latency for checkout endpoint
sli: 95th percentile response time
target: < 1000ms
measurement_window: 30 days
alerts:
- name: Checkout Error Rate Burn
condition: Error budget burn rate > 5× for 1 hour
severity: page (wake someone up)
runbook: docs/runbooks/checkout-errors.md
DO NOT
- DO NOT set up monitoring after the first incident — it's too late then
- DO NOT log PII (email, name, phone, card numbers) — mask or hash before logging
- DO NOT create alerts that can't be acted on — each alert needs a clear response procedure
- DO NOT use average latency as your primary metric — use P95 and P99
- DO NOT monitor only infrastructure (CPU, memory) and ignore application-level health
- DO NOT alert with CRITICAL severity on things that are informational
- DO NOT delete old logs without understanding retention requirements (GDPR, compliance)
OUTPUT FORMAT
For each monitoring setup, deliver:
Observability Coverage Matrix:
| Signal | Tool | Coverage | Retention | Alert? |
|--------|------|----------|-----------|--------|
| App logs | Loki | INFO+ in prod. DEBUG in staging | 30 days | On ERROR+ |
| Infrastructure metrics | Prometheus | CPU, memory, disk, network | 90 days | On threshold |
| App metrics | Prometheus | Request rate, error rate, latency histograms | 90 days | On SLO burn |
| Traces | Jaeger/Tempo | Sample rate: 10% prod, 100% staging | 7 days | On trace error |
| Uptime | Healthchecks.io | External probe every 60s | 365 days | On unavailable |
**Runbook Inventory:**
| Alert | Runbook | Owner | Last Verified |
|-------|---------|-------|---------------|
| High error rate | docs/runbooks/high-error-rate.md | Platform | 2024-01-01 |
QUALITY GATES
1---2name: tpl-situacao-monitoramento-observabilidade3description: Template do pack (situacao/14-monitoramento-observabilidade.md). Orienta o agente em tarefas situacionais como debug, seguranca e refactor alinhado a esse contexto.4---56# SITUATION: Monitoring & Observability Setup78Skill gerado a partir do pack `templates-claude-code`. Arquivo de origem: `situacao/14-monitoramento-observabilidade.md`. Use como baseline e adapte ao projeto antes de mudancas grandes.910## Conteudo do template1112## CONTEXT13Use this configuration when setting up or improving observability for a live application — adding metrics, structured logging, distributed tracing, alerting, SLO definitions, or runbook creation. Observability is not an add-on — an unmonitored production service is a liability you can't reason about. This configuration ensures you can answer: "Is the service healthy right now? How do I know? What do I do when it's not?"1415## OBJECTIVES16- Instrument all three pillars of observability: metrics, logs, and traces17- Define SLOs/SLIs before incidents reveal you needed them18- Alert on symptoms (what users experience) not just causes (CPU usage)19- Create actionable runbooks linked from every alert20- Enable a developer to debug a production incident using only the observability stack2122## APPROACH RULES23241. **Three pillars: metrics, logs, traces.** Metrics tell you something is wrong. Logs tell you what happened. Traces tell you where in the system it happened. You need all three for full observability.25262. **Alert on symptoms, not causes.** "Error rate > 1%" is a symptom — users are affected. "CPU > 80%" is a cause — but doesn't necessarily mean users are impacted. Start with symptom-based alerts.27283. **Every alert must have a runbook.** If an alert fires, the on-call engineer must know what to do. An alert without a runbook is just noise that trains people to ignore alerts. Runbook link must be in the alert body.29304. **SLO > SLA.** Define Service Level Objectives (SLOs) and track your error budget. If you're burning through error budget, slow down feature development and fix reliability. If you have budget headroom, safe to move fast.31325. **Structured logging everywhere.** Never `console.log("user logged in")`. Always `logger.info({ event: "user_login", userId: user.id, ip: req.ip })`. Structured logs are queryable. Strings are not.33346. **Correlation IDs across services.** Every request gets a unique trace/correlation ID at the entry point. Every log line, every downstream call includes that ID. Without this, distributed debugging is a nightmare.35367. **Monitor what users do, not just what systems do.** Business metrics (conversion rate, checkout completions, active users) belong on the same dashboard as technical metrics. When technical metrics are fine but business metrics drop, you have a bug.3738## ROUTING TABLE3940| If you encounter | Then |41|-----------------|------|42| No logging in the application | Add structured logger first (pino, winston, structlog). All existing console.log → logger.info/error. |43| console.log in production code | Replace with structured logger. console.log is not queryable, not leveled, not structured. |44| Alert without runbook | Write the runbook before the alert goes live. Template: What is this alert? What is the impact? What are the steps? |45| P99 latency not tracked | Add P95 and P99 latency histograms. Averages hide worst-case user experience. |46| Errors logged without stack trace | Fix immediately. Logs without stack traces are unusable for debugging. |47| Alert firing with no action taken | Either fix the condition or remove the alert. Alert fatigue kills incident response. |48| No distributed tracing | Add OpenTelemetry SDK. Instrument HTTP calls, DB queries, queue operations. |49| SLO not defined | Define Error Rate SLO (e.g., 99.9% of requests succeed) and Latency SLO (P95 < 500ms). |50| Logs growing without limit | Add log rotation. Set retention policies. Production logs: 30 days hot, 90 days cold. |51| Multiple services with no centralized logging | Implement log aggregation: ELK stack, Loki + Grafana, or Datadog/New Relic. |5253## Three Pillars Implementation Guide5455### Metrics (What is happening)56```typescript57// Using prom-client (Node.js)58import { Counter, Histogram, register } from 'prom-client'5960const httpRequestDuration = new Histogram({61 name: 'http_request_duration_seconds',62 help: 'HTTP request duration in seconds',63 labelNames: ['method', 'route', 'status_code'],64 buckets: [0.01, 0.05, 0.1, 0.25, 0.5, 1, 2.5, 5]65})6667const httpRequestTotal = new Counter({68 name: 'http_requests_total',69 help: 'Total HTTP requests',70 labelNames: ['method', 'route', 'status_code']71})7273// Expose at /metrics endpoint74app.get('/metrics', async (req, res) => {75 res.set('Content-Type', register.contentType)76 res.send(await register.metrics())77})78```7980### Logs (What happened)81```typescript82// Using pino (Node.js)83import pino from 'pino'8485const logger = pino({86 level: process.env.LOG_LEVEL || 'info',87 base: { service: 'api', version: process.env.APP_VERSION },88 timestamp: pino.stdTimeFunctions.isoTime89})9091// ✅ Good: structured, queryable92logger.info({ event: 'payment_processed', userId, amount, currency }, 'Payment processed')93logger.error({ err: error, userId, requestId }, 'Payment failed')9495// ❌ Bad: string only, not queryable96console.log(`Payment failed for user ${userId}: ${error.message}`)97```9899### Traces (Where it happened)100```typescript101// Using OpenTelemetry102import { trace, context, propagation } from '@opentelemetry/api'103104const tracer = trace.getTracer('my-service')105106async function processOrder(orderId: string) {107 const span = tracer.startSpan('processOrder', {108 attributes: { 'order.id': orderId }109 })110 111 try {112 await span.setAttribute('order.status', 'processing')113 const result = await doWork(orderId)114 span.setStatus({ code: SpanStatusCode.OK })115 return result116 } catch (err) {117 span.recordException(err)118 span.setStatus({ code: SpanStatusCode.ERROR })119 throw err120 } finally {121 span.end()122 }123}124```125126## SLO Definition Template127128```yaml129service: checkout-api130slos:131 - name: Checkout Success Rate132 description: Percentage of checkout requests that succeed133 sli: (sum of successful checkout requests) / (sum of all checkout requests)134 target: 99.5%135 measurement_window: 30 days136 error_budget: 0.5% → 216 minutes/month allowable downtime137138 - name: Checkout Latency139 description: P95 latency for checkout endpoint140 sli: 95th percentile response time141 target: < 1000ms142 measurement_window: 30 days143144alerts:145 - name: Checkout Error Rate Burn146 condition: Error budget burn rate > 5× for 1 hour147 severity: page (wake someone up)148 runbook: docs/runbooks/checkout-errors.md149```150151## DO NOT152153- **DO NOT** set up monitoring after the first incident — it's too late then154- **DO NOT** log PII (email, name, phone, card numbers) — mask or hash before logging155- **DO NOT** create alerts that can't be acted on — each alert needs a clear response procedure156- **DO NOT** use average latency as your primary metric — use P95 and P99157- **DO NOT** monitor only infrastructure (CPU, memory) and ignore application-level health158- **DO NOT** alert with CRITICAL severity on things that are informational159- **DO NOT** delete old logs without understanding retention requirements (GDPR, compliance)160161## OUTPUT FORMAT162163For each monitoring setup, deliver:164165**Observability Coverage Matrix:**166```markdown167| Signal | Tool | Coverage | Retention | Alert? |168|--------|------|----------|-----------|--------|169| App logs | Loki | INFO+ in prod. DEBUG in staging | 30 days | On ERROR+ |170| Infrastructure metrics | Prometheus | CPU, memory, disk, network | 90 days | On threshold |171| App metrics | Prometheus | Request rate, error rate, latency histograms | 90 days | On SLO burn |172| Traces | Jaeger/Tempo | Sample rate: 10% prod, 100% staging | 7 days | On trace error |173| Uptime | Healthchecks.io | External probe every 60s | 365 days | On unavailable |174175**Runbook Inventory:**176| Alert | Runbook | Owner | Last Verified |177|-------|---------|-------|---------------|178| High error rate | docs/runbooks/high-error-rate.md | Platform | 2024-01-01 |179```180181## QUALITY GATES182183- [ ] P95 and P99 latency tracked for all user-facing endpoints184- [ ] Error rate (5xx rate) tracked and alerting at > 1%185- [ ] Every log entry is structured JSON with at minimum: timestamp, level, service, message186- [ ] Correlation/trace IDs present in all logs for every request187- [ ] At least one SLO defined with error budget tracked on dashboard188- [ ] Every production alert has a linked runbook189- [ ] Runbook tested: follow it cold and confirm it resolves the described scenario190- [ ] PII not present in any log line (verified by audit sample)191- [ ] On-call rotation documented and first responder always identified192- [ ] Post-mortem template ready before first incident (not after)