Complexity Levels
| Level |
Tools |
Setup Time |
Best For |
| Minimal |
UptimeRobot, Healthchecks.io |
15 min |
Side projects, MVPs |
| Standard |
Uptime Kuma, Sentry, basic Grafana |
1-2 hours |
Small teams, startups |
| Professional |
Prometheus, Grafana, Loki, Alertmanager |
1-2 days |
Production systems |
| Enterprise |
Datadog, New Relic, or full OSS stack |
Ongoing |
Large-scale operations |
The Three Pillars
| Pillar |
What It Answers |
Tools |
| Metrics |
"How is the system performing?" |
Prometheus, Grafana, Datadog |
| Logs |
"What happened?" |
Loki, ELK, CloudWatch |
| Traces |
"Why is this request slow?" |
Jaeger, Tempo, Sentry |
Quick Start by Use Case
"I just want to know if it's down"
→ UptimeRobot (free) or Uptime Kuma (self-hosted). See simple.md.
"I need to debug production errors"
→ Sentry with your framework SDK. 5-minute setup. See apm.md.
"I want real observability"
→ Prometheus + Grafana + Loki. See prometheus.md.
"I need to centralize logs"
→ Loki for simple, ELK for complex queries. See logs.md.
What to Monitor
Applications (RED Method)
- Rate — requests per second
- Errors — error rate by endpoint
- Duration — latency (p50, p95, p99)
Infrastructure (USE Method)
- Utilization — CPU, memory, disk usage
- Saturation — queue depth, load average
- Errors — hardware/system errors
Alerting Principles
| Do |
Don't |
| Alert on symptoms (user impact) |
Alert on causes (CPU high) |
| Include runbook link |
Require investigation to understand |
| Set appropriate severity |
Make everything P1 |
| Require action |
Alert on "interesting" metrics |
Alert fatigue kills monitoring. If alerts are ignored, you have no monitoring.
For alert configuration, severities, and on-call setup, see alerting.md.
Cost Comparison
| Solution |
Monthly Cost (small) |
Monthly Cost (medium) |
| UptimeRobot |
Free |
$7 |
| Uptime Kuma |
$5 (VPS) |
$5 (VPS) |
| Sentry |
Free / $26 |
$80 |
| Grafana Cloud |
Free tier |
$50+ |
| Datadog |
$15/host |
$23/host + features |
| Self-hosted stack |
$10-20 (VPS) |
$50-100 (VPS) |
Common Mistakes
- Starting with Prometheus/Grafana when Uptime Kuma would suffice
- No alerting (dashboards nobody watches)
- Too many alerts (alert fatigue → ignored)
- Missing runbooks (alert fires, nobody knows what to do)
- Not monitoring from outside (only internal checks)
- Storing logs forever (cost explodes)
1---2name: monitoring3description: Set up observability for applications and infrastructure with metrics, logs, traces, and alerts.4---5
6## Complexity Levels
7
8| Level | Tools | Setup Time | Best For |
9|-------|-------|------------|----------|
10| **Minimal** | UptimeRobot, Healthchecks.io | 15 min | Side projects, MVPs |
11| **Standard** | Uptime Kuma, Sentry, basic Grafana | 1-2 hours | Small teams, startups |
12| **Professional** | Prometheus, Grafana, Loki, Alertmanager | 1-2 days | Production systems |
13| **Enterprise** | Datadog, New Relic, or full OSS stack | Ongoing | Large-scale operations |
14
15## The Three Pillars
16
17| Pillar | What It Answers | Tools |
18|--------|-----------------|-------|
19| **Metrics** | "How is the system performing?" | Prometheus, Grafana, Datadog |
20| **Logs** | "What happened?" | Loki, ELK, CloudWatch |
21| **Traces** | "Why is this request slow?" | Jaeger, Tempo, Sentry |
22
23## Quick Start by Use Case
24
25**"I just want to know if it's down"**
26→ UptimeRobot (free) or Uptime Kuma (self-hosted). See `simple.md`.
27
28**"I need to debug production errors"**
29→ Sentry with your framework SDK. 5-minute setup. See `apm.md`.
30
31**"I want real observability"**
32→ Prometheus + Grafana + Loki. See `prometheus.md`.
33
34**"I need to centralize logs"**
35→ Loki for simple, ELK for complex queries. See `logs.md`.
36
37## What to Monitor
38
39### Applications (RED Method)
40- **R**ate — requests per second
41- **E**rrors — error rate by endpoint
42- **D**uration — latency (p50, p95, p99)
43
44### Infrastructure (USE Method)
45- **U**tilization — CPU, memory, disk usage
46- **S**aturation — queue depth, load average
47- **E**rrors — hardware/system errors
48
49## Alerting Principles
50
51| Do | Don't |
52|----|-------|
53| Alert on symptoms (user impact) | Alert on causes (CPU high) |
54| Include runbook link | Require investigation to understand |
55| Set appropriate severity | Make everything P1 |
56| Require action | Alert on "interesting" metrics |
57
58**Alert fatigue kills monitoring.** If alerts are ignored, you have no monitoring.
59
60For alert configuration, severities, and on-call setup, see `alerting.md`.
61
62## Cost Comparison
63
64| Solution | Monthly Cost (small) | Monthly Cost (medium) |
65|----------|---------------------|----------------------|
66| UptimeRobot | Free | $7 |
67| Uptime Kuma | $5 (VPS) | $5 (VPS) |
68| Sentry | Free / $26 | $80 |
69| Grafana Cloud | Free tier | $50+ |
70| Datadog | $15/host | $23/host + features |
71| Self-hosted stack | $10-20 (VPS) | $50-100 (VPS) |
72
73## Common Mistakes
74
75- Starting with Prometheus/Grafana when Uptime Kuma would suffice
76- No alerting (dashboards nobody watches)
77- Too many alerts (alert fatigue → ignored)
78- Missing runbooks (alert fires, nobody knows what to do)
79- Not monitoring from outside (only internal checks)
80- Storing logs forever (cost explodes)