Datadog Observability
Overview
Datadog is a SaaS observability platform providing unified monitoring across infrastructure, applications, logs, and user experience. It offers AI-powered anomaly detection, 1000+ integrations, and OpenTelemetry compatibility.
Core Capabilities:
- APM: Distributed tracing with automatic instrumentation for 8+ languages
- Infrastructure: Host, container, and cloud service monitoring
- Logs: Centralized collection with processing pipelines and 15-month retention
- Metrics: Custom metrics via DogStatsD with cardinality management
- Synthetics: Proactive API and browser testing from 29+ global locations
- RUM: Frontend performance with Core Web Vitals and session replay
When to Use This Skill
Activate when:
- Setting up production monitoring and observability
- Implementing distributed tracing across microservices
- Configuring log aggregation and analysis pipelines
- Creating custom metrics and dashboards
- Setting up alerting and anomaly detection
- Optimizing Datadog costs
Do not use when:
- Building with open-source stack (use Prometheus/Grafana instead)
- Cost is primary concern and budget is limited
- Need maximum customization over managed solution
Quick Start
1. Install Datadog Agent
Docker (simplest):
docker run -d --name dd-agent \
-e DD_API_KEY=<YOUR_API_KEY> \
-e DD_SITE="datadoghq.com" \
-v /var/run/docker.sock:/var/run/docker.sock:ro \
-v /proc/:/host/proc/:ro \
-v /sys/fs/cgroup/:/host/sys/fs/cgroup:ro \
gcr.io/datadoghq/agent:7
Kubernetes (Helm):
helm repo add datadog https://helm.datadoghq.com
helm install datadog-agent datadog/datadog \
--set datadog.apiKey=<YOUR_API_KEY> \
--set datadog.apm.enabled=true \
--set datadog.logs.enabled=true
2. Instrument Your Application
Python:
from ddtrace import tracer, patch_all
# Automatic instrumentation for common libraries
patch_all()
# Manual span for custom operations
with tracer.trace("custom.operation", service="my-service") as span:
span.set_tag("user.id", user_id)
# your code here
Node.js:
// Must be first import
const tracer = require('dd-trace').init({
service: 'my-service',
env: 'production',
version: '1.0.0',
});
3. Verify in Datadog UI
- Go to Infrastructure > Host Map to verify agent
- Go to APM > Services to see traced services
- Go to Logs > Search to verify log collection
Core Concepts
Tagging Strategy
Tags enable filtering, aggregation, and cost attribution. Use consistent tags across all telemetry.
Required Tags:
| Tag |
Purpose |
Example |
env |
Environment |
env:production |
service |
Service name |
service:api-gateway |
version |
Deployment version |
version:1.2.3 |
team |
Owning team |
team:platform |
Avoid High-Cardinality Tags:
- User IDs, request IDs, timestamps
- Pod IDs in Kubernetes
- Build numbers, commit hashes
Unified Observability
Datadog correlates metrics, traces, and logs automatically:
- Traces include span tags that link to metrics
- Logs inject trace IDs for correlation
- Dashboards combine all data sources
Best Practices
Start Simple
- Install Agent with basic configuration
- Enable automatic instrumentation
- Verify data in Datadog UI
- Add custom spans/metrics as needed
Progressive Enhancement
Basic → APM tracing → Custom spans → Custom metrics → Profiling → RUM
Key Instrumentation Points
- HTTP entry/exit points
- Database queries
- External service calls
- Message queue operations
- Business-critical flows
Common Mistakes
- High-cardinality tags: Using user IDs or request IDs as tags creates millions of unique metrics
- Missing log index quotas: Leads to unexpected bills from log volume spikes
- Over-alerting: Creates alert fatigue; alert on symptoms, not causes
- Missing service tags: Prevents correlation between metrics, traces, and logs
- No sampling for high-volume traces: Ingests everything, causing cost explosion
Navigation
For detailed implementation:
- Agent Installation: Docker, Kubernetes, Linux, Windows, and cloud-specific setup
- APM Instrumentation: Python, Node.js, Go, Java instrumentation with code examples
- Log Management: Pipelines, Grok parsing, standard attributes, archives
- Custom Metrics: DogStatsD patterns, metric types, tagging best practices
- Alerting: Monitor types, anomaly detection, alert hygiene
- Cost Optimization: Metrics without Limits, sampling, index quotas
- Kubernetes: DaemonSet, Cluster Agent, autodiscovery
Complementary Skills
When using this skill, consider these related skills (if deployed):
- docker: Container instrumentation patterns
- kubernetes: K8s-native monitoring patterns
- python/nodejs/go: Language-specific APM setup
Resources
Official Documentation:
Cost Management:
1---2name: datadog3description: Full-stack observability with Datadog APM, logs, metrics, synthetics, and RUM. Use when implementing monitoring, tracing, alerting, or cost optimization for production systems.4license: MIT5---6
7# Datadog Observability
8
9## Overview
10
11Datadog is a SaaS observability platform providing unified monitoring across infrastructure, applications, logs, and user experience. It offers AI-powered anomaly detection, 1000+ integrations, and OpenTelemetry compatibility.
12
13**Core Capabilities:**
14- **APM**: Distributed tracing with automatic instrumentation for 8+ languages
15- **Infrastructure**: Host, container, and cloud service monitoring
16- **Logs**: Centralized collection with processing pipelines and 15-month retention
17- **Metrics**: Custom metrics via DogStatsD with cardinality management
18- **Synthetics**: Proactive API and browser testing from 29+ global locations
19- **RUM**: Frontend performance with Core Web Vitals and session replay
20
21## When to Use This Skill
22
23**Activate when:**
24- Setting up production monitoring and observability
25- Implementing distributed tracing across microservices
26- Configuring log aggregation and analysis pipelines
27- Creating custom metrics and dashboards
28- Setting up alerting and anomaly detection
29- Optimizing Datadog costs
30
31**Do not use when:**
32- Building with open-source stack (use Prometheus/Grafana instead)
33- Cost is primary concern and budget is limited
34- Need maximum customization over managed solution
35
36## Quick Start
37
38### 1. Install Datadog Agent
39
40**Docker (simplest):**
41```bash
42docker run -d --name dd-agent \
43 -e DD_API_KEY=<YOUR_API_KEY> \
44 -e DD_SITE="datadoghq.com" \
45 -v /var/run/docker.sock:/var/run/docker.sock:ro \
46 -v /proc/:/host/proc/:ro \
47 -v /sys/fs/cgroup/:/host/sys/fs/cgroup:ro \
48 gcr.io/datadoghq/agent:7
49```
50
51**Kubernetes (Helm):**
52```bash
53helm repo add datadog https://helm.datadoghq.com
54helm install datadog-agent datadog/datadog \
55 --set datadog.apiKey=<YOUR_API_KEY> \
56 --set datadog.apm.enabled=true \
57 --set datadog.logs.enabled=true
58```
59
60### 2. Instrument Your Application
61
62**Python:**
63```python
64from ddtrace import tracer, patch_all
65
66# Automatic instrumentation for common libraries
67patch_all()
68
69# Manual span for custom operations
70with tracer.trace("custom.operation", service="my-service") as span:
71 span.set_tag("user.id", user_id)
72 # your code here
73```
74
75**Node.js:**
76```javascript
77// Must be first import
78const tracer = require('dd-trace').init({
79 service: 'my-service',
80 env: 'production',
81 version: '1.0.0',
82});
83```
84
85### 3. Verify in Datadog UI
86
871. Go to Infrastructure > Host Map to verify agent
882. Go to APM > Services to see traced services
893. Go to Logs > Search to verify log collection
90
91## Core Concepts
92
93### Tagging Strategy
94
95Tags enable filtering, aggregation, and cost attribution. Use consistent tags across all telemetry.
96
97**Required Tags:**
98| Tag | Purpose | Example |
99|-----|---------|---------|
100| `env` | Environment | `env:production` |
101| `service` | Service name | `service:api-gateway` |
102| `version` | Deployment version | `version:1.2.3` |
103| `team` | Owning team | `team:platform` |
104
105**Avoid High-Cardinality Tags:**
106- User IDs, request IDs, timestamps
107- Pod IDs in Kubernetes
108- Build numbers, commit hashes
109
110### Unified Observability
111
112Datadog correlates metrics, traces, and logs automatically:
113- Traces include span tags that link to metrics
114- Logs inject trace IDs for correlation
115- Dashboards combine all data sources
116
117## Best Practices
118
119### Start Simple
1201. Install Agent with basic configuration
1212. Enable automatic instrumentation
1223. Verify data in Datadog UI
1234. Add custom spans/metrics as needed
124
125### Progressive Enhancement
126```
127Basic → APM tracing → Custom spans → Custom metrics → Profiling → RUM
128```
129
130### Key Instrumentation Points
131- HTTP entry/exit points
132- Database queries
133- External service calls
134- Message queue operations
135- Business-critical flows
136
137## Common Mistakes
138
1391. **High-cardinality tags**: Using user IDs or request IDs as tags creates millions of unique metrics
1402. **Missing log index quotas**: Leads to unexpected bills from log volume spikes
1413. **Over-alerting**: Creates alert fatigue; alert on symptoms, not causes
1424. **Missing service tags**: Prevents correlation between metrics, traces, and logs
1435. **No sampling for high-volume traces**: Ingests everything, causing cost explosion
144
145## Navigation
146
147For detailed implementation:
148
149- **[Agent Installation](references/agent-installation.md)**: Docker, Kubernetes, Linux, Windows, and cloud-specific setup
150- **[APM Instrumentation](references/apm-instrumentation.md)**: Python, Node.js, Go, Java instrumentation with code examples
151- **[Log Management](references/log-management.md)**: Pipelines, Grok parsing, standard attributes, archives
152- **[Custom Metrics](references/custom-metrics.md)**: DogStatsD patterns, metric types, tagging best practices
153- **[Alerting](references/alerting.md)**: Monitor types, anomaly detection, alert hygiene
154- **[Cost Optimization](references/cost-optimization.md)**: Metrics without Limits, sampling, index quotas
155- **[Kubernetes](references/kubernetes.md)**: DaemonSet, Cluster Agent, autodiscovery
156
157## Complementary Skills
158
159When using this skill, consider these related skills (if deployed):
160
161- **docker**: Container instrumentation patterns
162- **kubernetes**: K8s-native monitoring patterns
163- **python/nodejs/go**: Language-specific APM setup
164
165## Resources
166
167**Official Documentation:**
168- APM: https://docs.datadoghq.com/tracing/
169- Logs: https://docs.datadoghq.com/logs/
170- Metrics: https://docs.datadoghq.com/metrics/
171- DogStatsD: https://docs.datadoghq.com/developers/dogstatsd/
172
173**Cost Management:**
174- Billing: https://docs.datadoghq.com/account_management/billing/
175- Usage Attribution: https://docs.datadoghq.com/account_management/billing/usage_attribution/