# Prometheus Grafana Otel Tracing

> Unified observability: OpenTelemetry Collector pipelines, Prometheus scraping and recording rules, Loki log correlation, Tempo distributed tracing, exemplars, sampling strategy, and cardinality control. Use when instrumenting services so a latency spike can be followed to the exact trace and log line, building the metrics-logs-traces stack, or fixing missing telemetry and cardinality blowups.

- Skill: `mchittineni/prometheus-grafana-otel-tracing` (Agent Skill, multi-file: 2 files)
- Install (CLI): `npx skillmds@latest add mchittineni/prometheus-grafana-otel-tracing`
- Raw SKILL.md: https://api.skillmd.com/api/skills/mchittineni/prometheus-grafana-otel-tracing/raw
- Safety review: pending
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Coding & Dev Tools
- Author: mchittineni (https://skillmd.com/u/mchittineni)
- Updated: 2026-09-17
- Page: https://skillmd.com/skills/mchittineni/prometheus-grafana-otel-tracing

---


# Unified Observability: OpenTelemetry, Prometheus & Distributed Tracing

## When to Use This Skill

**Triggers — load this skill when:**

- A service or platform needs metrics, logs, and traces wired end to end
- Trace sampling, exemplars, or log-to-trace correlation must be configured
- Cardinality or collector resource problems are degrading the telemetry stack

**Route elsewhere when:**

- Objective and burn-rate alert definition -> `sli-slo-error-budget-design`
- Node/container infrastructure metrics -> `infrastructure-host-monitoring`

## 1. OpenTelemetry Collector Pipeline Configuration (`otel-collector.yaml`)

```yaml
receivers:
  otlp:
    protocols:
      grpc:
        endpoint: 0.0.0.0:4317
      http:
        endpoint: 0.0.0.0:4318

processors:
  batch:
    timeout: 1s
    send_batch_size: 1024
  memory_limiter:
    check_interval: 1s
    limit_percentage: 75
    spike_limit_percentage: 20
  resourcedetection:
    detectors: [env, gcp, ecs, ec2, azure]
    timeout: 2s

exporters:
  prometheus:
    endpoint: "0.0.0.0:8889"
    namespace: "otel"
  otlp/tempo:
    endpoint: "tempo.monitoring:4317"
    tls:
      insecure: true
  loki:
    endpoint: "http://loki.monitoring:3100/loki/api/v1/push"

service:
  pipelines:
    traces:
      receivers: [otlp]
      processors: [memory_limiter, batch, resourcedetection]
      exporters: [otlp/tempo]
    metrics:
      receivers: [otlp]
      processors: [memory_limiter, batch]
      exporters: [prometheus]
    logs:
      receivers: [otlp]
      processors: [memory_limiter, batch]
      exporters: [loki]
```

---

## 2. Distributed Tracing Best Practices

- **Trace Context Propagation**: Pass `traceparent` headers via W3C TraceContext standards across all inter-service REST, gRPC, and Kafka messages.
- **Span Attributes**: Enrich spans with high-cardinality metadata (`customer_id`, `order_id`, `http.status_code`).
- **Tail-based Sampling**: Sample 100% of error traces and slow traces (> 95th percentile latency) while dropping 90% of routine fast transactions to control storage costs.

---

## 3. Anti-Patterns

| Anti-pattern | Why it fails in production |
| --- | --- |
| Putting a user ID, request ID, email or full URL in a metric label | Cardinality is multiplicative: one unbounded label multiplied by an existing one takes Prometheus OOM in hours. High-cardinality identifiers belong on traces and logs, never on metrics. |
| Sampling traces at the SDK with a fixed low rate | The 1% you keep is random, so the slow and failing requests you actually needed are gone. Use tail sampling in the Collector, keeping all errors and slow traces plus a small baseline of successes. |
| Instrumenting every service with a different SDK config | Inconsistent resource attributes (`service.name`, `deployment.environment`) make cross-service queries impossible. Standardise resource attributes centrally and inject them via the Collector. |
| Logging without a trace ID | Three services' logs stay three unrelated streams, and every incident starts with manual correlation by timestamp. Propagate `traceparent` and emit `trace_id` on every structured log line. |
| Running the Collector as a single deployment for everything | One noisy service's spans starve the pipeline for all others. Split agent (per-node, cheap) from gateway (central, batching and tail sampling), and set memory limiter processors. |
| Dashboards built per team with no shared definitions | "Latency" means p50 on one dashboard and mean on another; incident calls argue about the data instead of the outage. Define recording rules once and build dashboards from them. |

