AWS Observability
Overview
Domain expertise for AWS observability across metrics, logs, and traces, covering the full lifecycle: enabling/onboarding a service to Application Signals using ADOT (AWS Distro for OpenTelemetry) auto-instrumentation SDKs and ServiceEvents — making the service show up in Application Signals — on EC2, ECS, EKS, and Lambda in Python, Node.js, Java, and .NET.
Works best with the AWS MCP server — enables running CLI commands, querying CloudWatch, and validating configurations directly. All guidance also works with standard AWS CLI access.
Note: Reference files contain specific runtime versions, quota values, and feature matrices that may change. When precision matters (e.g., deploying to production, choosing a runtime, or checking a quota), confirm values against current AWS documentation rather than relying solely on the values in these files.
Routing
| User need |
Action |
| Enabling/onboarding a service to Application Signals (auto-instrumentation) |
Read application-signals-onboarding.md |
| Propagating ServiceEvents git/deployment metadata through CI/CD |
Read application-signals-cicd-metadata.md |
| Per-platform/per-language enablement steps |
Read the matching references/appsignals-guides/<platform>-<language>.md (e.g. eks-python.md) |
| Writing Log Insights queries |
Read log-insights.md |
| Configuring alarms (metric, composite, anomaly) |
Read alarms.md |
| Publishing custom metrics or using EMF |
Read metrics.md |
| Setting up X-Ray tracing or ADOT |
Read tracing.md |
| Building dashboards |
Read dashboards.md |
| Debugging observability issues |
Read troubleshooting.md — starts with the 5 most common fixes |
| Debugging canary failures |
Read synthetics.md — see Common failures table |
| CloudTrail operational auditing |
Read cloudtrail.md |
| Setting up Lambda monitoring with CDK |
Use alarm-template.ts as a starting point |
| Creating synthetic canaries |
Read synthetics.md |
| Configuring ADOT collector |
Use otel-config.yaml as a starting point |
| Debugging a running service with breakpoints/snapshots — Dynamic Instrumentation (modifies live services and capture live data) |
Read dynamic-instrumentation.md in full before acting. Confirm with the user before any create/delete, and narrate before significant actions: observation → hypothesis → proposed action → expected result. Diagnosing running-service root cause from source/code inspection. Source inspection alone identifies hypotheses, not confirmed root causes. Keep suspected causes tentative until runtime evidence confirms them. |
| Spans multiple areas |
Read the most specific reference first, then consult others as needed |
Files
| File |
Content |
| application-signals-onboarding.md |
Enable Application Signals auto-instrumentation: EKS add-on, CloudWatch Agent IAM, OTLP endpoints, ServiceEvents env vars, Dynamic Instrumentation — two-tier scope by platform/language |
| application-signals-cicd-metadata.md |
ServiceEvents git & deployment metadata propagation through CI/CD (the 5 OTEL_AWS_SERVICE_EVENTS_* vars) |
references/appsignals-guides/ (e.g. eks-python.md) |
16 per-platform × per-language enablement guides (EC2/ECS/EKS/Lambda × Python/Node.js/Java/.NET) |
| alarms.md |
Metric, composite, anomaly detection alarms — configuration, constraints, recommended defaults |
| log-insights.md |
Complete query syntax, commands, functions, known issues, reusable query library |
| metrics.md |
Custom metrics, EMF spec, metric filters, high-resolution, retention |
| tracing.md |
X-Ray → ADOT migration, sampling rules, annotations vs metadata, collector config |
| dashboards.md |
Widget types, cross-account/region, dynamic labels, sharing |
| troubleshooting.md |
Error → cause → fix for all observability services |
| cloudtrail.md |
Operational auditing, event types, S3+Athena queries |
| synthetics.md |
Canary runtime/blueprint constraints, VPC networking, common failures |
| alarm-template.ts |
Best-practice CDK Lambda monitoring (alarms + dashboard) |
| otel-config.yaml |
ADOT collector config for X-Ray traces + CloudWatch EMF metrics |
| dynamic-instrumentation.md |
Dynamic Instrumentation debugging loop — breakpoints/probes on live code, snapshot capture + correlation analysis, create/delete gating, snapshot PII handling. Runs via scripts/di_instrumentation.py + scripts/di_snapshots.py. |
1---2name: aws-observability3description: Builds, configures, debugs, and optimizes AWS observability with CloudWatch (Log Insights, Metrics, Alarms, Dashboards, EMF), X-Ray, CloudTrail, and ADOT (AWS Distro for OpenTelemetry), AND enables/onboards services to Application Signals using ADOT auto-instrumentation SDKs. Covers Log Insights queries, alarms (metric, composite, anomaly), dashboards, custom metrics/EMF, X-Ray tracing and sampling, ADOT collector config, CloudTrail auditing, and end-to-end Application Signals enablement via ADOT SDKs (CloudWatch Observability EKS add-on, CloudWatch Agent IAM, OTLP endpoints, ServiceEvents, Dynamic Instrumentation), breakpoint and snapshot in Dynamic Instrumentation, live data capture in running service, debug without redeploying. Applies to CloudWatch, alarms, dashboards, EMF, X-Ray, traces, CloudTrail, ADOT, monitoring, synthetics/canaries, OR enabling/onboarding/instrumenting a service for Application Signals. Not for app logging or security threat detection.4license: Apache-2.05---67# AWS Observability89## Overview1011Domain expertise for AWS observability across metrics, logs, and traces, covering the full lifecycle: **enabling/onboarding** a service to Application Signals using ADOT (AWS Distro for OpenTelemetry) auto-instrumentation SDKs and ServiceEvents — making the service show up in Application Signals — on EC2, ECS, EKS, and Lambda in Python, Node.js, Java, and .NET.1213**Works best with** the [AWS MCP server](https://docs.aws.amazon.com/aws-mcp/) — enables running CLI commands, querying CloudWatch, and validating configurations directly. All guidance also works with standard AWS CLI access.1415**Note:** Reference files contain specific runtime versions, quota values, and feature matrices that may change. When precision matters (e.g., deploying to production, choosing a runtime, or checking a quota), confirm values against current AWS documentation rather than relying solely on the values in these files.1617## Routing1819| User need | Action |20|-----------|--------|21| Enabling/onboarding a service to Application Signals (auto-instrumentation) | Read [application-signals-onboarding.md](references/application-signals-onboarding.md) |22| Propagating ServiceEvents git/deployment metadata through CI/CD | Read [application-signals-cicd-metadata.md](references/application-signals-cicd-metadata.md) |23| Per-platform/per-language enablement steps | Read the matching `references/appsignals-guides/<platform>-<language>.md` (e.g. [eks-python.md](references/appsignals-guides/eks-python.md)) |24| Writing Log Insights queries | Read [log-insights.md](references/log-insights.md) |25| Configuring alarms (metric, composite, anomaly) | Read [alarms.md](references/alarms.md) |26| Publishing custom metrics or using EMF | Read [metrics.md](references/metrics.md) |27| Setting up X-Ray tracing or ADOT | Read [tracing.md](references/tracing.md) |28| Building dashboards | Read [dashboards.md](references/dashboards.md) |29| Debugging observability issues | Read [troubleshooting.md](references/troubleshooting.md) — starts with the 5 most common fixes |30| Debugging canary failures | Read [synthetics.md](references/synthetics.md) — see Common failures table |31| CloudTrail operational auditing | Read [cloudtrail.md](references/cloudtrail.md) |32| Setting up Lambda monitoring with CDK | Use [alarm-template.ts](assets/alarm-template.ts) as a starting point |33| Creating synthetic canaries | Read [synthetics.md](references/synthetics.md) |34| Configuring ADOT collector | Use [otel-config.yaml](assets/otel-config.yaml) as a starting point |35| Debugging a running service with breakpoints/snapshots — Dynamic Instrumentation (**modifies live services and capture live data**) | Read [dynamic-instrumentation.md](references/dynamic-instrumentation.md) in full before acting. Confirm with the user before any create/delete, and narrate before significant actions: observation → hypothesis → proposed action → expected result. Diagnosing running-service root cause from source/code inspection. Source inspection alone identifies hypotheses, not confirmed root causes. Keep suspected causes tentative until runtime evidence confirms them. |36| Spans multiple areas | Read the most specific reference first, then consult others as needed |3738## Files3940| File | Content |41|------|---------|42| [application-signals-onboarding.md](references/application-signals-onboarding.md) | Enable Application Signals auto-instrumentation: EKS add-on, CloudWatch Agent IAM, OTLP endpoints, ServiceEvents env vars, Dynamic Instrumentation — two-tier scope by platform/language |43| [application-signals-cicd-metadata.md](references/application-signals-cicd-metadata.md) | ServiceEvents git & deployment metadata propagation through CI/CD (the 5 `OTEL_AWS_SERVICE_EVENTS_*` vars) |44| `references/appsignals-guides/` (e.g. [eks-python.md](references/appsignals-guides/eks-python.md)) | 16 per-platform × per-language enablement guides (EC2/ECS/EKS/Lambda × Python/Node.js/Java/.NET) |45| [alarms.md](references/alarms.md) | Metric, composite, anomaly detection alarms — configuration, constraints, recommended defaults |46| [log-insights.md](references/log-insights.md) | Complete query syntax, commands, functions, known issues, reusable query library |47| [metrics.md](references/metrics.md) | Custom metrics, EMF spec, metric filters, high-resolution, retention |48| [tracing.md](references/tracing.md) | X-Ray → ADOT migration, sampling rules, annotations vs metadata, collector config |49| [dashboards.md](references/dashboards.md) | Widget types, cross-account/region, dynamic labels, sharing |50| [troubleshooting.md](references/troubleshooting.md) | Error → cause → fix for all observability services |51| [cloudtrail.md](references/cloudtrail.md) | Operational auditing, event types, S3+Athena queries |52| [synthetics.md](references/synthetics.md) | Canary runtime/blueprint constraints, VPC networking, common failures |53| [alarm-template.ts](assets/alarm-template.ts) | Best-practice CDK Lambda monitoring (alarms + dashboard) |54| [otel-config.yaml](assets/otel-config.yaml) | ADOT collector config for X-Ray traces + CloudWatch EMF metrics |55| [dynamic-instrumentation.md](references/dynamic-instrumentation.md) | Dynamic Instrumentation debugging loop — breakpoints/probes on live code, snapshot capture + correlation analysis, create/delete gating, snapshot PII handling. Runs via `scripts/di_instrumentation.py` + `scripts/di_snapshots.py`. |