AWS Observability
Overview
Domain expertise for AWS observability across metrics, logs, and traces, covering the full lifecycle: enabling/onboarding Application Signals on a service using ADOT (AWS Distro for OpenTelemetry) auto-instrumentation SDKs through operating it (CloudWatch alarms, dashboards, Log Insights, custom metrics, EMF, X-Ray trace analysis, CloudTrail auditing, ADOT collector config).
Works best with the AWS MCP server — enables running CLI commands, querying CloudWatch, and validating configurations directly. All guidance also works with standard AWS CLI access.
Note: Reference files contain specific runtime versions, quota values, and feature matrices that may change. When precision matters (e.g., deploying to production, choosing a runtime, or checking a quota), confirm values against current AWS documentation rather than relying solely on the values in these files.
Logs Insights query gotchas (common, easy to misdiagnose)
- A query whose time window is entirely before a log group's creation time fails with
MalformedQueryException: Query's end date and time is either before the log groups creation time or exceeds the log groups log retention settings. This bites synthetic/seeded data and freshly-created groups: back-dating events into a just-created group, then querying that past window, is rejected. Ensure the query window overlaps[log-group-creation-time, now]. - Logs Insights has ingestion latency — just-written events are not immediately queryable. After
PutLogEvents, a query can returnstatus: CompletewithrecordsMatched = 0for events that are physically present in the stream. Back-dated events are especially unreliable. For any test/automation that depends on log evidence: seed events at ~now, then poll the query (retry untilrecordsMatched > 0or a timeout) before asserting downstream behavior. Budget a settling period (often minutes). - Metrics (
GetMetricData) have no equivalent lag — custom metrics accept back-dated timestamps and are queryable immediately by absolute time. When both signals exist, metrics are the reliable/fast path; treat logs as eventually-queryable. - When a collector wraps Insights, make its diagnostics distinguish query failed vs returned zero rows vs still running — collapsing all three into one "unavailable" message hides which of the above you hit.
Routing
| User need | Action |
|---|---|
| Enabling/onboarding a service to Application Signals (auto-instrumentation) | Read application-signals-onboarding.md |
| Propagating ServiceEvents git/deployment metadata through CI/CD | Read application-signals-cicd-metadata.md |
| Per-platform/per-language enablement steps | Read the matching references/appsignals-guides/<platform>-<language>.md (e.g. eks-python.md) |
| Writing Log Insights queries | Read log-insights.md |
| Configuring alarms (metric, composite, anomaly) | Read alarms.md |
| Publishing custom metrics or using EMF | Read metrics.md |
| Setting up X-Ray tracing or ADOT | Read tracing.md |
| Building dashboards | Read dashboards.md |
| Debugging observability issues | Read troubleshooting.md — starts with the 5 most common fixes |
| Debugging canary failures | Read synthetics.md — see Common failures table |
| CloudTrail operational auditing | Read cloudtrail.md |
| Setting up Lambda monitoring with CDK | Use alarm-template.ts as a starting point |
| Creating synthetic canaries | Read synthetics.md |
| Configuring ADOT collector | Use otel-config.yaml as a starting point |
| Spans multiple areas | Read the most specific reference first, then consult others as needed |
Files
| File | Content |
|---|---|
| application-signals-onboarding.md | Enable Application Signals auto-instrumentation: EKS add-on, CloudWatch Agent IAM, OTLP endpoints, ServiceEvents env vars, Dynamic Instrumentation — two-tier scope by platform/language |
| application-signals-cicd-metadata.md | ServiceEvents git & deployment metadata propagation through CI/CD (the 5 OTEL_AWS_SERVICE_EVENTS_* vars) |
references/appsignals-guides/ (e.g. eks-python.md) |
16 per-platform × per-language enablement guides (EC2/ECS/EKS/Lambda × Python/Node.js/Java/.NET) |
| alarms.md | Metric, composite, anomaly detection alarms — configuration, constraints, recommended defaults |
| log-insights.md | Complete query syntax, commands, functions, known issues, reusable query library |
| metrics.md | Custom metrics, EMF spec, metric filters, high-resolution, retention |
| tracing.md | X-Ray → ADOT migration, sampling rules, annotations vs metadata, collector config |
| dashboards.md | Widget types, cross-account/region, dynamic labels, sharing |
| troubleshooting.md | Error → cause → fix for all observability services |
| cloudtrail.md | Operational auditing, event types, S3+Athena queries |
| synthetics.md | Canary runtime/blueprint constraints, VPC networking, common failures |
| alarm-template.ts | Best-practice CDK Lambda monitoring (alarms + dashboard) |
| otel-config.yaml | ADOT collector config for X-Ray traces + CloudWatch EMF metrics |