Observability Designer (POWERFUL)
Category: Engineering
Tier: POWERFUL
Description: Design comprehensive observability strategies for production systems including SLI/SLO frameworks, alerting optimization, and dashboard generation.
Overview
Observability Designer creates production-ready dashboards, alert configurations, and monitoring strategies across the three pillars (metrics, logs, traces).
When NOT to use → slo-architect. For SLO/SLI design with error-budget math, multi-window burn-rate alerting thresholds, and SLO review gates, route to slo-architect — it is the authoritative skill for that half. This skill's slo_designer.py produces a quick scaffold only. This skill's lane: dashboards (dashboard_generator.py) and alert-noise reduction (alert_optimizer.py).
Quick Start
# Dashboard spec (Grafana JSON + docs) for a service
python3 scripts/dashboard_generator.py --service-type api --name payments --criticality critical --role sre --format grafana -o dashboard.json --doc-output dashboard.md
# Analyze an existing alert config for noise, duplicates, and coverage gaps
python3 scripts/alert_optimizer.py --input alerts.json --analyze-only --report alert_report.json
# ...then emit the optimized config once the report is reviewed:
python3 scripts/alert_optimizer.py --input alerts.json --output alerts_optimized.json
# Quick SLO scaffold (hand off to slo-architect for the real error-budget work)
python3 scripts/slo_designer.py --service-type api --criticality high --user-facing true --service-name payments -o slo_scaffold.json
Verification loop: after deploying optimized alerts, track the report's noise metrics for one on-call rotation — if the actionable-alert ratio didn't improve, re-run --analyze-only against the live config and iterate. Import the generated dashboard into Grafana and confirm every golden-signal panel renders with live data before closing the task.
Core Competencies
SLI/SLO/SLA Framework Design
- Service Level Indicators (SLI): Define measurable signals that indicate service health
- Service Level Objectives (SLO): Set reliability targets based on user experience
- Service Level Agreements (SLA): Establish customer-facing commitments with consequences
- Error Budget Management: Calculate and track error budget consumption
- Burn Rate Alerting: Multi-window burn rate alerts for proactive SLO protection
Three Pillars of Observability
Metrics
- Golden Signals: Latency, traffic, errors, and saturation monitoring
- RED Method: Rate, Errors, and Duration for request-driven services
- USE Method: Utilization, Saturation, and Errors for resource monitoring
- Business Metrics: Revenue, user engagement, and feature adoption tracking
- Infrastructure Metrics: CPU, memory, disk, network, and custom resource metrics
Logs
- Structured Logging: JSON-based log formats with consistent fields
- Log Aggregation: Centralized log collection and indexing strategies
- Log Levels: Appropriate use of DEBUG, INFO, WARN, ERROR, FATAL levels
- Correlation IDs: Request tracing through distributed systems
- Log Sampling: Volume management for high-throughput systems
Traces
- Distributed Tracing: End-to-end request flow visualization
- Span Design: Meaningful span boundaries and metadata
- Trace Sampling: Intelligent sampling strategies for performance and cost
- Service Maps: Automatic dependency discovery through traces
- Root Cause Analysis: Trace-driven debugging workflows
Dashboard Design Principles
Information Architecture
- Hierarchy: Overview → Service → Component → Instance drill-down paths
- Golden Ratio: 80% operational metrics, 20% exploratory metrics
- Cognitive Load: Maximum 7±2 panels per dashboard screen
- User Journey: Role-based dashboard personas (SRE, Developer, Executive)
Visualization Best Practices
- Chart Selection: Time series for trends, heatmaps for distributions, gauges for status
- Color Theory: Red for critical, amber for warning, green for healthy states
- Reference Lines: SLO targets, capacity thresholds, and historical baselines
- Time Ranges: Default to meaningful windows (4h for incidents, 7d for trends)
Panel Design
- Metric Queries: Efficient Prometheus/InfluxDB queries with proper aggregation
- Alerting Integration: Visual alert state indicators on relevant panels
- Interactive Elements: Template variables, drill-down links, and annotation overlays
- Performance: Sub-second render times through query optimization
Alert Design and Optimization
Alert Classification
- Severity Levels:
- Critical: Service down, SLO burn rate high
- Warning: Approaching thresholds, non-user-facing issues
- Info: Deployment notifications, capacity planning alerts
- Actionability: Every alert must have a clear response action
- Alert Routing: Escalation policies based on severity and team ownership
Alert Fatigue Prevention
- Signal vs Noise: High precision (few false positives) over high recall
- Hysteresis: Different thresholds for firing and resolving alerts
- Suppression: Dependent alert suppression during known outages
- Grouping: Related alerts grouped into single notifications
Alert Rule Design
- Threshold Selection: Statistical methods for threshold determination
- Window Functions: Appropriate averaging windows and percentile calculations
- Alert Lifecycle: Clear firing conditions and automatic resolution criteria
- Testing: Alert rule validation against historical data
Runbook Generation and Incident Response
Runbook Structure
- Alert Context: What the alert means and why it fired
- Impact Assessment: User-facing vs internal impact evaluation
- Investigation Steps: Ordered troubleshooting procedures with time estimates
- Resolution Actions: Common fixes and escalation procedures
- Post-Incident: Follow-up tasks and prevention measures
Incident Detection Patterns
- Anomaly Detection: Statistical methods for detecting unusual patterns
- Composite Alerts: Multi-signal alerts for complex failure modes
- Predictive Alerts: Capacity and trend-based forward-looking alerts
- Canary Monitoring: Early detection through progressive deployment monitoring
Golden Signals Framework
Latency Monitoring
- Request Latency: P50, P95, P99 response time tracking
- Queue Latency: Time spent waiting in processing queues
- Network Latency: Inter-service communication delays
- Database Latency: Query execution and connection pool metrics
Traffic Monitoring
- Request Rate: Requests per second with burst detection
- Bandwidth Usage: Network throughput and capacity utilization
- User Sessions: Active user tracking and session duration
- Feature Usage: API endpoint and feature adoption metrics
Error Monitoring
- Error Rate: 4xx and 5xx HTTP response code tracking
- Error Budget: SLO-based error rate targets and consumption
- Error Distribution: Error type classification and trending
- Silent Failures: Detection of processing failures without HTTP errors
Saturation Monitoring
- Resource Utilization: CPU, memory, disk, and network usage
- Queue Depth: Processing queue length and wait times
- Connection Pools: Database and service connection saturation
- Rate Limiting: API throttling and quota exhaustion tracking
Distributed Tracing Strategies
Trace Architecture
- Sampling Strategy: Head-based, tail-based, and adaptive sampling
- Trace Propagation: Context propagation across service boundaries
- Span Correlation: Parent-child relationship modeling
- Trace Storage: Retention policies and storage optimization
Service Instrumentation
- Auto-Instrumentation: Framework-based automatic trace generation
- Manual Instrumentation: Custom span creation for business logic
- Baggage Handling: Cross-cutting concern propagation
- Performance Impact: Instrumentation overhead measurement and optimization
Log Aggregation Patterns
Collection Architecture
- Agent Deployment: Log shipping agent strategies (push vs pull)
- Log Routing: Topic-based routing and filtering
- Parsing Strategies: Structured vs unstructured log handling
- Schema Evolution: Log format versioning and migration
Storage and Indexing
- Index Design: Optimized field indexing for common query patterns
- Retention Policies: Time and volume-based log retention
- Compression: Log data compression and archival strategies
- Search Performance: Query optimization and result caching
Cost Optimization for Observability
Data Management
- Metric Retention: Tiered retention based on metric importance
- Log Sampling: Intelligent sampling to reduce ingestion costs
- Trace Sampling: Cost-effective trace collection strategies
- Data Archival: Cold storage for historical observability data
Resource Optimization
- Query Efficiency: Optimized metric and log queries
- Storage Costs: Appropriate storage tiers for different data types
- Ingestion Rate Limiting: Controlled data ingestion to manage costs
- Cardinality Management: High-cardinality metric detection and mitigation
Scripts Overview
This skill includes three powerful Python scripts for comprehensive observability design:
1. SLO Designer (slo_designer.py)
Generates complete SLI/SLO frameworks based on service characteristics:
- Input: Service description JSON (type, criticality, dependencies)
- Output: SLI definitions, SLO targets, error budgets, burn rate alerts, SLA recommendations
- Features: Multi-window burn rate calculations, error budget policies, alert rule generation
2. Alert Optimizer (alert_optimizer.py)
Analyzes and optimizes existing alert configurations:
- Input: Alert configuration JSON with rules, thresholds, and routing
- Output: Optimization report and improved alert configuration
- Features: Noise detection, coverage gaps, duplicate identification, threshold optimization
3. Dashboard Generator (dashboard_generator.py)
Creates comprehensive dashboard specifications:
- Input: Service/system description JSON
- Output: Grafana-compatible dashboard JSON and documentation
- Features: Golden signals coverage, RED/USE methods, drill-down paths, role-based views
Integration Patterns
Monitoring Stack Integration
- Prometheus: Metric collection and alerting rule generation
- Grafana: Dashboard creation and visualization configuration
- Elasticsearch/Kibana: Log analysis and dashboard integration
- Jaeger/Zipkin: Distributed tracing configuration and analysis
CI/CD Integration
- Pipeline Monitoring: Build, test, and deployment observability
- Deployment Correlation: Release impact tracking and rollback triggers
- Feature Flag Monitoring: A/B test and feature rollout observability
- Performance Regression: Automated performance monitoring in pipelines
Incident Management Integration
- PagerDuty/VictorOps: Alert routing and escalation policies
- Slack/Teams: Notification and collaboration integration
- JIRA/ServiceNow: Incident tracking and resolution workflows
- Post-Mortem: Automated incident analysis and improvement tracking
Advanced Patterns
Multi-Cloud Observability
- Cross-Cloud Metrics: Unified metrics across AWS, GCP, Azure
- Network Observability: Inter-cloud connectivity monitoring
- Cost Attribution: Cloud resource cost tracking and optimization
- Compliance Monitoring: Security and compliance posture tracking
Microservices Observability
- Service Mesh Integration: Istio/Linkerd observability configuration
- API Gateway Monitoring: Request routing and rate limiting observability
- Container Orchestration: Kubernetes cluster and workload monitoring
- Service Discovery: Dynamic service monitoring and health checks
Machine Learning Observability
- Model Performance: Accuracy, drift, and bias monitoring
- Feature Store Monitoring: Feature quality and freshness tracking
- Pipeline Observability: ML pipeline execution and performance monitoring
- A/B Test Analysis: Statistical significance and business impact measurement
Best Practices
Organizational Alignment
- SLO Setting: Collaborative target setting between product and engineering
- Alert Ownership: Clear escalation paths and team responsibilities
- Dashboard Governance: Centralized dashboard management and standards
- Training Programs: Team education on observability tools and practices
Technical Excellence
- Infrastructure as Code: Observability configuration version control
- Testing Strategy: Alert rule testing and dashboard validation
- Performance Monitoring: Observability system performance tracking
- Security Considerations: Access control and data privacy in observability
Continuous Improvement
- Metrics Review: Regular SLI/SLO effectiveness assessment
- Alert Tuning: Ongoing alert threshold and routing optimization
- Dashboard Evolution: User feedback-driven dashboard improvements
- Tool Evaluation: Regular assessment of observability tool effectiveness
1---2name: observability-designer3description: Design production-ready observability strategies combining metrics, logs, and traces. Includes SLI/SLO design, golden-signals monitoring, alert optimization. Use when adding observability to a new service, refactoring alerting that is too noisy, or designing an SLO program before scaling production load.4---56# Observability Designer (POWERFUL)78**Category:** Engineering 9**Tier:** POWERFUL 10**Description:** Design comprehensive observability strategies for production systems including SLI/SLO frameworks, alerting optimization, and dashboard generation.1112## Overview1314Observability Designer creates production-ready dashboards, alert configurations, and monitoring strategies across the three pillars (metrics, logs, traces).1516**When NOT to use → slo-architect.** For SLO/SLI design with error-budget math, multi-window burn-rate alerting thresholds, and SLO review gates, route to `slo-architect` — it is the authoritative skill for that half. This skill's `slo_designer.py` produces a quick scaffold only. This skill's lane: dashboards (`dashboard_generator.py`) and alert-noise reduction (`alert_optimizer.py`).1718## Quick Start1920```bash21# Dashboard spec (Grafana JSON + docs) for a service22python3 scripts/dashboard_generator.py --service-type api --name payments --criticality critical --role sre --format grafana -o dashboard.json --doc-output dashboard.md2324# Analyze an existing alert config for noise, duplicates, and coverage gaps25python3 scripts/alert_optimizer.py --input alerts.json --analyze-only --report alert_report.json26# ...then emit the optimized config once the report is reviewed:27python3 scripts/alert_optimizer.py --input alerts.json --output alerts_optimized.json2829# Quick SLO scaffold (hand off to slo-architect for the real error-budget work)30python3 scripts/slo_designer.py --service-type api --criticality high --user-facing true --service-name payments -o slo_scaffold.json31```3233**Verification loop:** after deploying optimized alerts, track the report's noise metrics for one on-call rotation — if the actionable-alert ratio didn't improve, re-run `--analyze-only` against the live config and iterate. Import the generated dashboard into Grafana and confirm every golden-signal panel renders with live data before closing the task.3435## Core Competencies3637### SLI/SLO/SLA Framework Design38- **Service Level Indicators (SLI):** Define measurable signals that indicate service health39- **Service Level Objectives (SLO):** Set reliability targets based on user experience40- **Service Level Agreements (SLA):** Establish customer-facing commitments with consequences41- **Error Budget Management:** Calculate and track error budget consumption42- **Burn Rate Alerting:** Multi-window burn rate alerts for proactive SLO protection4344### Three Pillars of Observability4546#### Metrics47- **Golden Signals:** Latency, traffic, errors, and saturation monitoring48- **RED Method:** Rate, Errors, and Duration for request-driven services49- **USE Method:** Utilization, Saturation, and Errors for resource monitoring50- **Business Metrics:** Revenue, user engagement, and feature adoption tracking51- **Infrastructure Metrics:** CPU, memory, disk, network, and custom resource metrics5253#### Logs54- **Structured Logging:** JSON-based log formats with consistent fields55- **Log Aggregation:** Centralized log collection and indexing strategies56- **Log Levels:** Appropriate use of DEBUG, INFO, WARN, ERROR, FATAL levels57- **Correlation IDs:** Request tracing through distributed systems58- **Log Sampling:** Volume management for high-throughput systems5960#### Traces61- **Distributed Tracing:** End-to-end request flow visualization62- **Span Design:** Meaningful span boundaries and metadata63- **Trace Sampling:** Intelligent sampling strategies for performance and cost64- **Service Maps:** Automatic dependency discovery through traces65- **Root Cause Analysis:** Trace-driven debugging workflows6667### Dashboard Design Principles6869#### Information Architecture70- **Hierarchy:** Overview → Service → Component → Instance drill-down paths71- **Golden Ratio:** 80% operational metrics, 20% exploratory metrics72- **Cognitive Load:** Maximum 7±2 panels per dashboard screen73- **User Journey:** Role-based dashboard personas (SRE, Developer, Executive)7475#### Visualization Best Practices76- **Chart Selection:** Time series for trends, heatmaps for distributions, gauges for status77- **Color Theory:** Red for critical, amber for warning, green for healthy states78- **Reference Lines:** SLO targets, capacity thresholds, and historical baselines79- **Time Ranges:** Default to meaningful windows (4h for incidents, 7d for trends)8081#### Panel Design82- **Metric Queries:** Efficient Prometheus/InfluxDB queries with proper aggregation83- **Alerting Integration:** Visual alert state indicators on relevant panels84- **Interactive Elements:** Template variables, drill-down links, and annotation overlays85- **Performance:** Sub-second render times through query optimization8687### Alert Design and Optimization8889#### Alert Classification90- **Severity Levels:** 91 - **Critical:** Service down, SLO burn rate high92 - **Warning:** Approaching thresholds, non-user-facing issues93 - **Info:** Deployment notifications, capacity planning alerts94- **Actionability:** Every alert must have a clear response action95- **Alert Routing:** Escalation policies based on severity and team ownership9697#### Alert Fatigue Prevention98- **Signal vs Noise:** High precision (few false positives) over high recall99- **Hysteresis:** Different thresholds for firing and resolving alerts100- **Suppression:** Dependent alert suppression during known outages101- **Grouping:** Related alerts grouped into single notifications102103#### Alert Rule Design104- **Threshold Selection:** Statistical methods for threshold determination105- **Window Functions:** Appropriate averaging windows and percentile calculations106- **Alert Lifecycle:** Clear firing conditions and automatic resolution criteria107- **Testing:** Alert rule validation against historical data108109### Runbook Generation and Incident Response110111#### Runbook Structure112- **Alert Context:** What the alert means and why it fired113- **Impact Assessment:** User-facing vs internal impact evaluation114- **Investigation Steps:** Ordered troubleshooting procedures with time estimates115- **Resolution Actions:** Common fixes and escalation procedures116- **Post-Incident:** Follow-up tasks and prevention measures117118#### Incident Detection Patterns119- **Anomaly Detection:** Statistical methods for detecting unusual patterns120- **Composite Alerts:** Multi-signal alerts for complex failure modes121- **Predictive Alerts:** Capacity and trend-based forward-looking alerts122- **Canary Monitoring:** Early detection through progressive deployment monitoring123124### Golden Signals Framework125126#### Latency Monitoring127- **Request Latency:** P50, P95, P99 response time tracking128- **Queue Latency:** Time spent waiting in processing queues129- **Network Latency:** Inter-service communication delays130- **Database Latency:** Query execution and connection pool metrics131132#### Traffic Monitoring133- **Request Rate:** Requests per second with burst detection134- **Bandwidth Usage:** Network throughput and capacity utilization135- **User Sessions:** Active user tracking and session duration136- **Feature Usage:** API endpoint and feature adoption metrics137138#### Error Monitoring139- **Error Rate:** 4xx and 5xx HTTP response code tracking140- **Error Budget:** SLO-based error rate targets and consumption141- **Error Distribution:** Error type classification and trending142- **Silent Failures:** Detection of processing failures without HTTP errors143144#### Saturation Monitoring145- **Resource Utilization:** CPU, memory, disk, and network usage146- **Queue Depth:** Processing queue length and wait times147- **Connection Pools:** Database and service connection saturation148- **Rate Limiting:** API throttling and quota exhaustion tracking149150### Distributed Tracing Strategies151152#### Trace Architecture153- **Sampling Strategy:** Head-based, tail-based, and adaptive sampling154- **Trace Propagation:** Context propagation across service boundaries155- **Span Correlation:** Parent-child relationship modeling156- **Trace Storage:** Retention policies and storage optimization157158#### Service Instrumentation159- **Auto-Instrumentation:** Framework-based automatic trace generation160- **Manual Instrumentation:** Custom span creation for business logic161- **Baggage Handling:** Cross-cutting concern propagation162- **Performance Impact:** Instrumentation overhead measurement and optimization163164### Log Aggregation Patterns165166#### Collection Architecture167- **Agent Deployment:** Log shipping agent strategies (push vs pull)168- **Log Routing:** Topic-based routing and filtering169- **Parsing Strategies:** Structured vs unstructured log handling170- **Schema Evolution:** Log format versioning and migration171172#### Storage and Indexing173- **Index Design:** Optimized field indexing for common query patterns174- **Retention Policies:** Time and volume-based log retention175- **Compression:** Log data compression and archival strategies176- **Search Performance:** Query optimization and result caching177178### Cost Optimization for Observability179180#### Data Management181- **Metric Retention:** Tiered retention based on metric importance182- **Log Sampling:** Intelligent sampling to reduce ingestion costs183- **Trace Sampling:** Cost-effective trace collection strategies184- **Data Archival:** Cold storage for historical observability data185186#### Resource Optimization187- **Query Efficiency:** Optimized metric and log queries188- **Storage Costs:** Appropriate storage tiers for different data types189- **Ingestion Rate Limiting:** Controlled data ingestion to manage costs190- **Cardinality Management:** High-cardinality metric detection and mitigation191192## Scripts Overview193194This skill includes three powerful Python scripts for comprehensive observability design:195196### 1. SLO Designer (`slo_designer.py`)197Generates complete SLI/SLO frameworks based on service characteristics:198- **Input:** Service description JSON (type, criticality, dependencies)199- **Output:** SLI definitions, SLO targets, error budgets, burn rate alerts, SLA recommendations200- **Features:** Multi-window burn rate calculations, error budget policies, alert rule generation201202### 2. Alert Optimizer (`alert_optimizer.py`)203Analyzes and optimizes existing alert configurations:204- **Input:** Alert configuration JSON with rules, thresholds, and routing205- **Output:** Optimization report and improved alert configuration206- **Features:** Noise detection, coverage gaps, duplicate identification, threshold optimization207208### 3. Dashboard Generator (`dashboard_generator.py`)209Creates comprehensive dashboard specifications:210- **Input:** Service/system description JSON211- **Output:** Grafana-compatible dashboard JSON and documentation212- **Features:** Golden signals coverage, RED/USE methods, drill-down paths, role-based views213214## Integration Patterns215216### Monitoring Stack Integration217- **Prometheus:** Metric collection and alerting rule generation218- **Grafana:** Dashboard creation and visualization configuration219- **Elasticsearch/Kibana:** Log analysis and dashboard integration220- **Jaeger/Zipkin:** Distributed tracing configuration and analysis221222### CI/CD Integration223- **Pipeline Monitoring:** Build, test, and deployment observability224- **Deployment Correlation:** Release impact tracking and rollback triggers225- **Feature Flag Monitoring:** A/B test and feature rollout observability226- **Performance Regression:** Automated performance monitoring in pipelines227228### Incident Management Integration229- **PagerDuty/VictorOps:** Alert routing and escalation policies230- **Slack/Teams:** Notification and collaboration integration231- **JIRA/ServiceNow:** Incident tracking and resolution workflows232- **Post-Mortem:** Automated incident analysis and improvement tracking233234## Advanced Patterns235236### Multi-Cloud Observability237- **Cross-Cloud Metrics:** Unified metrics across AWS, GCP, Azure238- **Network Observability:** Inter-cloud connectivity monitoring239- **Cost Attribution:** Cloud resource cost tracking and optimization240- **Compliance Monitoring:** Security and compliance posture tracking241242### Microservices Observability243- **Service Mesh Integration:** Istio/Linkerd observability configuration244- **API Gateway Monitoring:** Request routing and rate limiting observability245- **Container Orchestration:** Kubernetes cluster and workload monitoring246- **Service Discovery:** Dynamic service monitoring and health checks247248### Machine Learning Observability249- **Model Performance:** Accuracy, drift, and bias monitoring250- **Feature Store Monitoring:** Feature quality and freshness tracking251- **Pipeline Observability:** ML pipeline execution and performance monitoring252- **A/B Test Analysis:** Statistical significance and business impact measurement253254## Best Practices255256### Organizational Alignment257- **SLO Setting:** Collaborative target setting between product and engineering258- **Alert Ownership:** Clear escalation paths and team responsibilities259- **Dashboard Governance:** Centralized dashboard management and standards260- **Training Programs:** Team education on observability tools and practices261262### Technical Excellence263- **Infrastructure as Code:** Observability configuration version control264- **Testing Strategy:** Alert rule testing and dashboard validation265- **Performance Monitoring:** Observability system performance tracking266- **Security Considerations:** Access control and data privacy in observability267268### Continuous Improvement269- **Metrics Review:** Regular SLI/SLO effectiveness assessment270- **Alert Tuning:** Ongoing alert threshold and routing optimization271- **Dashboard Evolution:** User feedback-driven dashboard improvements272- **Tool Evaluation:** Regular assessment of observability tool effectiveness