Monitoring & Observability Agent Skills
Monitoring & Observability
234 skillsslo-architect
Design, review, and operate SLOs, SLIs, and error budgets with multi-window burn-rate alerts and automated validation.
20.4k · bundle
vpe-advisor
Analyze engineering delivery throughput, hiring funnel health, team structure, and production discipline for startup VPEs and founders.
20.4k · bundle
chaos-engineering
Design, run, and learn from safe chaos engineering experiments with structured plans, blast radius calculation, and postmortem generation.
20.4k · bundle
runbook-generator
Generate operational runbooks from a service name — deployment, incident response, maintenance, and rollback workflows. Templated structure customizable per environment.
20.4k · bundle
performance-profiler
Systematically profile Node.js, Python, and Go applications to identify CPU, memory, and I/O bottlenecks, generate flamegraphs, analyze bundle sizes, optimize database queries, and run load tests with k6 and Artillery.
20.4k · bundle
observability-designer
Design production-ready observability strategies combining metrics, logs, and traces, including SLI/SLO design, golden-signals monitoring, and alert optimization.
20.4k · bundle
google-cloud-waf-reliability
Evaluates Google Cloud workloads against the Reliability pillar of the Well-Architected Framework, providing actionable recommendations for building, deploying, and managing reliable systems.
14.4k
agent-platform-alert-configuration
Configures dynamic threshold alerting policies for Google Cloud Vertex AI Agent Platform agents, monitoring latency, error rates, and quality metrics using Terraform and PromQL.
14.4k · bundle
google-cloud-networking-observability
Investigates Google Cloud networking issues by analyzing logs, metrics, and diagnostics, including VPC Flow Logs, NAT, firewall, threat logs, latency, throughput, and Connectivity Tests.
14.4k · bundle
google-cloud-waf-operational-excellence
Generates operations-focused guidance for Google Cloud workloads based on the Operational Excellence pillar of the Well-Architected Framework, including assessment questions and validation checklists.
14.4k
debugview
Captures and analyzes Windows debug output (OutputDebugString, DbgPrint/KdPrint) from the command line, with filtering, bounded execution, and remote monitoring.
2.7k · bundle
azure-diagnostics
Debug and troubleshoot Azure production issues using AppLens, Azure Monitor, resource health, and systematic diagnosis flows for services like App Service, Functions, AKS, Container Apps, and Messaging.
2.7k · bundle
azure-monitor-query-py
Query logs and metrics from Azure Monitor and Log Analytics workspaces using the Python SDK.
2.7k
appinsights-instrumentation
Provides guidance and reference material for instrumenting web applications with Azure Application Insights, including telemetry patterns, SDK setup, and configuration references.
2.7k · bundle
azure-monitor-ingestion-java
Send custom logs to Azure Monitor via Data Collection Rules and Data Collection Endpoints using the Java SDK.
2.7k · bundle
azure-monitor-ingestion-py
Send custom logs to Azure Monitor Log Analytics workspace using the Logs Ingestion API and Python SDK.
2.7k
azure-monitor-opentelemetry-py
Configures Azure Monitor Application Insights with OpenTelemetry auto-instrumentation for Python applications in one line.
2.7k
azure-monitor-opentelemetry-ts
Auto-instrument Node.js applications with distributed tracing, metrics, and logs using Azure Monitor and OpenTelemetry.
2.7k
azure-monitor-opentelemetry-exporter-java
Export OpenTelemetry traces, metrics, and logs to Azure Monitor/Application Insights using the deprecated exporter or the recommended autoconfigure package.
2.7k · bundle
azure-monitor-opentelemetry-exporter-py
Export OpenTelemetry traces, metrics, and logs to Azure Application Insights using Python.
2.7k
cloudwatch
Monitor AWS resources and applications using CloudWatch logs, metrics, alarms, and dashboards with CLI and boto3 examples.
1.1k · bundle
rag-perf
Run config-driven performance benchmarks against a deployed NVIDIA RAG Blueprint server, including profiling and load testing, with a unified report.
2.2k · bundle
jetson-diagnostic
Captures a read-only health snapshot from a Jetson device, reporting identity, memory, GPU, thermal, power, storage, services, and top processes.
2.2k · bundle
vss-manage-alerts
Operate the VSS alert pipeline for real-time monitoring, Alert-Bridge subscriptions, Slack notifications, incident queries, and camera onboarding.
2.2k · bundle
dynamo-troubleshoot
Diagnose failed or unhealthy Dynamo deployments by collecting a read-only debug bundle, classifying failures, and providing step-by-step remediation guidance.
2.2k · bundle
jetson-memory-audit
Measure Jetson DRAM and NvMap usage, capture before/after baselines, and verify memory reclamation with live audit data.
2.2k · bundle
browser-trace
Capture a full DevTools-protocol trace of any browser automation, bisect the stream into per-page searchable buckets, and attach a trace to an in-progress session for debugging.
3.6k · bundle
canary-watch
Monitors a deployed URL for regressions after releases by checking HTTP status, console errors, network failures, performance metrics, content integrity, API health, static assets, and SSE streams.
226k
context-budget
Audits Claude Code context window consumption across agents, skills, MCP servers, and rules. Identifies bloat, redundant components, and produces prioritized token-savings recommendations.
226k
dashboard-builder
Build monitoring dashboards that answer real operator questions for Grafana, SigNoz, and similar platforms, turning metrics into actionable operational views.
226k
automation-audit-ops
Produces an evidence-backed inventory of live automations (jobs, hooks, connectors, MCP servers, wrappers) and recommends keep/merge/cut/fix-next actions before rewriting anything.
226k
ecc-tools-cost-audit
Audits ECC Tools GitHub App for cost issues like runaway PR creation, quota bypass, premium-model leakage, and duplicate jobs, using an evidence-first workflow.
226k
enterprise-agent-ops
Operate long-lived agent workloads with observability, security boundaries, and lifecycle management.
226k
langfuse
Provides expertise in Langfuse for LLM observability, including tracing, prompt management, evaluation, and integration with LangChain, LlamaIndex, and OpenAI.
42.4k
sentry
Inspect Sentry issues, events, and production errors using the Sentry CLI for read-only observability queries.
23.3k · bundle
dump-collect
Configure and collect crash dumps for modern .NET applications (CoreCLR and NativeAOT) on Linux, macOS, and Windows, including containers.
4k · bundle