Monitoring & Observability Agent Skills
Monitoring & Observability
234 skillsaws-waf
Analyzes AWS WAF web ACLs, rules, IP sets, and logging configurations, and retrieves blocked/allowed request metrics from CloudWatch.
7
aws-lambda
Analyzes AWS Lambda functions, including configuration, invocation metrics, cold starts, error rates, layers, and concurrency, producing structured reports.
7
aws-rds-deep
Deep-dive analysis of AWS RDS instances using Performance Insights, event subscriptions, proxy health, and global database status, with parallel execution and anti-hallucination guardrails.
7
managing-dbt
Manages and monitors dbt projects, model runs, and test results via dbt CLI and dbt Cloud API, covering run status, test failures, source freshness, and manifest analysis.
7
managing-tyk
Manage Tyk API gateway infrastructure by inspecting API definitions, policies, keys, and analytics through the Gateway and Dashboard APIs.
7
load-test-plan
Designs and executes load tests, covering scenario design, baseline capture, execution configuration, results analysis, and reporting for k6, Locust, Gatling, and JMeter.
7
azure-monitor
Queries Azure Monitor metrics, manages Log Analytics workspaces, audits alert rules and diagnostic settings, and configures action groups via Azure CLI.
7
managing-ably
Manages and analyzes Ably real-time messaging resources, covering channels, presence, connections, usage analytics, and account health via REST and Control APIs.
7
managing-etcd
Monitors and analyzes etcd clusters: checks endpoint health, member status, alarms, key-space usage, leases, and Raft consensus state, then reports issues and recommendations.
7
managing-grpc
Inspects gRPC services via reflection, health checks, channelz, and grpcurl to discover methods, monitor health, and analyze load balancing.
7
managing-kuma
Manage Kuma service mesh by discovering meshes, dataplanes, and policies via the Kuma API, then analyzing configuration and producing structured reports.
7
managing-nats
Analyze NATS clusters with read-only operations: discover JetStream streams and consumers, check cluster health, and inspect message flow.
7
managing-prtg
Queries and analyzes PRTG Network Monitor data via its HTTP API, covering sensor status, device trees, alerts, and bandwidth metrics.
7
aws-cloudfront
Analyzes AWS CloudFront distributions, including cache hit ratios, origin health, SSL certificate status, and invalidation history, producing structured reports.
7
aws-cloudtrail
Analyzes AWS CloudTrail events and trails, including trail health, API activity, security investigations, resource changes, and event selector audits, with parallel execution and anti-hallucination guardrails.
7
managing-100ms
Manages and analyzes 100ms live video/audio rooms, sessions, recordings, and usage via the management API, including discovery, analytics, and troubleshooting.
7
managing-dokku
Audits self-hosted Dokku PaaS deployments by inventorying apps, containers, domains, plugins, storage mounts, environment variable counts, proxy settings, and SSL status via read-only SSH commands.
7
050-bash-9160b3bf
Optimize Kubernetes cluster costs by right-sizing resources, configuring autoscalers, using spot instances, and monitoring spend.
7 · bundle
idrac
Monitors and manages Dell PowerEdge servers via the iDRAC Redfish API, covering hardware status, health, sensors, power operations, inventory, and event logs.
1 · bundle
gohome
Operate and validate a GoHome service via gRPC discovery, read-only RPC calls, and Prometheus metrics exposed over HTTP.
1 · bundle
langfuse
Instruments LLM applications with Langfuse for tracing, observability, and evaluation, covering setup, OpenAI and LangChain integrations, and best practices.
5
manifest
Installs and configures the Manifest observability plugin for AI agents, including API key setup, endpoint configuration, and connection verification.
5
phoenix-observability
Self-hosted observability platform for LLM applications, providing tracing, evaluation, datasets, experiments, and real-time monitoring to debug and improve AI systems.
3 · bundle
log-analysis
Routes runtime-log requests into an evidence packet to isolate the first actionable blocker, repeated signature, blast radius, or safest next read-only check.
42 · bundle
infrastructure-drift-detection
Detect and triage infrastructure drift by comparing declared Terraform state against live cloud resources using scheduled pipelines and audit logs.
2
observability-and-instrumentation
Adds logging, metrics, tracing, and alerting to make production behavior visible and diagnosable.
69.5k
manifest
Sets up the Manifest observability plugin for AI agents, including installation, API key configuration, and verification.
42.4k
phoenix-tracing
Instrument LLM applications with OpenInference tracing for Phoenix AI observability, covering setup, custom spans, and production deployment.
36.2k · bundle
qdrant-monitoring
Guides monitoring and observability setup for Qdrant vector search deployments, including Prometheus scraping, health checks, and metric-based debugging of production issues.
36.2k
debian-linux-triage
Diagnose and resolve Debian Linux issues using apt, systemd, and AppArmor-aware guidance.
36.2k
arize-instrumentation
Adds Arize AX tracing to LLM applications using a two-phase agent-assisted flow that analyzes the codebase before implementing instrumentation.
36.2k · bundle
qdrant-monitoring-debugging
Diagnoses Qdrant production issues using metrics and observability tools, covering optimizer problems, memory spikes, and slow queries.
36.2k
sre-engineer
Defines service level objectives, creates error budget policies, designs incident response procedures, develops capacity models, and produces monitoring configurations and automation scripts for production systems.
10.4k · bundle
eas-observe
Add EAS Observe to an Expo project for startup, navigation, and custom-event performance metrics, query them via the EAS CLI, and interpret the results.
2.2k · bundle
wifi-optimizer
Diagnose intermittent Wi-Fi issues like buffering, lag, packet loss, and weak coverage through read-only analysis, then safely optimize authorized router settings when evidence supports a change.
53 · bundle
weights-biases-run-monitor
Streams live training metrics, system stats, and gradients from active W&B runs, alerts on metric regressions, and posts summaries to Slack.
28