Monitoring & Observability Agent Skills

Monitoring & Observability

234 skills
cloudthinker-ai
aws-waf
Analyzes AWS WAF web ACLs, rules, IP sets, and logging configurations, and retrieves blocked/allowed request metrics from CloudWatch.
7
cloudthinker-ai
aws-lambda
Analyzes AWS Lambda functions, including configuration, invocation metrics, cold starts, error rates, layers, and concurrency, producing structured reports.
7
cloudthinker-ai
aws-rds-deep
Deep-dive analysis of AWS RDS instances using Performance Insights, event subscriptions, proxy health, and global database status, with parallel execution and anti-hallucination guardrails.
7
cloudthinker-ai
managing-dbt
Manages and monitors dbt projects, model runs, and test results via dbt CLI and dbt Cloud API, covering run status, test failures, source freshness, and manifest analysis.
7
cloudthinker-ai
managing-tyk
Manage Tyk API gateway infrastructure by inspecting API definitions, policies, keys, and analytics through the Gateway and Dashboard APIs.
7
cloudthinker-ai
load-test-plan
Designs and executes load tests, covering scenario design, baseline capture, execution configuration, results analysis, and reporting for k6, Locust, Gatling, and JMeter.
7
cloudthinker-ai
azure-monitor
Queries Azure Monitor metrics, manages Log Analytics workspaces, audits alert rules and diagnostic settings, and configures action groups via Azure CLI.
7
cloudthinker-ai
managing-ably
Manages and analyzes Ably real-time messaging resources, covering channels, presence, connections, usage analytics, and account health via REST and Control APIs.
7
cloudthinker-ai
managing-etcd
Monitors and analyzes etcd clusters: checks endpoint health, member status, alarms, key-space usage, leases, and Raft consensus state, then reports issues and recommendations.
7
cloudthinker-ai
managing-grpc
Inspects gRPC services via reflection, health checks, channelz, and grpcurl to discover methods, monitor health, and analyze load balancing.
7
cloudthinker-ai
managing-kuma
Manage Kuma service mesh by discovering meshes, dataplanes, and policies via the Kuma API, then analyzing configuration and producing structured reports.
7
cloudthinker-ai
managing-nats
Analyze NATS clusters with read-only operations: discover JetStream streams and consumers, check cluster health, and inspect message flow.
7
cloudthinker-ai
managing-prtg
Queries and analyzes PRTG Network Monitor data via its HTTP API, covering sensor status, device trees, alerts, and bandwidth metrics.
7
cloudthinker-ai
aws-cloudfront
Analyzes AWS CloudFront distributions, including cache hit ratios, origin health, SSL certificate status, and invalidation history, producing structured reports.
7
cloudthinker-ai
aws-cloudtrail
Analyzes AWS CloudTrail events and trails, including trail health, API activity, security investigations, resource changes, and event selector audits, with parallel execution and anti-hallucination guardrails.
7
cloudthinker-ai
managing-100ms
Manages and analyzes 100ms live video/audio rooms, sessions, recordings, and usage via the management API, including discovery, analytics, and troubleshooting.
7
cloudthinker-ai
managing-dokku
Audits self-hosted Dokku PaaS deployments by inventorying apps, containers, domains, plugins, storage mounts, environment variable counts, proxy settings, and SSL status via read-only SSH commands.
7
tools-only
050-bash-9160b3bf
Optimize Kubernetes cluster costs by right-sizing resources, configuring autoscalers, using spot instances, and monitoring spend.
7 · bundle
johnalbertini14-glitch
idrac
Monitors and manages Dell PowerEdge servers via the iDRAC Redfish API, covering hardware status, health, sensors, power operations, inventory, and event logs.
1 · bundle
johnalbertini14-glitch
gohome
Operate and validate a GoHome service via gRPC discovery, read-only RPC calls, and Prometheus metrics exposed over HTTP.
1 · bundle
lucaspmarie-a11y
langfuse
Instruments LLM applications with Langfuse for tracing, observability, and evaluation, covering setup, OpenAI and LangChain integrations, and best practices.
5
lucaspmarie-a11y
manifest
Installs and configures the Manifest observability plugin for AI agents, including API key setup, endpoint configuration, and connection verification.
5
qhjqhj00
phoenix-observability
Self-hosted observability platform for LLM applications, providing tracing, evaluation, datasets, experiments, and real-time monitoring to debug and improve AI systems.
3 · bundle
akillness
log-analysis
Routes runtime-log requests into an evidence packet to isolate the first actionable blocker, repeated signature, blast radius, or safest next read-only check.
42 · bundle
snoodleboot-io
infrastructure-drift-detection
Detect and triage infrastructure drift by comparing declared Terraform state against live cloud resources using scheduled pipelines and audit logs.
2
addyosmani
observability-and-instrumentation
Adds logging, metrics, tracing, and alerting to make production behavior visible and diagnosable.
69.5k
antigravity
manifest
Sets up the Manifest observability plugin for AI agents, including installation, API key configuration, and verification.
42.4k
github
phoenix-tracing
Instrument LLM applications with OpenInference tracing for Phoenix AI observability, covering setup, custom spans, and production deployment.
36.2k · bundle
github
qdrant-monitoring
Guides monitoring and observability setup for Qdrant vector search deployments, including Prometheus scraping, health checks, and metric-based debugging of production issues.
36.2k
github
debian-linux-triage
Diagnose and resolve Debian Linux issues using apt, systemd, and AppArmor-aware guidance.
36.2k
github
arize-instrumentation
Adds Arize AX tracing to LLM applications using a two-phase agent-assisted flow that analyzes the codebase before implementing instrumentation.
36.2k · bundle
github
qdrant-monitoring-debugging
Diagnoses Qdrant production issues using metrics and observability tools, covering optimizer problems, memory spikes, and slow queries.
36.2k
jeffallan
sre-engineer
Defines service level objectives, creates error budget policies, designs incident response procedures, develops capacity models, and produces monitoring configurations and automation scripts for production systems.
10.4k · bundle
expo
eas-observe
Add EAS Observe to an Expo project for startup, navigation, and custom-event performance metrics, query them via the EAS CLI, and interpret the results.
2.2k · bundle
iamcorey
wifi-optimizer
Diagnose intermittent Wi-Fi issues like buffering, lag, packet loss, and weak coverage through read-only analysis, then safely optimize authorized router settings when evidence supports a change.
53 · bundle
agentskillexchange
weights-biases-run-monitor
Streams live training metrics, system stats, and gradients from active W&B runs, alerts on metric regressions, and posts summaries to Slack.
28