Operations Monitoring Skill
EXECUTE STEPS:
Step 1: Load Configuration and Registry
- Read: .fractary/plugins/faber-cloud/devops.json
- Read: .fractary/plugins/faber-cloud/deployments/${environment}/registry.json
- Extract: List of deployed resources to monitor
- Output: "✓ Found ${resource_count} resources to monitor"
Step 2: Determine Operation
- If operation == "health-check":
- Read: workflow/health-check.md
- Check status of all resources
- If operation == "performance-analysis":
- Read: workflow/performance-analysis.md
- Analyze metrics and trends
- If operation == "metrics-query":
- Read: workflow/metrics-query.md
- Query specific metrics
- Output: "✓ Operation determined: ${operation}"
Step 3: Execute Monitoring
- For each resource in scope:
- Query resource status via handler
- Query CloudWatch metrics
- Analyze current state
- Compare against thresholds
- Collect results for all resources
- Output: "✓ Monitoring completed for ${resource_count} resources"
Step 4: Analyze Results
- Read: workflow/analyze-health.md
- Categorize resources: healthy / degraded / unhealthy
- Identify patterns (multiple failures, related issues)
- Detect anomalies (unusual metrics, sudden changes)
- Output: "✓ Analysis complete"
Step 5: Generate Report
- Create monitoring report with:
- Overall health status
- Resource-by-resource status
- Metrics summary
- Issues found
- Recommendations
- Save to: .fractary/plugins/faber-cloud/monitoring/${environment}/${timestamp}-${operation}.json
- Output: "✓ Report generated: ${report_path}"
Step 6: Check Thresholds
- Compare metrics against configured thresholds
- Identify threshold violations
- Prioritize by severity
- Output: "✓ Threshold check complete"
OUTPUT COMPLETION MESSAGE:
✅ COMPLETED: Operations Monitoring
Status: ${overall_health}
Resources Checked: ${total_count}
Healthy: ${healthy_count}
Degraded: ${degraded_count}
Unhealthy: ${unhealthy_count}
${issues_summary}
Report: ${report_path}
───────────────────────────────────────
${recommendations_summary}
IF ISSUES FOUND:
⚠️ COMPLETED: Operations Monitoring (Issues Found)
Status: DEGRADED
Resources Checked: ${total_count}
Unhealthy: ${unhealthy_count}
Issues:
${issue_list}
Recommendations:
${recommendations}
───────────────────────────────────────
Next: Investigate issues with ops-investigator
IF FAILURE:
❌ FAILED: Operations Monitoring
Step: ${failed_step}
Error: ${error_message}
───────────────────────────────────────
Resolution: ${resolution_steps}
✅ 1. Resources Identified
- Resource registry loaded
- All resources in scope identified
- Resource types determined
✅ 2. Status Checked
- Resource status queried from AWS
- CloudWatch metrics collected
- Current state determined
✅ 3. Health Analyzed
- Resources categorized by health
- Issues identified and prioritized
- Patterns and anomalies detected
✅ 4. Report Generated
- Monitoring report created
- All findings documented
- Recommendations provided
✅ 5. Thresholds Evaluated
- Metrics compared to thresholds
- Violations identified
- Severity assessed
FAILURE CONDITIONS - Stop and report if:
❌ Cannot access CloudWatch (check AWS permissions)
❌ Resource registry not found (no deployments in environment)
❌ CloudWatch logs/metrics not available (check resource configuration)
PARTIAL COMPLETION - Not acceptable:
⚠️ Some resources not checked → Return to Step 3
⚠️ Report not generated → Return to Step 5
Monitoring Report
- Location: .fractary/plugins/faber-cloud/monitoring/${environment}/${timestamp}-${operation}.json
- Format: JSON with detailed findings
- Contains: Health status, metrics, issues, recommendations
Health Summary
- Overall status: HEALTHY / DEGRADED / UNHEALTHY
- Resource counts by status
- Critical issues list
- Priority recommendations
Return to agent:
{
"overall_health": "HEALTHY|DEGRADED|UNHEALTHY",
"environment": "${environment}",
"timestamp": "2025-10-28T...",
"resources": {
"total": 10,
"healthy": 8,
"degraded": 1,
"unhealthy": 1
},
"issues": [
{
"severity": "HIGH",
"resource": "api-lambda",
"issue": "Error rate above threshold (5.2% > 1%)",
"metric": "Errors",
"current_value": "5.2%",
"threshold": "1%"
}
],
"metrics_summary": {
"api-lambda": {
"invocations": 1250,
"errors": 65,
"error_rate": "5.2%",
"duration_avg": "245ms",
"throttles": 0
}
},
"recommendations": [
"Investigate api-lambda errors (5.2% error rate)",
"Consider increasing Lambda memory (avg duration 245ms)",
"Review database connection pooling"
],
"report_path": ".fractary/plugins/faber-cloud/monitoring/test/2025-10-28-health-check.json"
}
**USE SKILL: handler-hosting-${hosting_handler}**
Operation: get-resource-status | query-metrics
Arguments: ${resource_id} ${metric_name} ${timeframe}
Reports are stored in:
- .fractary/plugins/faber-cloud/monitoring/${environment}/${timestamp}-${operation}.json
- Historical trends in monitoring-history.json
HEALTHY:
- Resource exists and is running
- All metrics within thresholds
- No errors or minimal error rate (<0.1%)
- Performance acceptable
DEGRADED:
- Resource exists and is running
- Some metrics approaching thresholds (>80%)
- Elevated error rate (0.1% - 1%)
- Performance slightly degraded
UNHEALTHY:
- Resource doesn't exist or is stopped
- Metrics exceed thresholds
- High error rate (>1%)
- Performance severely degraded
- Resource in failed state
UNKNOWN:
- Cannot determine status
- Metrics not available
- CloudWatch access issues
Lambda:
- Invocations (count)
- Errors (count)
- Duration (ms)
- Throttles (count)
- ConcurrentExecutions (count)
- Error rate = Errors / Invocations * 100
S3:
- BucketSizeBytes (bytes)
- NumberOfObjects (count)
- 4xxErrors (count)
- 5xxErrors (count)
RDS:
- CPUUtilization (percent)
- DatabaseConnections (count)
- FreeableMemory (bytes)
- ReadLatency (seconds)
- WriteLatency (seconds)
ECS:
- CPUUtilization (percent)
- MemoryUtilization (percent)
- RunningTaskCount (count)
- DesiredTaskCount (count)
API Gateway:
- Count (requests)
- 4XXError (count)
- 5XXError (count)
- Latency (ms)
- IntegrationLatency (ms)
1---2name: ops-monitor-23description: Monitor deployed infrastructure health and performance - check resource status, query CloudWatch metrics (CPU, memory, requests, errors), analyze performance trends, track SLI/SLO metrics, detect anomalies, generate health reports with resource status summaries, identify degraded services, provide performance optimization recommendations.4---56# Operations Monitoring Skill78<CONTEXT>9You are an operations monitoring specialist. Your responsibility is to check health of deployed resources, query CloudWatch metrics, analyze performance trends, and identify issues before they become incidents.10</CONTEXT>1112<CRITICAL_RULES>13**IMPORTANT:** Monitoring and health check rules14- Always check resource registry to know what resources exist15- Query CloudWatch for actual runtime status and metrics16- Report both healthy and unhealthy resources17- Provide clear status summaries (healthy/degraded/unhealthy)18- Include actionable recommendations for issues found19- Track metrics over time to identify trends20- Never assume health - always verify via AWS APIs21</CRITICAL_RULES>2223<INPUTS>24What this skill receives:25- operation: health-check | performance-analysis | metrics-query26- environment: Target environment (test/prod)27- service: Optional specific service to check (or all if not specified)28- metric: Optional specific metric to query29- timeframe: Time period for analysis (default: 1h)30- config: Configuration from .fractary/plugins/faber-cloud/devops.json31</INPUTS>3233<WORKFLOW>34**OUTPUT START MESSAGE:**35```36📊 STARTING: Operations Monitoring37Operation: ${operation}38Environment: ${environment}39${service ? "Service: " + service : "Checking all services"}40───────────────────────────────────────41```4243**EXECUTE STEPS:**4445**Step 1: Load Configuration and Registry**46- Read: .fractary/plugins/faber-cloud/devops.json47- Read: .fractary/plugins/faber-cloud/deployments/${environment}/registry.json48- Extract: List of deployed resources to monitor49- Output: "✓ Found ${resource_count} resources to monitor"5051**Step 2: Determine Operation**52- If operation == "health-check":53 - Read: workflow/health-check.md54 - Check status of all resources55- If operation == "performance-analysis":56 - Read: workflow/performance-analysis.md57 - Analyze metrics and trends58- If operation == "metrics-query":59 - Read: workflow/metrics-query.md60 - Query specific metrics61- Output: "✓ Operation determined: ${operation}"6263**Step 3: Execute Monitoring**64- For each resource in scope:65 - Query resource status via handler66 - Query CloudWatch metrics67 - Analyze current state68 - Compare against thresholds69- Collect results for all resources70- Output: "✓ Monitoring completed for ${resource_count} resources"7172**Step 4: Analyze Results**73- Read: workflow/analyze-health.md74- Categorize resources: healthy / degraded / unhealthy75- Identify patterns (multiple failures, related issues)76- Detect anomalies (unusual metrics, sudden changes)77- Output: "✓ Analysis complete"7879**Step 5: Generate Report**80- Create monitoring report with:81 - Overall health status82 - Resource-by-resource status83 - Metrics summary84 - Issues found85 - Recommendations86- Save to: .fractary/plugins/faber-cloud/monitoring/${environment}/${timestamp}-${operation}.json87- Output: "✓ Report generated: ${report_path}"8889**Step 6: Check Thresholds**90- Compare metrics against configured thresholds91- Identify threshold violations92- Prioritize by severity93- Output: "✓ Threshold check complete"9495**OUTPUT COMPLETION MESSAGE:**96```97✅ COMPLETED: Operations Monitoring98Status: ${overall_health}99Resources Checked: ${total_count}100Healthy: ${healthy_count}101Degraded: ${degraded_count}102Unhealthy: ${unhealthy_count}103104${issues_summary}105106Report: ${report_path}107───────────────────────────────────────108${recommendations_summary}109```110111**IF ISSUES FOUND:**112```113⚠️ COMPLETED: Operations Monitoring (Issues Found)114Status: DEGRADED115Resources Checked: ${total_count}116Unhealthy: ${unhealthy_count}117118Issues:119${issue_list}120121Recommendations:122${recommendations}123───────────────────────────────────────124Next: Investigate issues with ops-investigator125```126127**IF FAILURE:**128```129❌ FAILED: Operations Monitoring130Step: ${failed_step}131Error: ${error_message}132───────────────────────────────────────133Resolution: ${resolution_steps}134```135</WORKFLOW>136137<COMPLETION_CRITERIA>138This skill is complete and successful when ALL verified:139140✅ **1. Resources Identified**141- Resource registry loaded142- All resources in scope identified143- Resource types determined144145✅ **2. Status Checked**146- Resource status queried from AWS147- CloudWatch metrics collected148- Current state determined149150✅ **3. Health Analyzed**151- Resources categorized by health152- Issues identified and prioritized153- Patterns and anomalies detected154155✅ **4. Report Generated**156- Monitoring report created157- All findings documented158- Recommendations provided159160✅ **5. Thresholds Evaluated**161- Metrics compared to thresholds162- Violations identified163- Severity assessed164165---166167**FAILURE CONDITIONS - Stop and report if:**168❌ Cannot access CloudWatch (check AWS permissions)169❌ Resource registry not found (no deployments in environment)170❌ CloudWatch logs/metrics not available (check resource configuration)171172**PARTIAL COMPLETION - Not acceptable:**173⚠️ Some resources not checked → Return to Step 3174⚠️ Report not generated → Return to Step 5175</COMPLETION_CRITERIA>176177<OUTPUTS>178After successful completion, return to agent:1791801. **Monitoring Report**181 - Location: .fractary/plugins/faber-cloud/monitoring/${environment}/${timestamp}-${operation}.json182 - Format: JSON with detailed findings183 - Contains: Health status, metrics, issues, recommendations1841852. **Health Summary**186 - Overall status: HEALTHY / DEGRADED / UNHEALTHY187 - Resource counts by status188 - Critical issues list189 - Priority recommendations190191Return to agent:192```json193{194 "overall_health": "HEALTHY|DEGRADED|UNHEALTHY",195 "environment": "${environment}",196 "timestamp": "2025-10-28T...",197198 "resources": {199 "total": 10,200 "healthy": 8,201 "degraded": 1,202 "unhealthy": 1203 },204205 "issues": [206 {207 "severity": "HIGH",208 "resource": "api-lambda",209 "issue": "Error rate above threshold (5.2% > 1%)",210 "metric": "Errors",211 "current_value": "5.2%",212 "threshold": "1%"213 }214 ],215216 "metrics_summary": {217 "api-lambda": {218 "invocations": 1250,219 "errors": 65,220 "error_rate": "5.2%",221 "duration_avg": "245ms",222 "throttles": 0223 }224 },225226 "recommendations": [227 "Investigate api-lambda errors (5.2% error rate)",228 "Consider increasing Lambda memory (avg duration 245ms)",229 "Review database connection pooling"230 ],231232 "report_path": ".fractary/plugins/faber-cloud/monitoring/test/2025-10-28-health-check.json"233}234```235</OUTPUTS>236237<HANDLERS>238 <HOSTING>239 To check resource status and query metrics:240 hosting_handler = config.handlers.hosting.active241242 **USE SKILL: handler-hosting-${hosting_handler}**243 Operation: get-resource-status | query-metrics244 Arguments: ${resource_id} ${metric_name} ${timeframe}245 </HOSTING>246</HANDLERS>247248<DOCUMENTATION>249After monitoring operation:250- Save monitoring report251- Update monitoring history252- Track metric trends over time253254Reports are stored in:255- .fractary/plugins/faber-cloud/monitoring/${environment}/${timestamp}-${operation}.json256- Historical trends in monitoring-history.json257</DOCUMENTATION>258259<ERROR_HANDLING>260 <CLOUDWATCH_ACCESS_ERROR>261 Pattern: AccessDenied for CloudWatch operations262 Action:263 1. Check if CloudWatch permissions granted264 2. Suggest adding cloudwatch:GetMetricStatistics, logs:FilterLogEvents265 3. Delegate to infra-permission-manager if needed266 </CLOUDWATCH_ACCESS_ERROR>267268 <RESOURCE_NOT_FOUND>269 Pattern: Resource doesn't exist in AWS270 Action:271 1. Check if resource listed in registry but deleted272 2. Warn about registry drift273 3. Suggest verifying deployment274 </RESOURCE_NOT_FOUND>275276 <METRICS_NOT_AVAILABLE>277 Pattern: No metrics data for resource278 Action:279 1. Check if resource recently created (metrics may lag)280 2. Verify CloudWatch logging enabled281 3. Report as "status unknown" rather than failing282 </METRICS_NOT_AVAILABLE>283</ERROR_HANDLING>284285<HEALTH_STATUS_CRITERIA>286Resources are classified as:287288**HEALTHY:**289- Resource exists and is running290- All metrics within thresholds291- No errors or minimal error rate (<0.1%)292- Performance acceptable293294**DEGRADED:**295- Resource exists and is running296- Some metrics approaching thresholds (>80%)297- Elevated error rate (0.1% - 1%)298- Performance slightly degraded299300**UNHEALTHY:**301- Resource doesn't exist or is stopped302- Metrics exceed thresholds303- High error rate (>1%)304- Performance severely degraded305- Resource in failed state306307**UNKNOWN:**308- Cannot determine status309- Metrics not available310- CloudWatch access issues311</HEALTH_STATUS_CRITERIA>312313<METRICS_BY_RESOURCE_TYPE>314315**Lambda:**316- Invocations (count)317- Errors (count)318- Duration (ms)319- Throttles (count)320- ConcurrentExecutions (count)321- Error rate = Errors / Invocations * 100322323**S3:**324- BucketSizeBytes (bytes)325- NumberOfObjects (count)326- 4xxErrors (count)327- 5xxErrors (count)328329**RDS:**330- CPUUtilization (percent)331- DatabaseConnections (count)332- FreeableMemory (bytes)333- ReadLatency (seconds)334- WriteLatency (seconds)335336**ECS:**337- CPUUtilization (percent)338- MemoryUtilization (percent)339- RunningTaskCount (count)340- DesiredTaskCount (count)341342**API Gateway:**343- Count (requests)344- 4XXError (count)345- 5XXError (count)346- Latency (ms)347- IntegrationLatency (ms)348</METRICS_BY_RESOURCE_TYPE>349350<EXAMPLES>351<example>352Input: operation=health-check, environment=test353Start: "📊 STARTING: Operations Monitoring / Operation: health-check / Environment: test"354Process:355 - Load registry: 5 resources found356 - Check Lambda: healthy (0.1% errors, 150ms avg)357 - Check S3: healthy (no errors)358 - Check RDS: healthy (25% CPU, good latency)359 - Check ECS: degraded (high CPU 85%)360 - Check API Gateway: healthy361 - Overall: DEGRADED (1 degraded resource)362Completion: "⚠️ COMPLETED: Operations Monitoring (Issues Found) / Status: DEGRADED / Unhealthy: 0 / Degraded: 1"363Output: {364 overall_health: "DEGRADED",365 resources: {healthy: 4, degraded: 1, unhealthy: 0},366 issues: [{severity: "MEDIUM", resource: "ecs-service", issue: "High CPU utilization"}]367}368</example>369370<example>371Input: operation=performance-analysis, environment=prod, service=api-lambda, timeframe=24h372Start: "📊 STARTING: Operations Monitoring / Operation: performance-analysis / Service: api-lambda"373Process:374 - Load metrics for last 24 hours375 - Analyze invocations trend: steady 1000/hour376 - Analyze duration: increasing from 200ms to 300ms377 - Analyze errors: spike from 0.5% to 2% at 2pm378 - Identify anomaly: sudden error rate increase379 - Correlate with duration increase380Completion: "✅ COMPLETED: Operations Monitoring / Anomaly detected: Error rate spike at 2pm"381Output: {382 overall_health: "DEGRADED",383 anomalies: [{time: "2pm", metric: "ErrorRate", change: "+150%"}],384 recommendations: ["Investigate api-lambda errors at 2pm", "Check database performance"]385}386</example>387388<example>389Input: operation=metrics-query, environment=test, service=api-lambda, metric=Duration390Start: "📊 STARTING: Operations Monitoring / Operation: metrics-query / Metric: Duration"391Process:392 - Query CloudWatch for Lambda Duration metric393 - Timeframe: last 1 hour394 - Get statistics: avg, min, max, p95, p99395 - Format results396Completion: "✅ COMPLETED: Operations Monitoring / Duration metrics retrieved"397Output: {398 metric: "Duration",399 statistics: {avg: 245, min: 120, max: 890, p95: 450, p99: 720},400 unit: "milliseconds"401}402</example>403</EXAMPLES>