IT Log Analysis & Monitoring
Overview
This skill covers log aggregation, pattern detection, anomaly identification, and monitoring alert configuration for the ConstructAI platform. It encompasses log ingestion pipelines, structured query generation, pattern matching algorithms, anomaly detection thresholds, and alert routing. Primary agent: 02050-004 Log Analyst. Supporting skills: it-error-discovery-classification.
Triggers
- New log source needs monitoring
- Alert thresholds require configuration or tuning
- Log patterns indicate potential systemic issues
- New component deployment requires monitoring setup
- Error rate exceeds configured thresholds
- Periodic log analysis audit is due
Prerequisites
- Access to log aggregation system
- Known baseline metrics for normal operation
- Component topology documentation
- Understanding of error classification from
it-error-discovery-classification
- Alert routing configuration access
Steps
Step 1: Log Source Identification
- Identify all log sources (application logs, server logs, database logs, API logs, infrastructure logs)
- Verify log format and parsing capabilities
- Confirm log ingestion pipeline connectivity
- Map log sources to system components
Step 2: Pattern Definition
- Define known normal patterns for each log source
- Define error patterns requiring alerting
- Define anomaly patterns indicating potential issues
- Create pattern match rules (regex, structured queries, ML-based detection)
Step 3: Baseline Establishment
- Collect baseline metrics for normal operation periods
- Establish normal error rate ranges per component
- Establish normal response time distributions
- Establish normal request volume patterns by time
Step 4: Threshold Configuration
- Configure error rate thresholds (warning at 2x baseline, critical at 5x baseline)
- Configure response time thresholds (warning at 150% baseline, critical at 300% baseline)
- Configure volume thresholds (warning at 150% expected, critical at 300% expected)
- Configure cascading failure detection rules
Step 5: Anomaly Detection
- Implement statistical anomaly detection (Z-score, IQR methods)
- Implement time-series anomaly detection (sudden spikes, gradual drifts)
- Implement correlation-based anomaly detection (multiple anomalies coinciding)
- Implement seasonal adjustment for periodic patterns
Step 6: Alert Configuration
- Configure alert levels (info, warning, critical, emergency)
- Configure alert routing (on-call, team leads, engineering director by level)
- Configure alert deduplication rules (same pattern, same component, same time window)
- Configure alert escalation rules (no acknowledgment within SLA → escalate)
Step 7: Dashboard Configuration
- Configure real-time dashboards with key metrics
- Configure trend charts for historical analysis
- Configure component health dashboards
- Configure business-relevant dashboards (user impact, revenue impact)
Step 8: Analysis & Action
- When anomaly detected, run root cause correlation
- When threshold breached, create error classification task
- When pattern indicates systemic issue, generate incident report
- When alert fires, provide context to responding team
Success Criteria
- All log sources identified and being ingested
- Pattern definitions cover 90%+ of known error scenarios
- Baseline metrics established for all components
- Thresholds configured and tuned (few false positives)
- Anomaly detection identifies 85%+ of real anomalies
- Alerts route correctly with appropriate deduplication
- Dashboards provide real-time visibility
Common Pitfalls
- Alert fatigue: Too many non-actionable alerts — tune thresholds regularly
- Missing context: Alerts without component or user impact context — enrich alert data
- Baseline staleness: Baselines drift over time — update baselines quarterly
- Log gaps: Not all components logging at appropriate levels — verify coverage
- Deduplication failures: Same alert firing repeatedly — improve deduplication windows
- No seasonal adjustment: Normal traffic patterns flagged as anomalies — implement seasonality
Cross-References
it-error-discovery-classification/SKILL.md — Error classification for detected anomalies
it-performance-monitoring-analytics/SKILL.md — Performance metrics integration
shared/systematic-debugging/SKILL.md — Root cause investigation from detected anomalies
Usage
Apply this skill when setting up monitoring for new components, configuring alert thresholds, investigating anomalous patterns, or performing periodic log analysis. This runs continuously for production monitoring as well as on-demand for investigations.
Metrics
- Pattern Coverage: 90%+ of known error patterns detected
- Anomaly Detection Rate: 85%+ of real anomalies detected
- False Positive Rate: <10% of alerts are false positives
- Alert Response Time: 95%+ of critical alerts acknowledged within SLA
- Log Coverage: 100% of critical components producing structured logs
1---2name: it-log-analysis-monitoring3description: Skill for log aggregation pipelines, pattern detection, anomaly identification, and alert configuration in the ConstructAI platform4---56# IT Log Analysis & Monitoring78## Overview910This skill covers log aggregation, pattern detection, anomaly identification, and monitoring alert configuration for the ConstructAI platform. It encompasses log ingestion pipelines, structured query generation, pattern matching algorithms, anomaly detection thresholds, and alert routing. Primary agent: 02050-004 Log Analyst. Supporting skills: `it-error-discovery-classification`.1112## Triggers1314- New log source needs monitoring15- Alert thresholds require configuration or tuning16- Log patterns indicate potential systemic issues17- New component deployment requires monitoring setup18- Error rate exceeds configured thresholds19- Periodic log analysis audit is due2021## Prerequisites2223- Access to log aggregation system24- Known baseline metrics for normal operation25- Component topology documentation26- Understanding of error classification from `it-error-discovery-classification`27- Alert routing configuration access2829## Steps3031### Step 1: Log Source Identification32- Identify all log sources (application logs, server logs, database logs, API logs, infrastructure logs)33- Verify log format and parsing capabilities34- Confirm log ingestion pipeline connectivity35- Map log sources to system components3637### Step 2: Pattern Definition38- Define known normal patterns for each log source39- Define error patterns requiring alerting40- Define anomaly patterns indicating potential issues41- Create pattern match rules (regex, structured queries, ML-based detection)4243### Step 3: Baseline Establishment44- Collect baseline metrics for normal operation periods45- Establish normal error rate ranges per component46- Establish normal response time distributions47- Establish normal request volume patterns by time4849### Step 4: Threshold Configuration50- Configure error rate thresholds (warning at 2x baseline, critical at 5x baseline)51- Configure response time thresholds (warning at 150% baseline, critical at 300% baseline)52- Configure volume thresholds (warning at 150% expected, critical at 300% expected)53- Configure cascading failure detection rules5455### Step 5: Anomaly Detection56- Implement statistical anomaly detection (Z-score, IQR methods)57- Implement time-series anomaly detection (sudden spikes, gradual drifts)58- Implement correlation-based anomaly detection (multiple anomalies coinciding)59- Implement seasonal adjustment for periodic patterns6061### Step 6: Alert Configuration62- Configure alert levels (info, warning, critical, emergency)63- Configure alert routing (on-call, team leads, engineering director by level)64- Configure alert deduplication rules (same pattern, same component, same time window)65- Configure alert escalation rules (no acknowledgment within SLA → escalate)6667### Step 7: Dashboard Configuration68- Configure real-time dashboards with key metrics69- Configure trend charts for historical analysis70- Configure component health dashboards71- Configure business-relevant dashboards (user impact, revenue impact)7273### Step 8: Analysis & Action74- When anomaly detected, run root cause correlation75- When threshold breached, create error classification task76- When pattern indicates systemic issue, generate incident report77- When alert fires, provide context to responding team7879## Success Criteria8081- All log sources identified and being ingested82- Pattern definitions cover 90%+ of known error scenarios83- Baseline metrics established for all components84- Thresholds configured and tuned (few false positives)85- Anomaly detection identifies 85%+ of real anomalies86- Alerts route correctly with appropriate deduplication87- Dashboards provide real-time visibility8889## Common Pitfalls90911. **Alert fatigue**: Too many non-actionable alerts — tune thresholds regularly922. **Missing context**: Alerts without component or user impact context — enrich alert data933. **Baseline staleness**: Baselines drift over time — update baselines quarterly944. **Log gaps**: Not all components logging at appropriate levels — verify coverage955. **Deduplication failures**: Same alert firing repeatedly — improve deduplication windows966. **No seasonal adjustment**: Normal traffic patterns flagged as anomalies — implement seasonality9798## Cross-References99100- `it-error-discovery-classification/SKILL.md` — Error classification for detected anomalies101- `it-performance-monitoring-analytics/SKILL.md` — Performance metrics integration102- `shared/systematic-debugging/SKILL.md` — Root cause investigation from detected anomalies103104## Usage105106Apply this skill when setting up monitoring for new components, configuring alert thresholds, investigating anomalous patterns, or performing periodic log analysis. This runs continuously for production monitoring as well as on-demand for investigations.107108## Metrics109110- **Pattern Coverage**: 90%+ of known error patterns detected111- **Anomaly Detection Rate**: 85%+ of real anomalies detected112- **False Positive Rate**: <10% of alerts are false positives113- **Alert Response Time**: 95%+ of critical alerts acknowledged within SLA114- **Log Coverage**: 100% of critical components producing structured logs