IT Error Discovery & Classification
Overview
This skill covers the systematic discovery, classification, and prioritization of errors in the ConstructAI platform. It encompasses error detection from multiple sources (logs, user reports, monitoring alerts), classification by type/severity/component, and prioritization based on impact and frequency. Primary agent: 02050-002 Error Detective. Supporting skills: shared/build-error-fix, shared/systematic-debugging.
Triggers
- New error detected in application logs
- User reports an issue via chatbot or error modal
- Monitoring system alerts on anomalous behavior
- Automated error discovery scan identifies new patterns
- 500-level server error occurs
- Client-side JavaScript exception is captured
Prerequisites
- Access to error tracking system and logs
- Knowledge of system components and their criticality
- Understanding of severity classification matrix
- Access to 02050 domain knowledge for context
Steps
Step 1: Error Detection & Capture
- Ingest error from source (log, user report, monitoring alert)
- Capture full context: stack trace, component, user action, timestamp
- Extract structured fields: error type, message, severity indicators
- Tag with source system and environment (dev, staging, prod)
Step 2: Component Identification
- Parse stack trace to identify originating component
- Map component to system module using component registry
- Identify affected subsystems and potential cascade effects
- Check if component is customer-facing or internal
Step 3: Severity Classification
- Classify severity using matrix:
- P0 (Critical): System down, data loss, security breach
- P1 (High): Feature broken, significant degradation
- P2 (Medium): Minor feature impact, workaround available
- P3 (Low): Cosmetic, non-functional, edge case
- Use error type, user impact, and component criticality
- Assign confidence score to classification
Step 4: Pattern Matching
- Search error history for similar patterns (last 30 days)
- Check if error is recurring or new
- Identify if same root cause has been resolved before
- Link related errors into pattern groups
Step 5: Prioritization
- Apply prioritization rules:
- P0 errors: immediate page/auto-escalate
- P1 errors: same-day resolution target
- P2 errors: next sprint resolution
- P3 errors: backlog with trend monitoring
- Consider error frequency (high-frequency P2 → escalate to P1)
- Consider user impact scope (all users → escalate one level)
Step 6: Assignment & Routing
- Route to appropriate agent based on component:
- Frontend errors → UI/UX agent
- API errors → Backend agent
- Database errors → Data agent
- AI/LLM errors → AI agent
- Include full context and classification in assignment
- Set SLA based on priority level
Step 7: Documentation & Tracking
- Create error record in tracking system
- Record all classification decisions with reasoning
- Tag with component, severity, pattern group
- Log audit trail for all actions taken
Success Criteria
- Error correctly identified and captured with full context
- Component accurately identified (95%+ accuracy)
- Severity classification matches actual impact
- Pattern matching identifies recurring issues
- Priority level matches severity and frequency
- Error routed to correct agent/team
- Complete audit trail maintained
Common Pitfalls
- Incomplete stack trace: Ensure full context is captured, not just error message
- Over-classification as Critical: Not every 500 error is P0 — consider user impact
- Missing cascade effects: Always check if error affects downstream systems
- Incorrect component mapping: Verify component registry is up to date
- Ignoring error frequency: High-frequency P2 errors should escalate
- Duplicate error creation: Check pattern matching before creating new records
Cross-References
shared/systematic-debugging/SKILL.md — Root cause investigation for classified errors
shared/build-error-fix/SKILL.md — Resolution workflow for build-related errors
it-log-analysis-monitoring/SKILL.md — Ongoing error pattern monitoring
it-performance-monitoring-analytics/SKILL.md — Error rate trend analysis
Usage
Use this skill when any error is detected in the system that needs classification, prioritization, and routing. This is the first skill applied in the error management pipeline before resolution workflows.
Metrics
- Classification Accuracy: 95%+ of errors correctly classified
- Component Identification: 95%+ accuracy in identifying error source
- Routing Correctness: 90%+ of errors routed to correct team
- Pattern Detection: 80%+ of recurring errors identified as patterns
- Time to Classification: <30 seconds for automated classification
1---2name: it-error-discovery-classification3description: Skill for automated error detection, classification by severity and component, and prioritization of error resolution workflows in the ConstructAI platform4---56# IT Error Discovery & Classification78## Overview910This skill covers the systematic discovery, classification, and prioritization of errors in the ConstructAI platform. It encompasses error detection from multiple sources (logs, user reports, monitoring alerts), classification by type/severity/component, and prioritization based on impact and frequency. Primary agent: 02050-002 Error Detective. Supporting skills: `shared/build-error-fix`, `shared/systematic-debugging`.1112## Triggers1314- New error detected in application logs15- User reports an issue via chatbot or error modal16- Monitoring system alerts on anomalous behavior17- Automated error discovery scan identifies new patterns18- 500-level server error occurs19- Client-side JavaScript exception is captured2021## Prerequisites2223- Access to error tracking system and logs24- Knowledge of system components and their criticality25- Understanding of severity classification matrix26- Access to 02050 domain knowledge for context2728## Steps2930### Step 1: Error Detection & Capture31- Ingest error from source (log, user report, monitoring alert)32- Capture full context: stack trace, component, user action, timestamp33- Extract structured fields: error type, message, severity indicators34- Tag with source system and environment (dev, staging, prod)3536### Step 2: Component Identification37- Parse stack trace to identify originating component38- Map component to system module using component registry39- Identify affected subsystems and potential cascade effects40- Check if component is customer-facing or internal4142### Step 3: Severity Classification43- Classify severity using matrix:44 - **P0 (Critical)**: System down, data loss, security breach45 - **P1 (High)**: Feature broken, significant degradation46 - **P2 (Medium)**: Minor feature impact, workaround available47 - **P3 (Low)**: Cosmetic, non-functional, edge case48- Use error type, user impact, and component criticality49- Assign confidence score to classification5051### Step 4: Pattern Matching52- Search error history for similar patterns (last 30 days)53- Check if error is recurring or new54- Identify if same root cause has been resolved before55- Link related errors into pattern groups5657### Step 5: Prioritization58- Apply prioritization rules:59 - P0 errors: immediate page/auto-escalate60 - P1 errors: same-day resolution target61 - P2 errors: next sprint resolution62 - P3 errors: backlog with trend monitoring63- Consider error frequency (high-frequency P2 → escalate to P1)64- Consider user impact scope (all users → escalate one level)6566### Step 6: Assignment & Routing67- Route to appropriate agent based on component:68 - Frontend errors → UI/UX agent69 - API errors → Backend agent70 - Database errors → Data agent71 - AI/LLM errors → AI agent72- Include full context and classification in assignment73- Set SLA based on priority level7475### Step 7: Documentation & Tracking76- Create error record in tracking system77- Record all classification decisions with reasoning78- Tag with component, severity, pattern group79- Log audit trail for all actions taken8081## Success Criteria8283- Error correctly identified and captured with full context84- Component accurately identified (95%+ accuracy)85- Severity classification matches actual impact86- Pattern matching identifies recurring issues87- Priority level matches severity and frequency88- Error routed to correct agent/team89- Complete audit trail maintained9091## Common Pitfalls92931. **Incomplete stack trace**: Ensure full context is captured, not just error message942. **Over-classification as Critical**: Not every 500 error is P0 — consider user impact953. **Missing cascade effects**: Always check if error affects downstream systems964. **Incorrect component mapping**: Verify component registry is up to date975. **Ignoring error frequency**: High-frequency P2 errors should escalate986. **Duplicate error creation**: Check pattern matching before creating new records99100## Cross-References101102- `shared/systematic-debugging/SKILL.md` — Root cause investigation for classified errors103- `shared/build-error-fix/SKILL.md` — Resolution workflow for build-related errors104- `it-log-analysis-monitoring/SKILL.md` — Ongoing error pattern monitoring105- `it-performance-monitoring-analytics/SKILL.md` — Error rate trend analysis106107## Usage108109Use this skill when any error is detected in the system that needs classification, prioritization, and routing. This is the first skill applied in the error management pipeline before resolution workflows.110111## Metrics112113- **Classification Accuracy**: 95%+ of errors correctly classified114- **Component Identification**: 95%+ accuracy in identifying error source115- **Routing Correctness**: 90%+ of errors routed to correct team116- **Pattern Detection**: 80%+ of recurring errors identified as patterns117- **Time to Classification**: <30 seconds for automated classification