Linux Troubleshooting Workflow
Shared Knowledge: This skill builds on the guidelines in
brain/knowledge/devops-operations.md. Always apply those principles alongside the specific guidance below.
Overview
Specialized workflow for diagnosing and resolving Linux system issues including performance problems, service failures, network issues, and resource constraints.
When to Use This Workflow
Use this workflow when:
- Diagnosing system performance issues
- Troubleshooting service failures
- Investigating network problems
- Resolving disk space issues
- Debugging application errors
Workflow Phases
Phase 1: Initial Assessment
Actions
- Check system uptime
- Review recent changes
- Identify symptoms
- Gather error messages
- Document findings
Commands
uptime
hostnamectl
cat /etc/os-release
dmesg | tail -50
Phase 2: Resource Analysis
Actions
- Check CPU usage
- Analyze memory
- Review disk space
- Monitor I/O
- Check network
Commands
top -bn1 | head -20
free -h
df -h
iostat -x 1 5
Phase 3: Process Investigation
Actions
- List running processes
- Identify resource hogs
- Check process status
- Review process trees
- Analyze strace output
Commands
ps aux --sort=-%cpu | head -10
pstree -p
lsof -p PID
strace -p PID
Phase 4: Log Analysis
Actions
- Check system logs
- Review application logs
- Search for errors
- Analyze log patterns
- Correlate events
Commands
journalctl -xe
tail -f /var/log/syslog
grep -i error /var/log/*
Phase 5: Network Diagnostics
Actions
- Check network interfaces
- Test connectivity
- Analyze connections
- Review firewall rules
- Check DNS resolution
Commands
ip addr show
ss -tulpn
curl -v http://target
dig domain
Phase 6: Service Troubleshooting
Actions
- Check service status
- Review service logs
- Test service restart
- Verify dependencies
- Check configuration
Commands
systemctl status service
journalctl -u service -f
systemctl restart service
Phase 7: Resolution
Follow the gather -> hypothesize -> test -> verify -> document methodology in brain/knowledge/devops-operations.md §3 to implement the fix, verify the resolution, monitor stability, and document root cause plus a prevention plan.
Troubleshooting Checklist
- System information gathered
- Resources analyzed
- Logs reviewed
- Network tested
- Services verified
- Root cause identified
- Issue resolved / fix verified
- Monitoring in place
- Documentation created