Kubernetes Troubleshooting
Expert debugging and diagnostics for Kubernetes clusters using kubectl-mcp-server tools.
When to Apply
Use this skill when:
- User mentions: "debug", "troubleshoot", "diagnose", "failing", "crash", "not starting", "broken"
- Pod states: Pending, CrashLoopBackOff, ImagePullBackOff, OOMKilled, Error, Unknown
- Node issues: NotReady, MemoryPressure, DiskPressure, NetworkUnavailable, PIDPressure
- Keywords: "logs", "events", "describe", "why isn't working", "stuck", "not responding"
Priority Rules
| Priority |
Rule |
Impact |
Tools |
| 1 |
Check pod status first |
CRITICAL |
get_pods, describe_pod |
| 2 |
View recent events |
CRITICAL |
get_events |
| 3 |
Inspect logs (including previous) |
HIGH |
get_pod_logs |
| 4 |
Check resource metrics |
HIGH |
get_pod_metrics |
| 5 |
Verify endpoints |
MEDIUM |
get_endpoints |
| 6 |
Review network policies |
MEDIUM |
get_network_policies |
| 7 |
Examine node status |
LOW |
get_nodes, describe_node |
Quick Reference
| Symptom |
First Tool |
Next Steps |
| Pod Pending |
describe_pod |
Check events, node capacity, resource requests |
| CrashLoopBackOff |
get_pod_logs(previous=True) |
Check exit code, resources, liveness probes |
| ImagePullBackOff |
describe_pod |
Verify image name, registry auth, network |
| OOMKilled |
get_pod_metrics |
Increase memory limits, check for memory leaks |
| ContainerCreating |
describe_pod |
Check PVC binding, secrets, configmaps |
| Terminating (stuck) |
describe_pod |
Check finalizers, PDBs, preStop hooks |
Diagnostic Workflows
Pod Not Starting
1. get_pods(namespace, label_selector) - Get pod status
2. describe_pod(name, namespace) - See events and conditions
3. get_events(namespace, field_selector="involvedObject.name=<pod>") - Check events
4. get_pod_logs(name, namespace, previous=True) - For crash loops
Common Pod States
| State |
Likely Cause |
Tools to Use |
| Pending |
Scheduling issues |
describe_pod, get_nodes, get_events |
| ImagePullBackOff |
Registry/auth |
describe_pod, check image name |
| CrashLoopBackOff |
App crash |
get_pod_logs(previous=True) |
| OOMKilled |
Memory limit |
get_pod_metrics, adjust limits |
| ContainerCreating |
Volume/network |
describe_pod, get_pvc |
Node Issues
1. get_nodes() - List nodes and status
2. describe_node(name) - See conditions and capacity
3. Check: Ready, MemoryPressure, DiskPressure, PIDPressure
4. node_logs_tool(name, "kubelet") - Kubelet logs
Deep Debugging Workflows
CrashLoopBackOff Investigation
1. get_pod_logs(name, namespace, previous=True) - See why it crashed
2. describe_pod(name, namespace) - Check resource limits, probes
3. get_pod_metrics(name, namespace) - Memory/CPU at crash time
4. If OOM: compare requests/limits to actual usage
5. If app error: check logs for stack trace
Networking Issues
1. get_services(namespace) - Verify service exists
2. get_endpoints(namespace) - Check endpoint backends
3. If empty endpoints: pods don't match selector
4. get_network_policies(namespace) - Check traffic rules
5. For Cilium: cilium_endpoints_list_tool(), hubble_flows_query_tool()
Storage Problems
1. get_pvc(namespace) - Check PVC status
2. describe_pvc(name, namespace) - See binding issues
3. get_storage_classes() - Verify provisioner exists
4. If Pending: check storage class, access modes
DNS Resolution
1. kubectl_exec(pod, namespace, "nslookup kubernetes.default") - Test DNS
2. If fails: check coredns pods in kube-system
3. get_pods(namespace="kube-system", label_selector="k8s-app=kube-dns")
4. get_pod_logs(name="coredns-*", namespace="kube-system")
Multi-Cluster Debugging
All tools support context parameter for targeting different clusters:
get_pods(namespace="kube-system", context="production-cluster")
get_events(namespace="default", context="staging-cluster")
describe_pod(name="myapp-xyz", namespace="prod", context="prod-east")
Diagnostic Scripts
For comprehensive diagnostics, run the bundled scripts:
- See scripts/diagnose-pod.py for automated pod analysis
- See scripts/health-check.sh for cluster health checks
Decision Tree
See references/DECISION-TREE.md for visual troubleshooting flowcharts.
Common Errors Reference
See references/COMMON-ERRORS.md for error message explanations and fixes.
Related Tools
Core Diagnostics
get_pods, describe_pod, get_pod_logs, get_pod_metrics
get_events, get_nodes, describe_node
get_resource_usage, compare_namespaces
Advanced (Ecosystem)
- Cilium:
cilium_endpoints_list_tool, hubble_flows_query_tool
- Istio:
istio_proxy_status_tool, istio_analyze_tool
Related Skills
1---2name: k8s-troubleshoot3description: Kubernetes Troubleshooting4---5# Kubernetes Troubleshooting67Expert debugging and diagnostics for Kubernetes clusters using kubectl-mcp-server tools.89## When to Apply1011Use this skill when:12- User mentions: "debug", "troubleshoot", "diagnose", "failing", "crash", "not starting", "broken"13- Pod states: Pending, CrashLoopBackOff, ImagePullBackOff, OOMKilled, Error, Unknown14- Node issues: NotReady, MemoryPressure, DiskPressure, NetworkUnavailable, PIDPressure15- Keywords: "logs", "events", "describe", "why isn't working", "stuck", "not responding"1617## Priority Rules1819| Priority | Rule | Impact | Tools |20|----------|------|--------|-------|21| 1 | Check pod status first | CRITICAL | `get_pods`, `describe_pod` |22| 2 | View recent events | CRITICAL | `get_events` |23| 3 | Inspect logs (including previous) | HIGH | `get_pod_logs` |24| 4 | Check resource metrics | HIGH | `get_pod_metrics` |25| 5 | Verify endpoints | MEDIUM | `get_endpoints` |26| 6 | Review network policies | MEDIUM | `get_network_policies` |27| 7 | Examine node status | LOW | `get_nodes`, `describe_node` |2829## Quick Reference3031| Symptom | First Tool | Next Steps |32|---------|------------|------------|33| Pod Pending | `describe_pod` | Check events, node capacity, resource requests |34| CrashLoopBackOff | `get_pod_logs(previous=True)` | Check exit code, resources, liveness probes |35| ImagePullBackOff | `describe_pod` | Verify image name, registry auth, network |36| OOMKilled | `get_pod_metrics` | Increase memory limits, check for memory leaks |37| ContainerCreating | `describe_pod` | Check PVC binding, secrets, configmaps |38| Terminating (stuck) | `describe_pod` | Check finalizers, PDBs, preStop hooks |3940## Diagnostic Workflows4142### Pod Not Starting4344```451. get_pods(namespace, label_selector) - Get pod status462. describe_pod(name, namespace) - See events and conditions473. get_events(namespace, field_selector="involvedObject.name=<pod>") - Check events484. get_pod_logs(name, namespace, previous=True) - For crash loops49```5051### Common Pod States5253| State | Likely Cause | Tools to Use |54|-------|-------------|--------------|55| Pending | Scheduling issues | `describe_pod`, `get_nodes`, `get_events` |56| ImagePullBackOff | Registry/auth | `describe_pod`, check image name |57| CrashLoopBackOff | App crash | `get_pod_logs(previous=True)` |58| OOMKilled | Memory limit | `get_pod_metrics`, adjust limits |59| ContainerCreating | Volume/network | `describe_pod`, `get_pvc` |6061### Node Issues6263```641. get_nodes() - List nodes and status652. describe_node(name) - See conditions and capacity663. Check: Ready, MemoryPressure, DiskPressure, PIDPressure674. node_logs_tool(name, "kubelet") - Kubelet logs68```6970## Deep Debugging Workflows7172### CrashLoopBackOff Investigation7374```751. get_pod_logs(name, namespace, previous=True) - See why it crashed762. describe_pod(name, namespace) - Check resource limits, probes773. get_pod_metrics(name, namespace) - Memory/CPU at crash time784. If OOM: compare requests/limits to actual usage795. If app error: check logs for stack trace80```8182### Networking Issues8384```851. get_services(namespace) - Verify service exists862. get_endpoints(namespace) - Check endpoint backends873. If empty endpoints: pods don't match selector884. get_network_policies(namespace) - Check traffic rules895. For Cilium: cilium_endpoints_list_tool(), hubble_flows_query_tool()90```9192### Storage Problems9394```951. get_pvc(namespace) - Check PVC status962. describe_pvc(name, namespace) - See binding issues973. get_storage_classes() - Verify provisioner exists984. If Pending: check storage class, access modes99```100101### DNS Resolution102103```1041. kubectl_exec(pod, namespace, "nslookup kubernetes.default") - Test DNS1052. If fails: check coredns pods in kube-system1063. get_pods(namespace="kube-system", label_selector="k8s-app=kube-dns")1074. get_pod_logs(name="coredns-*", namespace="kube-system")108```109110## Multi-Cluster Debugging111112All tools support `context` parameter for targeting different clusters:113114```python115get_pods(namespace="kube-system", context="production-cluster")116get_events(namespace="default", context="staging-cluster")117describe_pod(name="myapp-xyz", namespace="prod", context="prod-east")118```119120## Diagnostic Scripts121122For comprehensive diagnostics, run the bundled scripts:123- See [scripts/diagnose-pod.py](scripts/diagnose-pod.py) for automated pod analysis124- See [scripts/health-check.sh](scripts/health-check.sh) for cluster health checks125126## Decision Tree127128See [references/DECISION-TREE.md](references/DECISION-TREE.md) for visual troubleshooting flowcharts.129130## Common Errors Reference131132See [references/COMMON-ERRORS.md](references/COMMON-ERRORS.md) for error message explanations and fixes.133134## Related Tools135136### Core Diagnostics137- `get_pods`, `describe_pod`, `get_pod_logs`, `get_pod_metrics`138- `get_events`, `get_nodes`, `describe_node`139- `get_resource_usage`, `compare_namespaces`140141### Advanced (Ecosystem)142- Cilium: `cilium_endpoints_list_tool`, `hubble_flows_query_tool`143- Istio: `istio_proxy_status_tool`, `istio_analyze_tool`144145## Related Skills146147- [k8s-diagnostics](../k8s-diagnostics/SKILL.md) - Metrics and health checks148- [k8s-incident](../k8s-incident/SKILL.md) - Emergency runbooks149- [k8s-networking](../k8s-networking/SKILL.md) - Network troubleshooting