K8S Troubleshooting

Expert Kubernetes troubleshooting assistant for diagnosing and resolving issues across the full stack - pods, control plane, nodes, networking, storage, and underlay infrastructure - in GPU cloud environments. Triggers on any report of a broken, degraded, or mysterious Kubernetes issue: pod crashes, OOMKills, scheduling failures, network problems, CRD errors, node NotReady, high latency, PVC issues, GPU/InfiniBand problems, workload hangs, or any cluster incident. Also triggers when the user pastes error messages, kubectl output, alert names, or incident-channel links and wants help understanding what's wrong. This skill works iteratively - it does NOT dump a wall of diagnostics all at once. It pauses after each step and asks the user how to proceed.

cfregly 4a0ac1b 11.3 KB Updated

File contents

cfregly/claude-gpu-perf-tune/tree/main/plugins/profile-and-optimize/skills/k8s-troubleshooting commit 4a0ac1b911

Frequently asked questions

npx skillmds@latest add cfregly/k8s-troubleshooting