Cluster Audit
Run a comprehensive audit of the Kaizen K8s cluster. Execute ALL of these in sequence:
- Nodes:
kubectl get nodes -o wide - All Pods:
kubectl get pods -A -o wide - Problem Pods:
kubectl get pods -A --field-selector=status.phase!=Running,status.phase!=Succeeded - Services:
kubectl get svc -A - PVCs:
kubectl get pvc -A - GPU Allocation:
kubectl describe nodes | grep -A5 "nvidia.com/gpu" - Recent Events:
kubectl get events -A --sort-by=.lastTimestamp --no-headers | tail -20 - Inference Health:
curl -s --max-time 3 http://10.10.10.10:30000/v1/modelsandcurl -s --max-time 3 http://10.10.10.10:30001/v1/models - Cognitive Health:
curl -s --max-time 3 http://10.10.10.10:30800/healthandcurl -s --max-time 3 http://10.10.10.10:30810/health
Produce a summary table:
| Component | Status | Details |
|---|
Flag any CrashLoopBackOff, OOMKilled, or Pending pods. Note GPU utilization vs allocation.