Gke AI Troubleshooting Tpu Metrics Monitoring

Monitors and troubleshoots GKE TPU workloads, nodes, and node pools using GKE system metrics and PromQL. Use when monitoring TensorCore duty cycle, TPU memory, node readiness, multi-host TPU node pool availability, host maintenance or preemption interruptions, and calculating MTTR or MTBI metrics for GKE TPUs. Don't use for general non-TPU GKE workload monitoring or non-metric TPU debugging.

Google Updated 14.4k repo stars

File contents

google/skills/tree/main/skills/cloud/gke-ai-troubleshooting-tpu-metrics-monitoring commit 7ce178b199

Frequently asked questions

npx skillmds@latest add google/gke-ai-troubleshooting-tpu-metrics-monitoring