Gke AI Troubleshooting Tpu Metrics Monitoring

Monitors and troubleshoots GKE TPU workloads, nodes, and node pools using GKE system metrics and PromQL. Use when monitoring TensorCore duty cycle, TPU memory, node readiness, multi-host TPU node pool availability, host maintenance or preemption interruptions, and calculating MTTR or MTBI metrics for GKE TPUs. Don't use for general non-TPU GKE workload monitoring or non-metric TPU debugging.

tuyv Updated

File contents

tuyv/ccpm/tree/main/preset-registry/skills/google-skills-gke-ai-troubleshooting-tpu-metrics-monitoring commit 1c18c2e441

Frequently asked questions

npx skillmds@latest add tuyv/gke-ai-troubleshooting-tpu-metrics-monitoring