Gke Tpu Metrics Monitoring

Monitors and troubleshoots GKE TPU workloads, nodes, and node pools using GKE system metrics and PromQL. Use when monitoring TensorCore duty cycle, TPU memory, node readiness, multi-host TPU node pool availability, host maintenance or preemption interruptions, and calculating MTTR or MTBI metrics for GKE TPUs. Don't use for general non-TPU GKE workload monitoring or non-metric TPU debugging.

Harlieunadjusted52 Updated

File contents

Harlieunadjusted52/skills/tree/main/skills/cloud/gke-tpu-metrics-monitoring commit 6ba22b8b2a

Frequently asked questions

npx skillmds@latest add harlieunadjusted52/gke-tpu-metrics-monitoring