Autoscaling Kubernetes

Expert mastery of autoscaling on Kubernetes across all layers — pod-horizontal (HPA), pod-vertical (VPA, in-place resize), node (Cluster Autoscaler, Karpenter, GKE Node Auto-Provisioning), and event-driven (KEDA). Use when designing or debugging HPA control loops (desiredReplicas, behavior, stabilization windows), custom/external metrics (custom.metrics.k8s.io, external.metrics.k8s.io, Prometheus Adapter), VPA modes and HPA-vs-VPA conflicts, Cluster Autoscaler vs Karpenter (NodePools, consolidation, disruption), KEDA ScaledObjects/ScaledJobs and scale-to-zero, Kueue ProvisioningRequest, or ML/GPU/LLM-inference autoscaling (GPU utilization, queue depth, TTFT/concurrency, scale-to-zero for accelerators, multi-host LWS). Covers tuning, metric-pipeline reliability, and autoscaler fights.

sanjeevrg89 Updated

File contents

sanjeevrg89/arete/tree/main/skills/autoscaling-kubernetes commit e45c2113c2

Frequently asked questions

npx skillmds@latest add sanjeevrg89/autoscaling-kubernetes