Why Your Kubernetes HPA Won't Scale Down (It's Probably Not Stuck)
You scaled up under load, traffic dropped ten minutes ago, and your HPA is still sitting at 8 replicas instead of 2. First instinct: something's broken. Second instinct, after kubectl describe hpa shows nothing wrong: confusion. It's probably not stuck. It's doing exactly what it's configured to do — you just never configured it. The default nobody sets on purpose If you don't define a…
Your Kubernetes Horizontal Pod Autoscaler (HPA) isn't scaling down because it's likely not configured to do so. The default behavior is a 300-second stabilization window, which prevents Kubernetes from removing replicas when metrics briefly dip below the target. This isn't a bug, but a deliberate measure to avoid rapid scaling changes, or "flapping." However, this window resets each time a metric sample exceeds the target, making it seem like the HPA is stuck in a scale-down state, especially with noisy CPU metrics.
The solution is to explicitly define the scale-down settings in your HPA configuration. By specifying a stabilization window of 120 seconds and a policy that caps the rate of replica removal, you create a scale-down window that aligns with your expectations. This way, you won't have to guess whether 300 seconds equals 300 seconds or 300 seconds from the last metric spike.
To illustrate this, I've created a minimal, runnable example with a demo Deployment, an HPA with a properly configured scale-down block, and a load generator script. You can view the full breakdown at github.com/polasamy-eng/devsaas-devops-examples — specifically, the kubernetes-hpa-scaling-demo folder. It demonstrates how the HPA behaves when the scale-down settings are explicitly defined, helping you distinguish between a genuinely stuck HPA and one that's just conservatively or accidentally configured.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.