Urgent.News

What's breaking now, across thousands of outlets.

Tech

A zero-disruption PodDisruptionBudget can block an AKS upgrade

Your AKS cluster upgrade is stuck. The control plane is waiting for a node to drain. The node is waiting for a pod to terminate. The pod is refusing to move. Look at the logs. The workload has a PodDisruptionBudget that allows exactly zero disruptions. It is doing exactly what it was told to do. Protect the application at all costs. Even if it means halting the cluster upgrade. To upgrade a node,…

Your AKS cluster upgrade has come to a grinding halt. The control plane is waiting for a node to be drained. Meanwhile, the node is in limbo, awaiting a pod to terminate. Unfortunately, the pod is being stubborn, refusing to move. Upon closer inspection of the logs, it becomes apparent that the workload has a PodDisruptionBudget that allows precisely zero disruptions.

This policy is diligently being adhered to, even at the expense of the cluster upgrade. To successfully upgrade a node, AKS must first drain it, a process that involves cajoling the node and evicting the pods within. However, the eviction process takes into account the disruption budgets set in place. If your deployment maintains three replicas, and your budget stipulates that three must remain operational, the eviction process will inevitably fail.

Consequently, the node remains tainted and the upgrade grinds to a standstill. It seems you did not merely configure high availability; you inadvertently created a deadlock scenario. The concept of replica headroom isn't merely a best practice for scaling; it is the indispensable space necessary for the control plane to carry out maintenance procedures.

Prior to your upcoming maintenance window, take a moment to review your calculations. Examine your minAvailable or maxUnavailable settings in relation to the actual quantity of running replicas. Keep in mind that Kubernetes rounds up when dealing with minAvailable values specified as a percentage. A setting of 100% translates to zero disruptions allowed, while a setting of 90% with three replicas effectively rounds up to three.

It is crucial to allocate sufficient space for the scheduler to operate effectively. Configure these values with deliberate intent, not merely to satisfy a compliance check. High availability signifies the ability to withstand a failure, yet it should not translate to obstructing your own operations team. Your availability policy ensures the application's survival in the event of a zone failure; however, have you tested its resilience against routine maintenance procedures?

If your zero-disruption configuration obstructs your upgrade pipeline, you have not constructed a resilient system. Instead, you have unwittingly created a hostage situation.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in Tech

More from Thursday 8 October →