Your cluster says 3/3. It will still go dark with one zone.
Every deployment in this three-zone cluster reports ready: $ kubectl get deployments NAME READY UP-TO-DATE AVAILABLE checkout-api 3/3 3 3 session-store 1/1 1 1 web 3/3 3 3 Now take us-east-1a away: $ kubectl survive-zone Losing us-east-1a -> 2 lost, 0 degraded, 1 impaired by a dependency WORKLOAD PLACEMENT VERDICT checkout-api us-east-1a:3 LOST session-store us-east-1a:1 LOST web us-east-1a:1…
In a three-zone Kubernetes cluster, the 'checkout-api' and 'web' deployments experienced a failure when zone 'us-east-1a' was lost. The 'session-store' deployment was also affected as it had only one replica in the affected zone. Despite the deployments being scheduled correctly, two workloads were lost, and the remaining workload stopped serving.
The issue is not due to a misconfiguration but rather the nature of topology spread constraints in Kubernetes. These constraints are applied during pod scheduling and are not re-checked. The manifests specify multi-AZ deployment, but the cluster behaves differently once pods move. The kubectl survive-zone command checks the dependency graph when a zone is lost, and a workload can be marked as impaired even if its pods appear healthy.
The command provides a patch that ensures the pods are schedulable and survive the zone loss, but it does not guarantee that the issue will not reoccur.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.