Urgent.News

What's breaking now, across thousands of outlets.

Tech

Article: High Availability Is Not Resilience: Why Cloud Systems Fail When It Matters Most

A routine TLS 1.3 upgrade silently broke Route 53 health checks, causing a CDN to stop routing traffic to a healthy region while internal dashboards showed nothing wrong. This article examines why HA and resilience are different problems, how control-plane dependencies create invisible failure modes, and why recovery capability erodes without explicit ownership. By Alexey Golev

High availability is not resilience, and understanding the distinction is crucial when cloud systems fail during critical moments. A team upgraded their public ingress load balancers to TLS 1.3 for compliance, but this change impacted their cloud architecture. The Route 53 HTTPS health checks required TLS 1.2 support, and with only TLS 1.3 available, the health checker marked the endpoint unhealthy.

The CDN deemed the region unhealthy, leading to traffic rerouting to another region, causing latency spikes. The failure was in the control plane, not the application data plane, and took about forty minutes to isolate. While the system was highly available, it was not resilient. High availability focuses on surviving expected failures, while resilience deals with recovering from unexpected conditions.

Measuring resilience is challenging, as it resists uptime percentages and replication lag. Modern cloud platforms make high availability relatively accessible, but real incidents often don't fit the assumptions of high-availability engineering. Resilience engineering starts with the belief that critical failures will occur in unmodelled ways.

Three patterns highlight the gap: software-layer correlation, such as simultaneous config pushes or dependency upgrades; correlated failures, like an entire availability zone going down; and the assumption that systems will degrade gracefully. Most systems are tested under graceful degradation, but few have been tested under realistic loads.

Recovery ownership is often ambiguous, and few organizations have dedicated structures for resilience testing. Recovery ownership should be separate from on-call, with dedicated engineering time and risk to demonstrate meaningful recovery capabilities. The first thirty minutes of a major incident are often lost due to dependency archaeology, as runbooks and runbooks reference outdated tools.

The fix requires explicit recovery ownership, which is distinct from on-call, and unglamorous, requiring preparation and proactive effort.

Written by urgent.news from InfoQ's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at infoq.com →

More in Tech

More from Thursday 1 October →