Urgent.News

What's breaking now, across thousands of outlets.

Tech

Availability Should Follow the Workload, Not the Infrastructure

Platform engineering has automated deployment but not recovery. The next maturity curve is intelligent operations that automate routine production decisions and resilience.

Availability Should Follow the Workload, Not the Infrastructure

I can no longer ignore the question that haunts me… How did we automate almost every aspect of software delivery, yet still rely on humans when production fails? Consider the advancements modern platform engineering has made. We create infrastructure through code, deploy pipelines capable of hundreds of deployments daily, and utilize GitOps to maintain consistency across environments.

Our systems can automatically scale based on demand. Kubernetes enables us to spin up an entire cluster in mere minutes. It's impressive progress. But when a production database crashes at 2:13 a.m., the Slack channel erupts. Individuals initiate bridge calls, while others sift through outdated runbooks. And everyone waits for the engineer who "understands this system."

Is this still our operating model? Over the past decade, we've taught our platforms how to deploy software. Yet, we've devoted far less time to teaching them how to react when production breaks. That's the next step in platform engineering maturity – not more deployment automation, but intelligent operational automation. One prevalent misconception is equating modern infrastructure with resilient infrastructure.

They are not synonymous. Kubernetes excels at managing workloads, restarting pods, replacing failed nodes, and maintaining desired state. However, it is unaware of whether customers can still place orders, process payments, or complete transactions. This is not a criticism of Kubernetes; its role isn't reliability. Organizations often conflate orchestration with a comprehensive reliability strategy.

It is not. Orchestration addresses the question, "Where should this workload run?" Reliability tackles a different query: "How can I ensure this workload remains available when something inevitably fails?" These are distinct problems. In contemporary settings, availability shouldn't be linked to a server, cluster, or even a cloud.

It should follow the workload. The question every platform team must ask is: When production incidents occur, what should happen next – and how much still depends on humans? If the answer is simply, "Someone gets paged, we jump on a call, and then we figure it out," your platform may not be as automated as it seems. It's not a criticism; it's an opportunity.

The objective isn't to eliminate engineers from the process. Instead, the aim is to eliminate the routine operational decisions they've already performed hundreds of times. Engineers should concentrate on novel issues rather than repeating the same recovery steps every time a server, node, or workload malfunctions. So, what should you do differently on Monday?

Here's where I would begin. Identify your 2 a.m. processes. Compile a list of all manual actions your team takes during a production incident. Not deployments or upgrades, but failures. If only one person comprehends how to recover a critical workload, that is not expertise; it's technical debt. Recovery should not reside within someone's mind.

It should be embedded in the platform. Measure application recovery, not infrastructure recovery. Most teams are familiar with how long it takes to deploy a release. Fewer know how long it takes for a critical application to recover from a real failure. Users do not care about how swiftly a pod restarts; they care about how quickly they can resume work.

Measure what they experience. More importantly, begin measuring how many recovery decisions still require human intervention. That may be the best gauge of your platform's maturity. Test recovery as frequently as you test deployments. Most organizations exercise their CI/CD pipelines consistently. How often do you deliberately disrupt production to observe its behavior?

If the answer is "almost never," then you may have identified your next engineering project. Stop presuming that Kubernetes resolved everything. Kubernetes is an exceptional orchestration platform, one of the best ever built. However, it is merely one layer of your platform. Stateful applications, databases, storage, networking, and application dependencies each have unique failure modes.

Ensure your operational strategy accommodates them. Modern platforms are not resilient merely because they run on Kubernetes. They are resilient because every layer – from infrastructure to the application itself – is engineered to recover intelligently when something fails. Eliminate decisions, not just steps. The goal isn't to make your outage runbook shorter.

The goal is to make it unnecessary. The most advanced platforms don't merely automate tasks; they automate decisions that have already been defined through policy and testing. That's how you minimize downtime and alleviate stress on your engineers concurrently. A straightforward test: if your most experienced engineer took two weeks off tomorrow, would your recovery process operate identically?

If the answer is "probably," your platform still relies on tribal knowledge. Mature platforms do not merely automate tasks; they automate operational decisions. The Next Frontier Extends Beyond Deployment. It's Intelligent Operations For years, we've evaluated platform engineering by one question: "How fast can we deploy?" I believe there's a superior question for the upcoming decade: "How intelligently can we recover?"

The answer relies less on where an application resides and more on how effortlessly it can continue running when conditions change. Contemporary enterprises do not choose between one environment over another; they operate across on-premises infrastructure, virtual machines, Kubernetes clusters, public clouds, and edge environments simultaneously.

Reliability can no longer be confined to a specific server, cluster, or cloud. It must follow the workload.

Written by urgent.news from DevOps.com's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at devops.com →

More in Tech

Spectranet Marks Customer Service Week, Expands Home Fiber and Launches New Website

As the global community celebrates Customer Service Week 2026, Nigeria’s leading Internet Service Provider Spectranet has reaffirmed its commitment to placing customers at the centre of its business…

  • Spectranet celebrates Customer Service Week 2026, focusing on 'The Extra Mile'
  • Launches new website spectranet.com.ng for smoother digital experience
  • Expands Home Fiber network for faster speeds and seamless streaming

More from Tuesday 6 October →