Urgent.News

What's breaking now, across thousands of outlets.

Tech

Amazon ECS now auto-repairs failing GPUs and instances. Here’s why it matters for SREs.

Running applications in production means maintaining an “always-on” posture through disruptions. Infrastructure fails; dependencies slow down, and networks partition, not The post Amazon ECS now auto-repairs failing GPUs and instances. Here’s why it matters for SREs. appeared first on The New Stack .

Amazon ECS now auto-repairs failing GPUs and instances. Here’s why it matters for SREs.

Amazon ECS now automatically repairs failing GPUs and instances, enhancing reliability for Serverless engineers. In production, infrastructure failures are inevitable – instances degrade, networks partition, and dependencies slow down. Amazon ECS attempts to handle these issues for you, eliminating the need to build detection loops and remediation runbooks for ECS-managed failures.

The resilience principles of ECS include uniform stability across Availability Zones, pre-scaling capacity, and workload isolation, with recovery mechanisms built atop these principles.

Two key features demonstrate this approach: Managed Instances and Fargate. Managed Instances take responsibility for resiliency of the underlying instance infrastructure, applying OS, kernel, and GPU driver patches without disrupting applications. AWS Fargate, meanwhile, relies on ECS to detect and recover from issues like bad instances, OS updates, or driver regressions.

This article focuses on how ECS manages specific failures: unhealthy instances, Availability Zone events, container or task failures, and degraded dependencies. When an instance becomes impaired, ECS detects the issue by integrating with NVIDIA’s Data Center GPU Manager (DCGM) to monitor GPU health. This allows ECS to identify genuine hardware faults that may not be visible to Amazon EC2 instance status checks, which often only surface as task failures or latency issues.

ECS also monitors the broader health of the data plane instance, including EC2 status checks and on-instance components such as the ECS agent and container runtime. When the ECS agent loses connectivity to the control plane, ECS treats the instance as impaired, running a continuous monitor-and-repair loop to ensure the affected tasks are replaced and limits are not exceeded. This automated recovery mechanism helps maintain service stability and prevents cascading failures.

Written by urgent.news from The New Stack's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at thenewstack.io →

More in Tech

Governance Attack Surface Review: ether.fi Stake

Governance Attack Surface Review: ether.fi Stake Target Protocol : ether.fi Stake (TVL: $4763.4M) Governance Attack‑Surface Review – ether.fi Stake Prepared by: [Your Firm] – Senior DeFi Security…

How a Tool Directory Stays Useful When It Has Too Many Tools

The hardest part of a small utility site is not adding one more calculator. It is helping a visitor find the right one without making them learn the site's architecture first.

  • Be Good Tool homepage uses Vue framework for navigation and search integration
  • Catalogue list imported as composable, maintaining separate data structures
  • Structured data and responsive visuals adapt to different device views

More from Friday 9 October →