Urgent.News

What's breaking now, across thousands of outlets.

Tech

Critical or non-critical: the failure contract of ECS daemons

Every container cluster carries a layer of software nobody asked for and everybody needs: the log shipper, the EDR agent, the metrics exporter, the mesh proxy. For years that layer was treated as a packaging detail, one more sidecar in the task definition, a DaemonSet on Kubernetes. The launch of non-critical daemons in Amazon ECS Managed Daemons, on September 3, 2026, exposes the question that…

Amazon's launch of non-critical daemons in Amazon ECS Managed Daemons raises an important question: what happens to a machine when its critical logging or security agent fails? The answer is now a new parameter called "critical". This simple addition to the daemon contract exposes a deeper problem with the current sidecar-per-task model in container clusters.

By treating log shippers, EDR agents, metrics exporters, and mesh proxies as separate sidecars, we duplicate overhead without truly understanding the economic and governance implications.

Traditionally, ECS had a DAEMON scheduling strategy per EC2 instance, but this provided no guarantee of the daemon's readiness before the application tasks started serving traffic. Managed Daemons resolves this by ensuring an instance only transitions to "RUNNING" status after its daemon is up and running, closing the observability gap and providing a security layer.

However, this guarantee comes at the cost of rigid coupling - any daemon failure immediately drains and replaces the entire instance, ideal for compliance but potentially absurd for lightweight metrics exporters.

The new non-critical mode introduces a switch that governs remediation behavior. When set to "fail open", the instance remains active even if the daemon fails, allowing existing tasks to continue running without interruption. This is suitable for non-critical daemons like metrics exporters. In contrast, the default "fail closed" mode blocks instance registration and replaces the whole instance if the daemon fails, providing stronger security guarantees at the expense of resilience.

The key takeaway is that the choice between critical and non-critical modes should be driven by observability and remediation requirements, not just the need for a daemon to be up before traffic. By inverting the relationship between instance readiness and daemon health, Managed Daemons offers a more robust foundation for managing sidecar agents in containerized environments.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in Tech

Cacti, LibreNMS, Icinga and Observium: measuring the monitoring tier nobody replaces

Cacti, LibreNMS, Icinga and Observium: measuring the monitoring tier nobody replaces Monitoring platforms accumulate. They are installed by an engineer who has since moved teams, they hold credentials…

  • Cacti and LibreNMS dominate monitoring platforms with 17,703 and 15,614 instances respectively
  • Icinga and Observium have fewer instances, indicating smaller user bases
  • All monitoring platforms store sensitive credentials, making them prime attack targets

We Built 27 Free, Zero-Auth Web Calculators for Developers & Founders

Most modern online calculators are bloated with intrusive banner ads, demand an email address, or send your sensitive financial numbers to remote cloud databases.

  • 27 free, zero-auth web calculators launched for developers and founders
  • All tools perform client-side processing with no data transmission
  • Notable calculators include SQLite storage estimator and S-Corp tax savings tool

More from Thursday 8 October →