{
  "id": 12920792,
  "title": "Critical or non-critical: the failure contract of ECS daemons",
  "url": "https://urgent.news/2026/10/08/critical-or-non-critical-the-failure-contract-of-ecs-daemons",
  "topic": "tech",
  "section": "Tech",
  "published": "2026-10-08T18:00:58.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/fernando_azevedo_6844e930/critical-or-non-critical-the-failure-contract-of-ecs-daemons-505i"
  },
  "original_language": "en",
  "account": "Amazon's launch of non-critical daemons in Amazon ECS Managed Daemons raises an important question: what happens to a machine when its critical logging or security agent fails? The answer is now a new parameter called \"critical\". This simple addition to the daemon contract exposes a deeper problem with the current sidecar-per-task model in container clusters. By treating log shippers, EDR agents, metrics exporters, and mesh proxies as separate sidecars, we duplicate overhead without truly understanding the economic and governance implications.\n\nTraditionally, ECS had a DAEMON scheduling strategy per EC2 instance, but this provided no guarantee of the daemon's readiness before the application tasks started serving traffic. Managed Daemons resolves this by ensuring an instance only transitions to \"RUNNING\" status after its daemon is up and running, closing the observability gap and providing a security layer. However, this guarantee comes at the cost of rigid coupling - any daemon failure immediately drains and replaces the entire instance, ideal for compliance but potentially absurd for lightweight metrics exporters.\n\nThe new non-critical mode introduces a switch that governs remediation behavior. When set to \"fail open\", the instance remains active even if the daemon fails, allowing existing tasks to continue running without interruption. This is suitable for non-critical daemons like metrics exporters. In contrast, the default \"fail closed\" mode blocks instance registration and replaces the whole instance if the daemon fails, providing stronger security guarantees at the expense of resilience.\n\nThe key takeaway is that the choice between critical and non-critical modes should be driven by observability and remediation requirements, not just the need for a daemon to be up before traffic. By inverting the relationship between instance readiness and daemon health, Managed Daemons offers a more robust foundation for managing sidecar agents in containerized environments.",
  "summary": "Every container cluster carries a layer of software nobody asked for and everybody needs: the log shipper, the EDR agent, the metrics exporter, the mesh proxy. For years that layer was treated as a packaging detail, one more sidecar in the task definition, a DaemonSet on Kubernetes. The launch of non-critical daemons in Amazon ECS Managed Daemons, on September 3, 2026, exposes the question that…",
  "key_points": [
    "Amazon introduces non-critical daemons in ECS Managed Daemons.",
    "Critical logging and security agents now have a 'critical' parameter.",
    "Non-critical daemons can 'fail open' to maintain instance activity."
  ],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}