Urgent.News

What's breaking now, across thousands of outlets.

Tech

Our ECS Fargate Task Was Silently Failing for Days — Here's Exactly How We Found It

Our ECS Fargate Task Was Silently Failing for Days — Here's Exactly How We Found It No alarm fired. No Slack alert. No PagerDuty page. Just a service quietly burning compute while cycling through thousands of failed tasks. Here's the full story. It started with a routine check. I opened the ECS console to verify a deployment and noticed something wrong in the numbers. The running task count was…

This article recounts a story where an ECS Fargate task was silently failing for days without any alarms or alerts. The service was cycling through thousands of failed tasks, burning compute resources. The root cause was a health check configuration mismatch - the health check path (/api/health) didn't match the actual application path (/health).

Three key signals pointed to the issue: the ECS console events tab filling up with failed tasks, the RunningTaskCount metric dropping to zero, and the absence of application logs. The reporter outlines three measures they've implemented to prevent such issues: validating health check paths before deployments, setting up an alarm on the RunningTaskCount metric, and streaming ECS service events to CloudWatch Logs.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in Tech

More from Wednesday 26 August →