{
  "id": 9328638,
  "title": "Why \"Wait Until It Breaks\" Doesn't Scale: The Math Nobody Runs on Reactive vs. Proactive Infrastructure",
  "url": "https://urgent.news/2026/09/23/why-wait-until-it-breaks-doesnt-scale-the-math-nobody-runs-on",
  "topic": "tech",
  "section": "Tech",
  "published": "2026-09-23T11:54:04.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/aaronsmithcs/why-wait-until-it-breaks-doesnt-scale-the-math-nobody-runs-on-reactive-vs-proactive-4bid"
  },
  "original_language": "en",
  "account": "Most teams can easily price a monitoring tool, but few understand the true cost of not having monitoring systems. Monitoring tools have a cost, but so does not having monitoring - a cost that most people only realize when it's too late. For instance, if nobody notices a certificate expiration due to lack of monitoring, it can lead to critical failures. As incidents unfold, overtime is spent fixing problems and managing fallout. In retrospectives, teams often wonder why monitoring wasn't invested in earlier. This highlights the cost of reactive, \"wait until it breaks\" infrastructure management.\n\nWhen asked to price proactive fixes, engineers can quickly provide a figure. However, when asked to price the alternative - inaction - the conversation stalls because no one has calculated the risk. Reactive spending feels free until it's too late, while proactive spending appears as a predictable monthly line item. Studies show that downtime can cost over $300,000 per hour for many enterprises, not including legal and regulatory costs. Additionally, 85% of major outages in the past three years were due to human error from skipped or broken procedures, not one-offs.\n\nProactive infrastructure management reduces recurring risks by shifting from reactive to preventive measures. The real math behind this isn't a single number but a sum of various costs. By calculating your own outage history and revenue, the total is often higher than the casual dismissal in planning meetings. Reactive (wait until it breaks) infrastructure has detection delays of 30-120 minutes, often reported by customers, while proactive infrastructure offers seconds to minutes via automated telemetry. Resolution for reactive issues involves hours of manual work, whereas proactive methods use automated failover, rollback, or restart. Engineering impact from reactive incidents leads to roadmap delays, on-call fatigue, and missed maintenance windows. Proactive infrastructure reduces toil, as seen in Google's container orchestration, where automatic restarts occur without human intervention. To improve, move away from reactive operations and embrace proactive infrastructure management, building systems that catch and correct problems before they require human intervention.",
  "summary": "Most teams can price a monitoring tool in five minutes. Almost none can price what an hour of downtime costs. Here is the math that changes it. Monitoring tools come at a cost. However, there’s also a cost to not having monitoring systems, and it’s a cost most people don’t understand until it happens to them. For example, if nobody in the company notices that a certificate has expired because no…",
  "key_points": [
    "Reactive infrastructure incurs hidden costs until failure occurs.",
    "Proactive infrastructure management offers predictable monthly expenses.",
    "Automated telemetry reduces detection delays to seconds, vs. reactive 30-120 minutes."
  ],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}