Urgent.News

What's breaking now, across thousands of outlets.

Tech

Your Uptime Monitor Says 200 OK and Your Site Is Still Broken

If you run production web apps, you have probably watched a green dashboard during an outage at least once. This post covers the gaps a basic uptime check leaves open, and the few settings that close them without flooding your alert channel. What a Basic Uptime Check Actually Proves A standard uptime monitor sends an HTTP request every one to five minutes and records whether a 2xx status came…

Many production web applications experience outages, and a basic uptime monitor may seem sufficient, but it falls short in detecting critical issues. A standard uptime check sends an HTTP request to verify if the server responds with a 2xx status code, but it cannot guarantee that the web page content is functioning correctly. A bad deployment, a broken template, or a cached error page can all return a 200 OK status, while the user sees nothing.

To address this, add keyword checks and response assertions to confirm that the page body contains the expected content. For instance, look for a product name or a footer line that should appear when the page renders real content. Furthermore, monitor the page size to detect blank pages caused by errors. A basic uptime check assumes a 99.9% uptime allows about 43.8 minutes of downtime per month, but a single bad deployment could consume the entire budget in one incident.

Achieving 99.99% uptime leaves only approximately 4.4 minutes of allowable downtime, necessitating sub-minute checks and automated failover to prevent significant disruptions. To reduce alert noise, run each check from at least three different geographic regions and only alert when two or more fail. Require consecutive failures before notifying the team.

Set response time thresholds at roughly two to three times the baseline measured during a stable week and adjust them as traffic increases. Define maintenance windows to avoid paging the on-call engineer during planned deployments. Certificate expiry should also have its own alert, warning at least 30 days in advance to give sufficient time for renewal, even if using Let's Encrypt with automatic renewal, as renewal can fail silently due to DNS changes, server migrations, or configuration drift.

To monitor the content rather than just availability, focus on the ten pages or flows that drive revenue instead of checking all hundred existing pages. Run synthetic checks, such as logging in, adding an item to the cart, and reaching checkout, to catch broken payment steps that page-level checks may overlook. Apply this approach to the most crucial parts of the website, and you'll have a more robust monitoring system that minimizes downtime and alerts the appropriate personnel to address issues promptly.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in Tech

How Did We Get Here? A Decade of Building on Qlik

Let me get something out of the way first. This is a frustration post. I have been working in software engineering for more than 13 years, and I have been building with and extending Qlik since Qlik…

  • Author frustrated by Qlik's evolution over a decade since 2016 Qlik Sense 2.2 release
  • Trusted Extension Developer community of ~30,000 members created open-source tools

More from Wednesday 30 September →