Urgent.News

What's breaking now, across thousands of outlets.

Tech

Noisy Uptime Checks: Feature-Flag Kill Switches for Retry Storms

A customer-support notification page fires: delivery failures are climbing, the checker is retrying, and the on-call sees enough noise to obscure the original fault. The immediate move is to disable the noisy uptime check with a feature flag, without waiting for a deployment, while leaving delivery itself alone. Short answer: put the probe execution behind a remotely evaluated boolean, poll at a…

A customer-support notification page reports a surge in delivery failures. The system's health checker is attempting repeated retries, causing additional noise that masks the initial problem. To address this, a feature flag is used to temporarily disable the noisy uptime check, bypassing the need for an immediate deployment. The key is to integrate the probe execution behind a remotely evaluated boolean, implementing a bounded polling interval, caching the most recent valid result, and treating the feature flag as a temporary kill switch rather than an alerting mechanism.

This approach allows for rapid mitigation of problematic probes, high-frequency checkers, or retry storms without compromising audit trails, evaluation analytics, parent-child dependencies, alert thresholds, or notification routing, which are separate operational controls. When introducing a new monitor path, gradual rollout should be employed, starting with one region or tenant before disabling it entirely, as deletion lacks a recycle bin.

Should a feature flag be responsible for halting noisy uptime checks? The notification page confirms this. By focusing on three events—customer notifications failing, the health check identifying failures, and the checker's retries exacerbating traffic—the distinction between user-impacting symptoms, evidence, and load generated by the safety system is crucial.

Reporting includes a delivery-attempt counter and a delivery-failure counter, with a failure ratio calculated over a window that resists short-term fluctuations. The observability surface supports metric reporting and querying, but lacks built-in threshold rules or notification routes, necessitating operator intervention to poll the query API and manage alert dispatching.

The query filters are undeclared, so verifying the discovery schema before constructing any filtered queries is essential. Ultimately, the SLO question revolves around the acceptable level of failed delivery within a specific window and the rate at which the current failure volume consumes that budget. Google SRE's four golden signals—errors, traffic, latency, and saturation—serve as the foundational framework for evaluating these metrics.

A raw count of failures lacks context without the corresponding number of attempts. The page emphasizes that a notification solely triggered by the checker's retries does not constitute a valid alert. In conclusion, implementing a kill switch at the probe boundary, placing the flag check immediately before the uptime probe initiates network operations, and adhering to the principles outlined above can effectively manage noisy uptime checks, ensuring stability and reliability.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in Tech

More from Sunday 11 October →