Urgent.News

What's breaking now, across thousands of outlets.

Tech

Cutting Alert Noise in Healthcare Operations: How We Took the Pager From Hundreds a Week to Dozens

Every on-call engineer in healthcare IT knows the 3 a.m. page that turns out to be nothing. A disk at 81% on a server that has sat at 81% for six months. A CPU spike on a batch job that spikes every night by design. A "service down" alert for an endpoint that was down for forty seconds during a scheduled restart. None of these are incidents. All of them wake someone up. And in a hospital…

Every on-call engineer in healthcare IT is familiar with the late-night paging that turns out to be inconsequential. Disk utilization at 81% on a server that has remained the same for months, CPU spikes on scheduled batch jobs, and service downtime alerts for endpoints that were down during scheduled restarts are all examples of such noise.

Not only do these alerts disrupt sleep, but they also condition staff to ignore critical alerts. This piece explains how our operations team reduced the number of alerts from hundreds per week to a few dozen, maintaining all vital monitorings. The methodology employs AppDynamics and LogicMonitor with ServiceNow, but the principles are applicable to any stack. The crux of the problem lies in the rules, not the tools.

First, tally alerts before tweaking thresholds. Exporting four weeks of alerts from ServiceNow and categorizing them by rule, configuration item, and outcome—human action, or flag as noise—revealed that a dozen rules generated the majority of alerts. Some rules never resulted in human intervention, such as a disk space warning that triggered over a thousand times monthly across the network.

The critical alerts, like interface queue depth, database log growth, and application login failures, constituted a small fraction of the total and were lost in the noise. The first takeaway is to perform this count, as it often reveals the real situation contrary to intuition.

The second step is to ensure every alert specifies an action for human response. Instead of generic alerts like "disk at 80%," the new standard is "disk at 80% and growing more than 2% per hour—extend the volume or purge the log directory." This not only eliminates third of the rules but also provides an initial runbook for each surviving alert.

Third, replace static thresholds with rate and duration criteria. Many alerts were triggered by single samples, such as a server at 95% CPU for a few seconds. Adjusting the logic in LogicMonitor from "value above X" to "value above X for N consecutive polls" and using rate-based alerts instead of level-based ones significantly reduced the noise.

In AppDynamics, dynamic baselines with a minimum deviation window were used to normalize spikes from nightly batch jobs. Applying rate-based triggers to hospital-specific issues, like interface queue depth and database log growth, cut the alert volume by roughly two-thirds.

Fourth, maintenance windows are mandatory. A considerable number of remaining alerts stemmed from planned changes, patching windows, and backup jobs. By integrating maintenance windows into the change process—automatically suppressing alerts during scheduled changes—overlapping noise from manual silencing was eliminated. This step also curbed the habit of forgetting to re-enable alerts post-maintenance.

Fifth, correlate alerts before paging. The final alerts were mostly genuine, but multiple alerts from the same issue caused redundant pages. Correlating alerts based on topology and time—grouping alerts from the same host or cluster within a short timeframe—allowed a single, comprehensive alert that included all affected services. This consolidation streamlined on-call responses and enabled quicker identification of root causes.

Lastly, monthly reviews of survived alerts are essential. Applications evolve, capacity changes, and thresholds can become outdated. A thirty-minute monthly review of the ten most noisy rules and their effectiveness ensures that rules remain relevant and impactful. Rules that do not lead to actionable outcomes are either tuned or retired, while those that do are maintained.

This ongoing maintenance process keeps the alert system efficient and relevant, reducing unnecessary interruptions while ensuring timely notifications of critical incidents.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in Tech

Treat Every ID as an Authorization Claim

Authentication answers one important question: who made this request? It does not answer a second question that often hides inside an ordinary-looking query parameter: may this caller name the record…

  • Treat identifiers as authorization claims, not just query inputs.
  • Verify identifier scope against caller's authorized access before querying.
  • Reject cross-scope identifiers outright for clear, observable security.

More from Wednesday 30 September →