Urgent.News

What's breaking now, across thousands of outlets.

Tech

Backend Uptime Alerts from One-Minute Metrics API Polling and Failure Evidence

A one-minute poller is only useful if it preserves enough evidence to explain an outage after the alert fires. For a gaming backend, the practical choice is to evaluate a low-cardinality availability metric every 60 seconds, notify through a separately owned channel, and fetch logs only after the metric crosses a threshold. Short answer: keep detection and diagnosis separate. Let aggregated…

The One-Minute Metrics API employs polling to monitor backend uptime, with the interval set to one minute. Polling preserves sufficient evidence to explain outages following alert activation. For gaming backends, it's practical to evaluate availability metrics every 60 seconds, notify through a dedicated channel, and fetch logs only when the metric exceeds a threshold.

Detection and diagnosis should remain separate, with aggregated metrics confirming service failure and logs providing timestamps, request identifiers, and error context for a specific incident window. Each notification should include an incident key to consolidate repeated polls into a single incident rather than overwhelming notification channels like Slack, email, or webhook consumers.

The design allows for a one-polling interval detection delay while avoiding the escalation of transient failed requests into alerts. To reconstruct player experiences, incident records should include the failed service and region, the time the condition first crossed the threshold, the number of consecutive evaluations failing, the rule version that made the decision, and the log window's start and end points.

Notification attempts should be recorded separately from health results, as "the service was down" and "the alert email was filtered" are distinct failures. In an OTP service within a game, a regional availability drop might appear as a login outage if the overall game remains healthy; thus, partitioning by service, environment, and region is essential.

Personal data like player IDs, match IDs, email addresses, or phone numbers should not be included in metric labels to avoid creating numerous time series in Prometheus. One minute is a sampling contract, not proof of continuous uptime. The source and evaluator timestamps must be stored so an investigator can differentiate late data from late workers.

A Node.js worker should poll uptime metrics using a state machine: healthy, pending, firing, and recovered. The number of consecutive failures before entering firing and the number of successful evaluations before recovery should be defined as business policy and versioned. The worker should handle different failure types, such as checkout failures during live events versus delayed leaderboard refreshes, by adjusting the trigger speed accordingly.

If the metrics query times out, label the result as unknown rather than silently interpreting it as healthy. On encountering an HTTP 429 response, honor the Retry-After header and apply exponential backoff with jitter. Avoid overloading the system with excessive retries due to the scheduler's minute-long wake cycles.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in Tech

Why I Use Umami Instead of Google Analytics (And Don't Have a Cookie Banner)

When I was setting up analytics for my blog, Google Analytics seemed like the obvious choice. But once I looked at what it actually requires — a tracking script, a cookie consent banner, a cookie…

  • Google Analytics requires tracking script, cookie consent banner, and GDPR compliance
  • Umami is open-source, self-hosted analytics platform without cookies or consent banners
  • Umami provides essential metrics with smaller script size and user-friendly dashboard

An Agent Fleet, Sorted By Job

"We have agents" is about as informative as "we have software". The question is what they do, when they run and who notices when one stalls. Ours grew into a fleet step by step.

  • Agent fleets are organized groups of specialized software agents
  • Leadership agents summarize agent work and generate reports
  • Review agents provide unbiased, critical analysis of agent work

Meet Mind

Introduction Have you ever entered a meeting and struggled to remember what was discussed in the previous conversation? Important details such as requirements, concerns, preferences, and follow-ups…

More from Tuesday 29 September →