Urgent.News

What's breaking now, across thousands of outlets.

Tech

Four in ten of our requests were our own health checks

In May we moved our account API onto a service mesh, and overnight its request rate fell by forty one percent. The traffic drop alert fired at two in the morning. Users were fine. Nothing had been lost that had ever been a user. The mesh had moved health checks to a separate port that bypasses our metrics middleware. So I counted what had been going through it. Kubelet readiness and liveness…

In May, the organization shifted its account API to a service mesh, resulting in a forty-one percent reduction in request rate. Health checks, including kubelet readiness and liveness probes every five seconds on each of twenty-four pods, load balancer checks from three zones every ten seconds per target, and an uptime service polling from six locations, were moved to a separate port.

These probes were bypassing the metrics middleware, which recorded request count, latency, and status code for the organization's objectives.

During peak daytime hours, non-user traffic accounted for approximately fifteen percent of the requests, whereas overnight, it reached seventy percent. This represented four out of ten requests across the entire day. The error rate objective was significantly diluted by the high volume of traffic that could not fail. During a partial outage in March, users experienced seven point eight percent errors, while the dashboard indicated four point six percent, which was below the five percent burn threshold. Thus, the issue was reported to a customer twenty-six minutes later.

The latency percentiles were also skewed due to the fast requests that were not being waited for by users. Probes were now residing on a distinct admin port that was not instrumented as traffic. Each request metric included a caller class label with three possible values: user, internal, and synthetic. Objectives were calculated based solely on user traffic. Even though synthetic checks continued to alert, they served their unique purpose.

A panel displayed the share of non-user traffic per service, allowing the organization to identify any dilution before it resulted in any issues. After excluding the probes, the team recomputed the error budget for the last quarter. They had utilized almost twice the error budget they had reported, and had failed to meet the objective in just two months out of three.

An objective is a ratio, and it became evident that there was a need to decide who should be counted in the bottom half, or alternatively, everything that could reach the endpoint would decide on its own.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in Tech

Getting the downloads was the easy part. Retention is the real game.

I'm Ayotemi, co-founder of The Rhyme Game, a freestyle rap training app. In my last post I broke down how we hit 400K downloads in four months by making the product generate its own content.

  • Rhyme Game initially focused on generating viral downloads
  • Retention became the true challenge after initial success
  • Streaks, visible progression, and useful tool drove retention

More from Monday 28 September →