Urgent.News

What's breaking now, across thousands of outlets.

Tech

Attaching Evidence to Alerts: an Enrichment Sidecar for Alertmanager

Every on-call engineer knows the ritual. A page lands in Slack: the 5xx rate on a service is above the threshold. The message tells you that something is wrong and nothing about what. So you open Kibana or Loki in another tab, set the time window to the last few minutes, filter by the service, […]

Attaching Evidence to Alerts: an Enrichment Sidecar for Alertmanager

Every on-call engineer experiences the same routine when a page arrives in Slack: a 5xx rate on a service exceeds the threshold, and the message provides little context about the issue. Engineers then open Kibana or Loki separately, set the time frame to the last few minutes, and filter by the service to begin diagnosing. This process often takes two or three minutes, placing the initial diagnosis phase at the worst moments of the incident.

The desire to embed the recent error lines directly into the alert message itself is understandable. However, implementing this within Alertmanager is not feasible, as it utilizes Go templates that only utilize the labels and annotations attached by Prometheus at the time of the rule firing. Since Prometheus lacks the capability to communicate with log stores, fetching data at send time is not a part of the design. As a result, logs must come from an external process operating alongside Alertmanager.

The solution takes the form of a small service known as a "sidecar" positioned next to Alertmanager. This design originated from two failed attempts. The initial attempt involved a bot that posted alerts to Slack with the attached logs. Although it appeared promising in a demonstration, this version became problematic because the bot assumed the role of the paging path. Any bugs or timeouts in the log store would prevent the team from being paged. Consequently, the bot was abandoned quickly.

The second attempt involved posting logs as a separate message in the same Slack channel as the alert. However, this caused the channel to become cluttered, with two messages appearing per alert. During an incident, these messages would be interleaved, making it challenging to identify the relevant logs. The solution was to reply to the alert message in the thread, preserving the one message per alert format while still providing the necessary context.

The enricher, which retrieves the relevant logs, must locate the corresponding Slack message based on a fingerprint of the alert. Initially, the matching was done by comparing the label set within the message text, and later, the fingerprint was rendered in the message to ensure an exact match. The enricher limits its search to about two minutes to prevent replying to old messages with the same labels, which would be counterproductive.

The enricher pulls the last few error-level log lines, typically around ten, based on the alert's labels, such as service, namespace, and severity. Even if no matching logs are found, the sidecar still posts an empty result, prompting the on-call engineer to broaden their investigation into infrastructure or the alerting rule itself. This additional information can be crucial in diagnosing the incident more efficiently.

The sidecar service is relatively straightforward, consisting of a few hundred lines of code. In practice, it has proven to be reliable and unobtrusive. The primary change for the team is that the first evidence of the incident appears directly in Slack, eliminating the need to navigate through multiple dashboards, indices, and time windows.

Diagnoses have been expedited across various incidents where logs were provided, without requiring any modifications to the services themselves. Moreover, the sidecar can extend its capabilities beyond just logging, delivering links to runbooks, markers of the latest service deployment, or even scoped dashboards directly relevant to the incident.

This enrichment significantly enhances the efficiency of incident response, placing the necessary context at the forefront of the alerting process.

Written by urgent.news from DevOps.com's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at devops.com →

More in Tech

More from Thursday 8 October →