When Your AI SRE Runs Out of Telemetry: The Missing-Evidence Problem in Production Debugging
AI SRE can reason over logs, traces, and metrics, but it cannot recover evidence that was never captured. Learn how runtime evidence closes that gap.
The rise of AI SRE systems has improved the initial stages of incident investigation significantly. They can receive alerts, inspect metrics, search logs, follow distributed traces, check recent deployments, read source code, retrieve runbooks, and compare the current failure with previous incidents. However, there is a fundamental limitation to this model. Sometimes, the knowledge required to understand an incident is not present anywhere in the system. The evidence may never have been collected.
For instance, in a production checkout service experiencing intermittent HTTP 500 errors, an alert triggers an AI SRE. The system identifies an increase in error rate from 0.2% to 7.4%. Deployments and traces indicate that the issue began after a recent release and follows the same code path. Application logs show repeated InvalidRegionException errors.
The first few minutes of investigation seem promising, but there's a missing piece: the value of customer.region at the moment of failure. If this value was never logged or captured, the AI SRE cannot retrieve it, leaving a gap in the understanding of the issue.
This problem is known as the missing-evidence problem, and it's a significant limitation when evaluating AI SRE systems. While observability helps gather clues and narrow down possibilities, it doesn't always provide the complete picture. The AI SRE can form reasonable hypotheses based on the available evidence, but without the missing facts, it's challenging to verify the root cause.
There are three main problems to consider when dealing with AI SRE systems: retrieval, reasoning, and observation. The retrieval problem involves finding the right information, while the reasoning problem focuses on understanding the relationships between the gathered data. The observation problem arises when the necessary information for forming plausible explanations is not present in the telemetry.
Modern AI SREs have access to various tools and systems, such as Prometheus metrics, Sentry issues, APM traces, Kubernetes state, Git history, and deployment metadata. While this expanded access is beneficial, it doesn't automatically mean more information will be available. The evidence ceiling is reached when the required facts are simply not captured in any of these sources.
When presented with a plausible explanation based on the available evidence, it's crucial to remember that the AI SRE has reached the edge of its telemetry. LLMs are skilled at constructing coherent explanations, but treating the output as a verified root cause can be dangerous. The missing-evidence problem highlights the importance of acquiring new evidence and understanding the limitations of existing data sources in production debugging.
Written by urgent.news from HackerNoon's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.