Urgent.News

What's breaking now, across thousands of outlets.

AI

Agents operate, humans govern: Scale your operations and reduce toil with Azure SRE Agent

What if engineers could spend their time building and optimizing systems rather than maintaining them? It’s 3 a.m., and the The post Agents operate, humans govern: Scale your operations and reduce toil with Azure SRE Agent appeared first on The New Stack .

Agents operate, humans govern: Scale your operations and reduce toil with Azure SRE Agent

In the fast-paced world of software engineering, the blaring pager at 3 a.m. signals a critical incident. The on-duty Site Reliability Engineer (SRE) must rapidly navigate through numerous monitoring tools, deployment histories, incident systems, and runbooks, attempting to discern the root cause and rapidly remedy the issue to prevent further customer impact.

However, the emergence of Azure SRE Agent promises to revolutionize this process. According to Sanchit Mehta, a senior engineer for Azure SRE Agent, the tool begins analyzing telemetry and correlates key factors such as blast radius, deployment changes, and recent rollouts to identify potential causes. Moreover, it can generate pull requests for the proposed fixes.

The support provided by Azure SRE Agent proves beneficial even during regular business hours. At InEight, correlating telemetry across tens of thousands of Azure resources can take days, but with Azure SRE Agent, identifying the affected product, tracing the performance issue to its root cause, and recommending scaling Redis can be achieved swiftly.

Microsoft has integrated Azure SRE Agent across over 3,000 service teams to streamline incident response, root cause analysis, and automated mitigations. The agent has already handled more than 1.8 million incidents within the company, mitigating many within minutes. Beyond incident management, Azure SRE Agent aids in developing and refining the service itself, including code review, deployment, evaluation, and monitoring.

This "agent-powered engineering" strategy capitalizes on ongoing AI advancements. By proactively detecting problems like quota issues, the agent can automatically generate support tickets to resolve them. For instance, the agent recently pinpointed the root cause of a synthetic test failure as soon as the change reached the first region, identifying an upstream PyPI package breakage and recommending immediate rollback and test addition.

While some internal teams have seen more than half of incidents autonomously managed by the SRE agent, such 'safe' operations typically require human governance to establish guidelines and provide initial coaching. As coding increasingly relies on agents, the reasoning loop has matured to the point where agents can autonomously handle complex problems, particularly in correlating across multiple data sources.

However, harness engineering, which combines enterprise-grade systems, auditing, and validation, is essential to ensure that agents are trusted, auditable, and governable.

Written by urgent.news from The New Stack's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at thenewstack.io →

More in AI

OpenAI, Anthropic warning to slow down AI development is encouraging, Technion professor to 'Post'

"I find it encouraging when the people leading the development of this technology speak openly about the risks and take steps to address them," Technion professor Yaniv Romano said.

  • OpenAI CEO Sam Altman and Anthropic CEO Dario Amodei warn against rapid AI development.
  • Technion professor Yaniv Romano praises leaders' candid discussions about AI risks.
  • Concerns about rogue AI agents accessing sensitive systems and causing harm.

More from Tuesday 15 September →