Urgent.News

What's breaking now, across thousands of outlets.

Tech

Trustworthy AI starts with surviving production failures

Reliable AI agents recover from failures, prove identity, and contain breaches before damage spreads.

Trustworthy AI starts with surviving production failures

Trustworthy AI systems must be designed with the expectation of encountering production failures. It's not enough to simply assess whether an agent completed a task accurately or whether a demo went well. The crucial factor is how the system behaves when things go wrong, accounting for the 30% of instances where an issue arises.

In financial services, a bot that mishandles money transfers could lead to legal liability, while in healthcare, unchecked data access could compromise patient safety and result in HIPAA violations. Industries with high stakes can't afford to consider failure a minor issue.

Currently, popular agent frameworks are often built by teams focused on integrating large language models with reasoning loops and evaluation. However, these frameworks lack the expertise in distributed systems engineering, such as recovery, consistency, and fault isolation. Many of these frameworks don't account for what happens when an agent fails during execution, which can be extremely costly once they begin handling real-world transactions.

Replaying from the beginning every time an agent fails is not an efficient solution. For instance, a sales agent working through a 10-step workflow that fails at step nine would require starting the entire process again. This approach not only wastes resources but also creates a poor experience for downstream systems and people.

In regulated workflows, it becomes a compliance and audit problem waiting to happen. The solution lies in durable execution: checkpointing that captures progress at each significant step. This way, recovery involves resuming from step nine instead of repeating the entire sequence. Distributed systems have long practiced this method, and agent orchestration should adopt it as well.

When evaluating agent frameworks, a key question to ask is: does the system resume from where it left off in case of a failure, or does it start over? The answer will reveal whether the framework is designed for production or merely for demonstrations.

Another issue is the access problem in agentic systems. Security in these systems presents a unique challenge, especially as agents act on behalf of other agents and humans. Each layer of delegation adds ambiguity about accountability. When MCP servers, which grant language models access to company records, patient data, and internal systems, are involved, the risk multiplies.

Most MCP servers in production today connect to databases, and a common mistake is granting broad access to the entire data store instead of narrowly defining what each server can access. This distinction becomes critical if a system with broad access suffers a breach or supply chain compromise, allowing the attacker to gain access to the entire environment.

Implementing scoped access, where an MCP server or agent can only reach the specific data needed for its task, is an overlooked but crucial design decision. It's relatively inexpensive to fix before deployment but becomes extremely expensive to rectify after a breach. Identity also plays a significant role. Many organizations still rely on traditional authentication protocols designed for human users, not for autonomous systems that operate at machine speed and scale.

Cryptographic attestation, a tamper-proof record of events tied to the specific identity that performed them, is needed. This record enables replaying the system's state after the fact to determine with certainty that a specific piece of code accessed a specific system at a specific moment, ensuring that malicious code can't impersonate a legitimate part of the system undetected.

Lastly, organizations often focus on patching known vulnerabilities (CVEs), but this reactive approach misses the bigger picture. A vulnerability only becomes a known CVE after a breach has occurred. Runtime enforcement, which focuses on containment rather than just prevention, is far more important. In the event of a compromise, can you detect abnormal behavior and immediately restrict a component's access, even before identifying or patching the underlying flaw?

Applying zero trust principles specifically to AI workloads—not just inheriting them from traditional cloud-native security protocols—is crucial. This means enforcing strict authentication and limiting access to only what's necessary, ensuring that AI systems are both secure and trustworthy in production environments.

Written by urgent.news from TechRadar's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at techradar.com →

More in Tech

More from Wednesday 12 August →