Urgent.News

What's breaking now, across thousands of outlets.

AI

Your AI Agent Didn't Break the Rules. One of Your Rules Was Missing.

Anatomy of an autonomy bug: when two valid decision paths create one invalid outcome. Part 1 — For everyone The thing about autonomous agents nobody tells you Building an autonomous agent is a bit like raising a very obedient, very literal child with a credit card. You write rules. The child follows them. Perfectly. The problem is that the child follows the rules you wrote , not the rules you…

A recent bug with an autonomous coding agent called Sentinel exposed a critical issue with the design of autonomous agents that make decisions, retry actions, or require spending resources like tokens. The bug allowed Sentinel to spend tokens without any trigger or reason in two separate instances.

Sentinel is designed to operate on a schedule, pick files in its codebase, assess their worth, and consult an LLM if local analysis isn't confident enough. Tokens are limited, and every spend is audited. When Sentinel woke up on October 7 at 9:55, it spent two tokens without any trigger, six and a half hours before its scheduled autonomy window. On October 8 at 15:01, an OpenAI API hiccuped, and Sentinel spent a token to retry anyway.

Both anomalies stemmed from two valid decision paths that allowed Sentinel to spend tokens without any reason. The system's architecture had two gates into the same decision-making room: the strict AutonomyBudget.admit() and the more lenient AutonomyBudget.consider(). The consider() path, added later, did not require a trigger or respect the daily window and cooldown, allowing Sentinel to spend tokens even when it had no reason to do so.

This bug highlights the potential risks of building systems with autonomous agents that can decide, retry, or spend resources. The issue lies not in a single flawed check but in two well-functioning checks that do not communicate with each other. The architecture should ensure that each function independently verifies preconditions before authorizing any expensive action, preventing such backdoors from forming.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

What Happens When Users Don't Behave as Expected?

In the previous article, we looked at why functional testing alone is not enough for AI applications. Traditional testing usually starts with a simple assumption: developers know how users are…

  • Users often behave unexpectedly, rephrasing requests or asking unrelated questions.
  • AI systems interpret natural language, leading to responses different from developer intentions.
  • Testing must go beyond predefined scenarios to evaluate unexpected user behavior.

Gemini & Claude: Google's AI Agent Gamble

The AI Tango: When Rivals Become Partners (or, My Morning with Gemini) My morning began with a deceptively simple request for Gemini.

  • Google and Anthropic form alliance, surprising tech world
  • Google's Gemini agents will support Anthropic's Claude 3.5 Sonnet
  • Strategy positions Google as foundational layer for AI agents

Getting Started with Voice AI Development in 2026

Why Voice AI Is the Next Big Thing in 2026 Voice is the most natural way humans communicate. Whether it’s powering smart assistants, creating immersive games, or generating on‑the‑fly narration, voice…

  • Voice AI is prevalent in modern applications like smart assistants and gaming.
  • ElevenLabs platform offers developer-friendly TTS and voice-cloning features.
  • Python example provided for creating basic TTS demo with ElevenLabs API.

Portable AI Agent Memory: What Should Move When Users Switch Agents?

Your AI Agent Knows You. What Happens When You Leave? Imagine using an AI assistant for two years. It knows how you prefer reports to be structured.

  • AI agents store active conversation, task state, and user preferences during transitions.
  • Useful context should be preserved without granting operational authority to new agents.

More from Friday 9 October →