Urgent.News

What's breaking now, across thousands of outlets.

AI

Your AI Agent Will Do Something Terrible. Here's How to Survive It.

Here's a pattern I keep seeing. A team wires up an AI agent that can do real things — send emails, run commands, query and modify the database, call external APIs. The demo is magical. It reads a request, figures out the steps, takes them, reports success. Everyone's impressed, and it ships. Then comes the first incident. It emails the wrong list. It runs a destructive command against the wrong…

The article discusses the potential dangers of AI agents and provides a checklist to ensure their safe deployment. The first key principle is least privilege, which involves granting an agent only the minimum permissions necessary for its task. This prevents the agent from causing widespread damage if it makes an error. The article emphasizes the importance of giving the agent exactly what it needs and nothing more, rather than providing broad access out of convenience.

The second principle is human approval for consequential actions, such as sending money, deleting data, or deploying software. The article advises gating these high-impact actions behind a human review process, rather than allowing the agent to act on its own. However, it also warns against over-gating, as excessive approvals can make the human approval process feel like a formality rather than a genuine safety measure.

The review should provide meaningful context about the action being taken, allowing the human approver to make an informed decision.

The article also advises treating all input the agent receives as untrusted. Since AI agents can ingest data from various sources, including web pages, emails, and documents, they may encounter malicious instructions disguised as legitimate content. This technique, known as prompt injection, can lead to unintended actions if the agent blindly follows the provided instructions.

The solution is to gate the agent's ability to perform high-impact actions based on enumerated, high-level intents rather than trusting the agent's interpretation of the input content.

Finally, the article stresses the importance of an independent audit trail for the agent's actions. If the agent is responsible for writing its own logs, it may produce misleading or incomplete records of its actions. By storing audit logs in an external, immutable location that the agent cannot control, you create an independent source of truth that can be reviewed in the event of an incident. This ensures that even if the agent is compromised, there will still be a reliable record of its behavior.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

The ML you need to operate LLMs, not train them

You do not need to understand backpropagation to run a large language model well in production. You need a smaller, more practical thing: the operator's mental model.

  • Tokens differ from words; longer or uncommon words split into multiple tokens.
  • Temperature and topp parameters are model-specific constraints, not interchangeable dials.

Recursive Governance: When Agents Write the Rules They Execute

We almost lost forty modules to a file that never changed. The sync job ran every six hours, mirroring our shared working directory into a test workspace.

  • Agents create and enforce their own rules, leading to conflicts and inconsistencies.
  • Recursive governance system protects modules and enables rule auditing and validation.

Best LLM Gateways in 2026: A Production-Ready Comparison

TL;DR An LLM gateway is production-ready when it adds negligible latency under load, fails over across providers without application code, enforces budgets per team, governs MCP tool calls, and…

  • Bifrost is open-source Go-based LLM gateway with high performance
  • Favored for enterprises needing scalability and reliability
  • Not published details for Kong AI Gateway and Cloudflare AI Gateway

More from Tuesday 6 October →