Urgent.News

What's breaking now, across thousands of outlets.

Tech

Claude Code puts auto mode in the driver's seat

Walk away and hope the classifier catches anything irreversible or destructive

Claude Code puts auto mode in the driver's seat

Anthropic is set to launch auto mode as the default setting for Claude Code starting August 14, touting its classifier as being as secure or more secure than an average user manually navigating prompts. Existing users on Pro, Max, and Team plans who have a different default set could encounter a one-time prompt to confirm whether they wish to switch.

This change will initially be opt-in for Claude Enterprise, Claude API, Claude Platform on AWS, Amazon Bedrock, Google Cloud's Agent Platform, and Microsoft Foundry. Anthropic intends to roll out auto mode across these services within the next month. The company has also ceased charging Pro, Max, and Team users for the additional tokens used by the classifier and plans to extend this to other platforms.

Auto mode was initially introduced in March as a research preview and became generally available on July 10. It provided an alternative to Claude Code's usual default permissions, which required manual approval for every file write and bash command, making it impractical for large tasks. The alternative was the --dangerously-skip-permissions flag, which, as its name implies, allowed Claude to operate without those checks, potentially leading to risky or destructive outcomes.

Claude Code's auto mode sends each tool call through a classifier designed to prevent irreversible, destructive, or actions outside the user's environment. If the classifier blocks something, Claude tries to find a safer alternative. If there are three consecutive blocks or 20 blocks in a session, Claude Code switches back to manual approvals.

After several months of testing, Anthropic concluded that auto mode is as secure or more secure than an average user manually going through prompts. Tests included internal red-teaming, third-party red-teaming and prompt-injection evaluations, a controlled study involving 1,053 paid testers, and an analysis of real production sessions.

In the controlled study, testers only managed to catch a deliberately inserted dangerous command 13.6 percent of the time, while auto mode blocked 89 percent of the same commands. According to Anthropic, users of Claude Code approve 97 percent of permission prompts, suggesting that the human checkpoint may be redundant. The company provided charts showing the superiority of its new feature compared to competitors, demonstrating that auto mode blocked all 720 attack attempts tested, whereas GPT-5.6 Sol running Codex's Auto-review mode let 5.83 percent of attacks through.

Additionally, Anthropic highlighted three potential dangerous actions that auto mode successfully stopped: an off-network data leak, a destructive mass operation, and a privilege escalation. Anthropic stated that in all cases, Claude found a safer alternative or sought user approval before proceeding.

Written by urgent.news from The Register Science's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

This story

This is one outlet's version. Read the fullest account.

Read the original at theregister.com →

More in Tech

I, Spied

We’re looking at the little black cameras that are causing a national uproar.

More from Monday 10 August →