Anthropic says auto mode will be the default in Claude Code for Pro, Max, Team plans, starting on Aug. 14, claiming it's good enough at catching harmful actions (Simon Willison/Simon Willison's Weblog)
This was one of the topics discussed in our Fireside Chat with Cat Wu and Thariq Shihipar at the AI Engineer World's Fair last month.
Anthropic has announced that auto mode will become the default setting for Claude Code in Pro, Max, and Team plans, starting on August 14th. According to Simon Willison, Anthropic is highly confident in the auto mode's safety, stating that almost every person in their organization uses auto mode. They claim that Claude Code's auto mode effectively mitigates most risks, such as prompt injection and data exfiltration, with only a small percentage of cases remaining unprotected.
However, Willison acknowledges that there are still two safety issues to address: accidental damaging actions and prompt injection. He highlights Anthropic's evaluation by Trajectory Labs, which tested various Claude Code models, including Claude Fable 5, Opus 5, and Sonnet 5, and found that none of the 720 attack attempts succeeded when running auto mode. Willison remains skeptical and calls for more independent confirmation of Anthropic's claims.
Written by urgent.news from Techmeme's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.
This story
This is one outlet's version. Read the fullest account.
