Urgent.News

What's breaking now, across thousands of outlets.

AI

Anthropic says auto mode will be the default in Claude Code for Pro, Max, Team plans, starting on Aug. 14, claiming it's good enough at catching harmful actions (Simon Willison/Simon Willison's Weblog)

This was one of the topics discussed in our Fireside Chat with Cat Wu and Thariq Shihipar at the AI Engineer World's Fair last month.

Anthropic has announced that auto mode will become the default setting for Claude Code in Pro, Max, and Team plans, starting on August 14th. According to Simon Willison, Anthropic is highly confident in the auto mode's safety, stating that almost every person in their organization uses auto mode. They claim that Claude Code's auto mode effectively mitigates most risks, such as prompt injection and data exfiltration, with only a small percentage of cases remaining unprotected.

However, Willison acknowledges that there are still two safety issues to address: accidental damaging actions and prompt injection. He highlights Anthropic's evaluation by Trajectory Labs, which tested various Claude Code models, including Claude Fable 5, Opus 5, and Sonnet 5, and found that none of the 720 attack attempts succeeded when running auto mode. Willison remains skeptical and calls for more independent confirmation of Anthropic's claims.

Written by urgent.news from Techmeme's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

This story

This is one outlet's version. Read the fullest account.

Read the original at simonwillison.net →

More in AI

Teaching AI to speak India

Artificial intelligence can chat, translate and answer questions but only if it understands your language. At IIT Madras, researchers at AI4Bharat are travelling across India, collecting voices…

More from Sunday 9 August →