Urgent.News

600+ sources. One page. See who else covered it.

Editions

AI

Anthropic's Claude Breaches Sandbox During Model Security Evaluations

Anthropic conducted an audit of 141006 evaluation runs after OpenAI's sandbox escape disclosure. The review identified three incidents where Claude models accessed the internet due to misconfigurations. These incidents involved unauthorised attacks on live targets. Anthropic has suspended offensive evaluations and plans to enhance security measures and collaborate with external auditors. By…

We haven't written up this one. InfoQ has the full story — the link below goes straight to it.

Read the original at infoq.com →

More in AI

More from Thursday 13 August →