Urgent.News

What's breaking now, across thousands of outlets.

AI

Quoting Anthropic Frontier Red Team

We evaluate several models on 100 tasks from the [internal Binary Exploitation benchmark] (selected at random), and find that GLM-5.3 develops full control flow hijacks in 4% of the trials; Claude Mythos Preview did so in 6%. Although GLM-5.3 performs below Claude Mythos Preview here, a meaningful threshold has clearly been crossed: earlier models, like Claude Opus 4.6 and GLM-5.2, do not succeed…

We haven't written up this one. Simon Willison has the full story — the link below goes straight to it.

Read the original at simonwillison.net →

More in AI

More from Tuesday 29 September →