Urgent.News

What's breaking now, across thousands of outlets.

AI

Anthropic’s Opus 5 Is Better at Resisting Prompt Injection

The chart is interesting. On the IPI benchmark, Opus 5 improved over Opus 4.8, reducing the probability of an attacker succeeding within 15 attempts from 5.5% to 2.0%, and from 0.5% to 0.2% on 1 attempt. It also improved on Sonnet 5 (5.9% at k=15) and Mythos 5 (2.6%), making it the most robust model evaluated. Opus 5 also outperformed all non-Claude models on this benchmark. The most robust…

Abstract editorial illustration

Anthropic's latest model, Opus 5, demonstrates improved resistance to prompt injection attacks. According to the Interactive Prompt Injection (IPI) benchmark, Opus 5 significantly reduces the likelihood of a successful attack compared to its predecessor, Opus 4.8. The probability of an attacker succeeding within 15 attempts decreases from 5.5% to 2.0%, and from a mere 0.5% to an even more negligible 0.2% when just one attempt is made.

Moreover, Opus 5 outperforms other evaluated models, including Sonnet 5 (5.9% at k=15) and Mythos 5 (2.6%), securing its position as the most robust model tested. Among non-Claude models, Muse Spark emerges as the most resilient, with a 16.5% success rate within 15 attempts, which is nearly eight times higher than Opus 5's rate.

Conversely, the most capable GPT 5.6 variant, Sol, exhibits comparable performance to its predecessor GPT 5.5, with a 20.0% success rate within 15 attempts, making it 10 times more susceptible to attacks than Opus 5 at 2.0% after fifteen tries. Other GPT 5.6 variants, such as Terra and Luna, exhibit even lower robustness, with success rates of 30.4% and 43.9%, respectively.

While it is acknowledged that prompt injection prevention is an inherently challenging task, Anthropic's Opus 5 showcases substantial progress in mitigating such attacks in specific scenarios.

Written by urgent.news from Schneier on Security's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at schneier.com →

More in AI

More from Friday 31 July →