Measuring LLMs’ Ability to Perform Cryptanalysis
There’s new benchmark measuring AI’s ability to perform mathematical cryptanalysis. Anthropic’s frontier model actually found new attacks. The benchmark: “ CryptanalysisBench: Can LLMs do Cryptanalysis? ” The idea is to benchmark the ability of LLMs to discover new mathematical cryptanalytic attacks against a series of historical algorithms. Abstract: Cryptanalysis—the task of finding attacks…
A new benchmark, titled "CryptanalysisBench: Can LLMs do Cryptanalysis?" has been introduced, measuring the ability of AI language models to perform mathematical cryptanalysis. This benchmark focuses on assessing LLMs' ability to discover new attacks against historical cryptographic algorithms. The benchmark presents three tiers, each with 191 tasks across six families of cryptographic primitives, such as block ciphers and hash functions.
Five frontier models, including Claude Opus 4.8, Sonnet 5, Mythos 5, GPT-5.5, and GLM-5.2, were tested on the CryptanalysisBench. The models demonstrated the ability to break 65-86% of Tier 1 schemes, 61-62% of Tier 2 schemes at their full strength, and 24-61% of scaled-down variants. Beyond applying known methods, the models also produced novel cryptanalysis, including key-recovery attacks against the SpoC AEAD and an error in KINDI's published CCA-security proof – both of which were not previously known.
The release of CryptanalysisBench aims to monitor if (or when) AI cryptanalysis becomes a serious factor and to serve as a platform for stress-testing cryptographic schemes before deployment. The new attacks discovered by the models provide an early glimpse of a rapidly evolving frontier that might soon match or surpass the current state of the art in cryptanalysis.
Anthropic used the benchmark to test Mythos Preview and found vulnerabilities in Hawk and reduced-round AES. However, these findings represent early results, and the field remains an area to watch closely.
Written by urgent.news from Schneier on Security's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.