Urgent.News

What's breaking now, across thousands of outlets.

AI

A fundamental flaw leaves LLMs strikingly vulnerable to attack

It is impossible to make large language models fully secure against hacks because of a fundamental flaw in how they work, a team of researchers argue in a paper presented at the International Conference on Machine Learning, a top AI conference, this month. The claim has huge implications for the safety of this technology, which…

Abstract editorial illustration

A team of researchers has uncovered a fundamental flaw in how large language models (LLMs) operate, making them vulnerable to attacks that could render the technology unsafe for various applications. Presented at a major AI conference, the researchers argue that this flaw is deeply rooted in the way LLMs identify the source of their instructions.

This vulnerability allows attackers to manipulate LLMs into providing sensitive information they were not trained to reveal, such as instructions for illicit activities or ways to compromise critical systems. The researchers demonstrated attacks against popular LLMs, including those from OpenAI, Anthropic, Alibaba, and DeepSeek, showing that the flaw is widespread across different models.

The core issue lies in LLMs' inability to accurately distinguish between different roles of text, such as user prompts, system instructions, or text generated internally by the model. The researchers refer to this type of attack as "chain-of-thought forgery." They discovered that by crafting text that mimics the style and content of LLM-generated chain-of-thought notes, attackers can trick the models into believing the malicious instructions originated from a trusted source, leading the models to comply with the attacker's requests.

This flaw stems from LLMs' reliance on tags to distinguish between different text roles. However, the researchers found that these tags have little impact on the model's interpretation of the text itself. Instead, LLMs seem to rely more on the style and content of the text to determine its role, making it easy for attackers to deceive the models. This weakness means that no matter how many rules or restrictions are put in place, there will always be some instructions the model can bypass.

The researchers emphasize that this vulnerability is likely unsolvable due to the nature of LLMs' architecture and their reliance on text patterns rather than explicit role tags. Companies handling sensitive applications must now grapple with the realization that LLMs may always be at risk of being compromised through cleverly crafted attacks.

Written by urgent.news from MIT Technology Review's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at technologyreview.com →

More in AI

More from Thursday 30 July →