A New Trick Reveals AI Models’ Inner Thoughts
Researchers devised a way to extract “reasoning traces” from Claude, GPT, and Gemini. What they found, they say, indicates that some Chinese AI may be trained on leading US models.
Researchers have discovered a new method that could reveal the inner workings of certain artificial intelligence models, potentially exposing sensitive information such as passwords and API keys. This technique, known as reasoning distillation, involves copying the reasoning patterns from one model to another, which is a common practice in the development of open-weight models.
The findings, however, are not conclusive proof that Chinese AI companies have been using this method to train their models. The researchers tested frontier models from OpenAI, Anthropic, and Google, which are accessed via an application programming interface (API), and found that they produced similar outputs to the hidden reasoning traces of two Chinese open-weight models, Kimi K3 and DeepSeek. However, other Chinese models, such as Inkling from Thinking Machines, did not exhibit this similarity.
While the researchers have demonstrated that distillation could be used to recover personal information from a model's inner reasoning, they note that this vulnerability has been fixed in the APIs of the major AI model providers. Nevertheless, they warn that using their method could potentially enable more information to be distilled from closed models than previously realized.
The issue raises concerns about the potential for large-scale reasoning distillation attacks, which could lead to the leakage of private information. The researchers have alerted OpenAI, Anthropic, and Google to the vulnerability, and each company has adjusted its API to mitigate the problem. However, some reasoning traces may still be uncovered using the same method, even though extracting private information is no longer possible.
The discovery of this new technique has sparked discussions about the geopolitical implications of AI model development. As US and Chinese companies compete for AI supremacy, distillation has become a contentious issue, with claims that Chinese companies are using it to copy US models. However, some experts argue that distillation is a widely used technique to enhance the capabilities of existing models, and that it is not necessarily indicative of illicit copying.
Written by urgent.news from Wired's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.
This story
This is one outlet's version. Read the fullest account.