Urgent.News

What's breaking now, across thousands of outlets.

AI

Can an AI model’s ‘reasoning’ be extracted? New research fuels US-China distillation row

Can an AI model’s ‘reasoning’ be extracted? New research fuels US-China distillation row

A recent research paper may intensify the ongoing debate surrounding 'distillation', an AI training method that has been under intense scrutiny following the emergence of Chinese open-weight models like Moonshot AI's Kimi K3. Researchers have uncovered a method to extract hidden reasoning traces from these advanced AI models before they produce their final outputs.

The study, published by a team of researchers from the University of Tubingen, the Max Planck Institute, the AI safety institute MATS Research, and security company Snyk, indicates that Kimi K3 generates similar outputs to the hidden reasoning traces of Anthropic's Claude Opus 4.8 and OpenAI's GPT 5.6 Sol for specific prompts. However, the paper does not definitively prove that Chinese firms like Moonshot have 'distilled' reasoning information from US models.

The researchers noted that the findings "cannot causally establish distillation." Additionally, an open-weight model called Inkling from US-based Thinking Machines and one from China's DeepSeek did not show similarities in reasoning with the Claude Opus model. This research comes amidst a fierce competition between the US and China, with distillation serving as a critical flashpoint.

The technique of distillation, which efficiently transfers the capabilities of existing models to new ones, has been a well-established practice in AI/ML for years. However, it has recently become a contentious issue. Earlier this year, OpenAI accused Chinese AI startup DeepSeek of copying one of its models to create the R1 reasoning model, and Anthropic accused Alibaba of the same.

In response, Meta CEO Mark Zuckerberg defended distillation, stating that it is a core principle of the open-source ecosystem and warned that restricting the practice would put the US at a disadvantage. The study reveals that chain-of-thought reasoning in closed, proprietary models like Claude Opus 4.8 remains concealed to prevent other models from training on it.

However, researchers discovered this information by feeding encrypted reasoning traces to a smaller model of the same family. These smaller models, which have not undergone extensive alignment training like their larger counterparts, are more likely to reveal the reasoning information. The researchers cautioned that this method could potentially be abused to extract sensitive user data, such as passwords and API keys, from models.

This vulnerability was reported to tech companies, who promptly patched the issue and adjusted their model APIs to mitigate the risks.

Written by urgent.news from The Indian Express's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at indianexpress.com →

More in AI

More from Friday 14 August →