Urgent.News

What's breaking now, across thousands of outlets.

AI

A New Trick Reveals AI Models’ Inner Thoughts

Researchers devised a way to extract “reasoning traces” from Claude, GPT, and Gemini. What they found, they say, indicates that some Chinese AI may be trained on leading US models.

A New Trick Reveals AI Models’ Inner Thoughts

A research team has discovered a method to reveal the inner workings of AI models, potentially exposing sensitive information. Researchers from the University of Tübingen in Germany, the Max Planck Institute, MATS Research, and security company Snyk tested major AI providers, finding that Chinese models may have been trained by "distilling" reasoning information from US models.

The technique, known as distillation, is commonly used to replicate the capabilities of existing models in open-weight or fully downloadable models. However, this research indicates that distillation could be used to recover personal information, such as passwords and API keys, from a model's inner reasoning. The vulnerability has been fixed in major models, but researchers believe it could still lead to large-scale reasoning distillation attacks.

The researchers used encrypted reasoning traces from smaller, cheaper models to uncover the hidden reasoning inside larger, more expensive models. This method has also revealed secret information, including API keys and passwords. While the threat has been mitigated by companies adjusting their APIs, researchers believe a complete fix would require a fundamental overhaul of API design.

The discovery has geopolitical implications as China and US companies compete for AI supremacy with increasingly powerful models.

Written by urgent.news from Wired Business's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Also reported by 1 other outlet

Read the original at wired.com →

More in AI

More from Tuesday 11 August →