Urgent.News

What's breaking now, across thousands of outlets.

AI

AI Reasoning Leak: Extracting Models' Inner Thoughts

What the Vulnerability Is A joint research effort led by Alexander Panfilov (University of Tübingen), Florian Tramer (ETH Zürich), Yarin Gal (Oxford), and Kyle Miller (Center for Security and Emerging Technologies) has uncovered a previously unknown side‑channel in frontier AI systems. The flaw allows an attacker to extract the hidden “inner thoughts” —the chain‑of‑thought reasoning traces that…

A recent joint research effort has exposed a previously unknown vulnerability in frontier AI systems that allows attackers to extract the inner reasoning processes of large, highly-aligned models. The flaw enables an attacker to intercept encrypted reasoning traces, replay them to a weaker model variant, and subsequently reconstruct the original model's chain-of-thought. This exposes sensitive information, such as passwords and API keys, that the original model was asked to reason about.

The attack exploits the design choice of sending encrypted reasoning traces to clients for local processing. Many AI providers offload heavy computation to the client's hardware, sending a short reasoning trace encrypted with a symmetric key. The decryption key is often shared across model families, creating a single point of failure: any model capable of decrypting the payload can read the raw reasoning steps.

The researchers demonstrated the attack on two proprietary models—Claude Opus 4.8 (Anthropic) and GPT 5.6 Sol (OpenAI)—and tested open-weight models like Kimi K3 (Moonshot AI), DeepSeek, and Inkling. Kimi K3 showed nearly identical reasoning traces to the closed models, suggesting it had distilled the reasoning capability of the proprietary systems. In contrast, DeepSeek and Inkling showed no such similarity.

The implications of this vulnerability are significant. Personal data, such as passwords and API keys, can be exposed if an attacker exploits the flaw. Furthermore, the fact that open models can reproduce the reasoning of closed systems raises concerns about intellectual property erosion and potential regulatory violations related to data privacy and export controls.

To mitigate the risk, providers have adjusted their APIs to redact sensitive tokens from traces, rotate decryption keys per session, and implement other security measures.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

A Week as an AI Integration Consultant

A Week as an AI Integration Consultant Most weeks I don't write much new code. I read other people's systems, draw arrows on a whiteboard, and try to figure out which of the twelve places customer…

More from Thursday 13 August →