Urgent.News

What's breaking now, across thousands of outlets.

AI

Stealing Reasoning Traces from Proprietary LLM APIs

Article URL: https://stolen-thoughts.com/ Comments URL: https://news.ycombinator.com/item?id=49257876 Points: 257 # Comments: 88

Research demonstrates that reasoning traces from proprietary large language model (LLM) APIs can be stolen. These traces closely mirror the number of hidden thinking tokens reported by the API, with a corresponding token count for the decoded reasoning. Using publicly available agent trajectories from GitHub and Hugging Face, produced by Claude, GPT, and Gemini models, the researchers decoded 315,320 reasoning blocks from 6,708 sources.

This process revealed sensitive information such as API keys, passwords, access tokens, and personal email addresses, alongside technical identifiers that were exclusively found in the reasoning blocks and not in the visible session.

Furthermore, the study found that prompting models like Kimi-K3 with the first 1% tokens of Opus 4.8's reasoning altered the visible answer to resemble Opus's wording, even though the answer itself was not prefilled. This indicates that sensitive knowledge is embedded within the hidden traces. By prompting a model to reason through harmful content without displaying the answer, the attack successfully recovers this hazardous knowledge in plaintext.

The API summary's inability to preserve the distinction between clean derivations and potentially hazardous reasoning was also evident in cases where Opus 4.8 sometimes stated the answer before deriving it.

Written by urgent.news from Hacker News Best's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Also reported by 1 other outlet

Read the original at stolen-thoughts.com →

More in AI

More from Tuesday 11 August →