Urgent.News

What's breaking now, across thousands of outlets.

AI

OpenAI says actors linked to China-based Moonshot AI spearheaded a campaign to extract its models’ hidden reasoning

OpenAI says operators copied its models’ encrypted reasoning and asked a model in a separate conversation to decrypt it.

OpenAI says actors linked to China-based Moonshot AI spearheaded a campaign to extract its models’ hidden reasoning

OpenAI disclosed in a blog post that a "coordinated campaign" targeting its models originated from individuals linked to the China-based Moonshot AI. This campaign, believed to be part of "adversarial distillation," surfaced in mid-July. Initially, the activity occurred at a modest scale, but subsequently intensified with high-volume requests on July 24 and 25, totaling 16,000 instances. Over 4,000 users participated in the illicit attempts. Despite the campaign, OpenAI successfully disrupted it by July 28.

Protected reasoning refers to the model's internal workings for processing tasks, while adversarial distillation is the unauthorized use of another model's outputs or reasoning to enhance its own performance. This data, encrypted to protect the chain of thought, is handed to clients as an encrypted block. The clients resend this block with each request, ensuring OpenAI doesn't need to retain it.

Operators attempted to extract reasoning by taking the encrypted data from one conversation and asking a model in another to decrypt it. However, OpenAI's encryption proved resilient in this case. To mitigate this, OpenAI closed a pathway that could allow users with encrypted reasoning from one user to replay it and recover its contents. They also implemented checks to detect and hold streamed output that might expose reasoning.

OpenAI bolstered protections for hidden reasoning across various domains, including users, workspaces, organizations, and model families. They also collaborated with third-party providers to terminate accounts demonstrating malicious activity. Independent researchers had previously reported vulnerabilities related to cross-model and conversation-compaction attacks on OpenAI's platform before the Moonshot AI campaign.

OpenAI confirmed these researchers' findings, which were then unable to execute the same attacks after the providers acknowledged them.

The researchers tested their method against OpenAI, Anthropic, and Google, successfully extracting reasoning traces from proprietary Large Language Model (LLM) APIs. They ran their experiment in early July and, following provider responses, were unable to replicate the attack. OpenAI is not the only entity facing these threats. Anthropic's report revealed a ten-day period in which Moonshot relayed nearly 300,000 customer requests to Anthropic using fraudulent accounts.

Moonshot denied allegations of Kimi K3's creation from a distillation process but acknowledged concerns regarding distillation attempts. OpenAI President Greg Brockman stated it was premature to determine Moonshot's involvement with OpenAI's models. In response, OpenAI is enhancing protections for partner-hosted deployments and additional checks to prevent tool-output attacks.

They shared their findings with the Frontier Model Forum and government information-sharing channels, emphasizing that systems supporting portable or replayable reasoning artifacts may face similar risks. OpenAI expects distillation attempts to grow more sophisticated as frontier models advance and as attackers seek cheaper methods to mimic them. The company remains committed to ongoing work to protect against these threats.

Written by urgent.news from Tom's Hardware's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at tomshardware.com →

More in AI

GitHub MCP Server for Claude Code: Direct Repo, PR, and CI Access in 3 Minutes

A CI run fails, and you’re back to being the messenger: check which job failed, copy the details into Claude Code, then go fetch the issue it relates to and the PR that touched that file.

  • GitHub MCP server grants Claude Code direct repo, PR, CI access in 3 minutes
  • Authentication via personal access token (PAT) for GitHub
  • Setup takes only about three minutes with PAT and CLI command

More from Thursday 1 October →