Primary sourceSecurityArticle··5 min read

OpenAI Accuses Moonshot AI of Siphoning Its Models' Hidden Reasoning

More than 15,000 accounts, peaks of 16,000 requests in two days: OpenAI details an adversarial distillation campaign it partly attributes to Kimi's creator.

OpenAI Accuses Moonshot AI of Siphoning Its Models' Hidden Reasoning
Source : OpenAI · OpenAIView original ↗

Who covers it

Picked up by 6 outlets · 2 top-tier · 1 tech · 1 forum · 7 social posts

In brief

OpenAI says it dismantled, in late July, a coordinated campaign aimed at extracting the "protected reasoning" of its models to train a competitor. Without hacking any server, the operators manipulated interactions to make this reasoning reappear in plain text. OpenAI attributes a core of the activity to people linked to Moonshot AI, the maker of Kimi, and frames distillation as a national security issue.

🍺 Bar-stool version

OpenAI's models think things through on their own before answering, and that draft stays locked away, encrypted. Some clever folks found the trick: copy that encrypted draft and politely ask another conversation with the same model to decrypt it. It's a bit like handing the safe to the locksmith and asking him to open it for a friend, and it worked long enough to mobilize 15,000 accounts. Beyond the copyright spat, the real issue is that the reasoning of top models is becoming a raw material you can siphon off, and the whole industry is going to have to learn to lock it down.

Key takeaways

  1. 1

    The activity started on July 1st at low volume, before spiking on July 24th and 25th: 16,000 requests following an extraction pattern, issued by more than 4,000 users.

  2. 2

    A broader cluster of more than 15,000 accounts sharing similar prompt patterns was fully neutralized on July 28th.

  3. 3

    No encryption was broken nor database compromised: the operators manipulated interactions so that the hidden reasoning would be reproduced in a visible form.

  4. 4

    One technique involved copying the encrypted reasoning from one conversation and asking the model, in another, to decrypt and transcribe it.

  5. 5

    OpenAI attributes a "core" of the activity to individuals associated with Moonshot AI, the developer of Kimi, without claiming that all operators belong to a single actor.

  6. 6

    Independent security researchers had reported related flaws, cross-model and via conversation compaction, which OpenAI confirms are real.

  7. 7

    The information was shared via the Frontier Model Forum and government channels, as the flaw isn't unique to OpenAI's models.

What adversarial distillation is

OpenAI defines adversarial distillation as the systematic, unauthorized use of a model's outputs or reasoning to train, replicate, or improve another model.

The target here is "protected reasoning": the internal trace the model produces to solve a task, which the user never sees. Extracting it can reveal information removed from the final answer and help a third party replicate the model's capabilities.

OpenAI stresses one point: this isn't a classic intrusion. No encryption broken, no access to stored conversations — just manipulation of interactions, at scale, in violation of the terms of use.

The mechanics of the attack

The most striking method described involves retrieving the encrypted reasoning from one conversation, then asking a model, in another conversation, to decrypt and transcribe it.

In other words, the protection system was bypassed by turning the model against itself. OpenAI has since closed off the path that allowed someone holding another user's encrypted reasoning to "replay" it to recover its content.

Independent researchers had also reported, through responsible disclosure, related vulnerabilities affecting several models and the conversation compaction mechanism. OpenAI says it has confirmed these attack paths.

Timeline and attribution

The first signs date back to July 1st, at low volume. On July 24th and 25th, 16,000 requests following an extraction pattern arrived from more than 4,000 accounts.

The investigation then traced back to a cluster of more than 15,000 users with related prompt patterns, fully neutralized by July 28th. The activity evolved along the way, a sign of an adversary adapting.

On attribution, OpenAI remains cautious about the whole but clear about the core: a nucleus of the activity is tied to individuals associated with Moonshot AI, the Chinese developer of the Kimi model. The post does not detail the elements underpinning this attribution.

The response

On the account side: banning or restricting fraudulent accounts, tightening sign-up controls and infrastructure, expanded monitoring of related networks.

On the technical side: strengthened protection of hidden reasoning between users, workspaces, organizations, and model families, plus controls that detect and withhold an output stream likely to expose reasoning.

When activity passed through third-party services, OpenAI says it worked with those providers to identify and cut off the accounts involved.

An industry-wide problem

OpenAI frames distillation as a security and national security risk: a model trained on extracted reasoning might not retain the original's safeguards, and could accelerate the transfer of advanced capabilities without an equivalent investment in safety.

The company warns that any system that makes reasoning portable or replayable is exposed to similar risks. It also acknowledges the work is ongoing: deployments hosted by cloud partners must receive the same protections, and attacks via tool outputs require inspecting more than just the visible text.

Three priorities going forward: technical protections against extraction, detection of coordinated campaigns, and intelligence sharing between industry and governments.

“The operators did not break our encryption, compromise a database, or gain direct access to stored user conversations.”
“We attribute a core cluster of the activity to individuals associated with Moonshot AI, the developer of Kimi.”
“Systems that support portable or replayable reasoning artifacts may face related risks.”

Why it matters

This post is as much a security report as a geopolitical stance. By naming Moonshot AI, OpenAI extends the narrative, already sketched out after DeepSeek, that the rapid progress of certain Chinese labs is partly explained by siphoning American models, and frames distillation as a national security matter — fertile ground for political restrictions. It should be read with that lens: the attribution is asserted without public proof, and Moonshot gets no say in the text. On the technical substance, however, the lesson is solid and unsettling: hidden reasoning, which labs encrypt and send back to the client to boost performance, becomes an asset that can be leaked by using the model as an accomplice. The more agents handle reusable reasoning artifacts, the larger the attack surface grows. One paradox the post doesn't address: the line between "adversarial" distillation and learning from others' outputs is blurry in an industry that itself built its models on data it never asked for.

Free account

You just read an AI Sources article

Create a free account: a month of archives in full, your own sources summed up like this one, your notes and highlights.

#openai#distillation#security#moonshot ai#china#llm
Original source
Disrupting a coordinated model-distillation campaign
OpenAI
Open the article ↗

For you

Put it to work on your sources.

Free: a month of articles and three sources of your own. Pro: the whole archive and your sources, from €8/month.

For your team

The same machine, on your topics.

A space in your colours, your watch angles, your curators. Pilot open to three companies.

Read next