Skip to content

OpenAI links reasoning-extraction campaign to Moonshot AI

OpenAI says a campaign it ties to Moonshot AI, maker of Kimi, sent 16,000 requests from 4,000+ users in July to copy its models' hidden reasoning.

By Tech AI Wire Team

4 min read

XLinkedIn
Two laptops on a dark desk at night, one showing a grid of chat windows and the other OpenAI's knot logo, linked by network cables to a switch.

By the numbers

extraction requests in the July 24-25 spike
16,000
users behind that spike
4,000+
users in the wider related cluster
15,000+
date OpenAI says it fully shut the campaign down
July 28

OpenAI said on September 30, 2026 that it had stopped a coordinated campaign to copy the hidden reasoning of its AI models. It blames a core group of the activity on people linked to Moonshot AI, the Chinese company behind the Kimi models. The attackers used a method OpenAI calls "novel": they had OpenAI's own model unlock its protected reasoning for them.

The stakes are about copying. Reasoning is the step-by-step thinking a model does before it answers, and it is valuable training data. A rival that collects enough of it can teach its own model to think the same way, without doing the same expensive work.

What OpenAI says happened

The campaign ran for four weeks in July, according to CyberScoop's report on OpenAI's disclosure.

DateWhat happened
July 1OpenAI first detects the activity
July 24-25A spike of 16,000 requests from more than 4,000 users
July 28OpenAI says the campaign is fully shut down

Looking wider, OpenAI found related prompt patterns across a cluster of more than 15,000 users, CyberScoop and AI Weekly report.

OpenAI calls the activity adversarial distillation. Distillation means training one model on the outputs of another. AI Weekly quotes OpenAI's definition: "the systematic and unauthorized use of one model's outputs or reasoning to help train, reproduce, or improve another model."

How the extraction worked

OpenAI hides its models' full reasoning from users. The reasoning is passed around in encrypted form, which a user cannot read.

The attackers found a way around that without breaking the encryption. They copied the encrypted reasoning out of one conversation. Then, in a separate conversation, they asked a model to decrypt it and write it out as plain text, CyberScoop reports. In other words, the model itself did the unlocking.

OpenAI stresses what the attackers did not do. "The operators did not break our encryption, compromise a database, or gain direct access to stored user conversations," the company said, according to CyberScoop.

The attribution, and its limits

OpenAI ties "a core cluster of the activity to individuals associated with Moonshot AI," SL Guardian reports. It could not tell whether all the operators worked for one group.

CyberScoop notes that OpenAI gave no technical evidence or reasoning for naming Moonshot. Moonshot did not immediately respond to a request for comment from CNBC, according to SL Guardian. OpenAI says it shared its findings with the Frontier Model Forum, an industry safety group, and with government contacts.

Moonshot has faced this kind of claim before. Anthropic has accused Chinese AI developers, including Moonshot and Alibaba, of using its Claude models to train competing systems, SL Guardian reports. CyberScoop adds that U.S. officials and American AI companies have accused Chinese firms of "systematic" distillation using thousands of purchased accounts.

OpenAI says it is not targeting open models. "Our concern is about violation of our terms of service, not open models or legitimate distillation," Caroline Zier of OpenAI said, according to AI Weekly.

What OpenAI changed

OpenAI fixed the flaw that let encrypted data move from one conversation to another, CyberScoop reports. It also tightened sign-up controls, hardened its infrastructure and added network monitoring. AI Weekly adds that OpenAI banned the accounts involved and added new protections for hidden reasoning.

What this means for developers

Expect stricter checks on OpenAI accounts. The attack relied on thousands of users, so tighter sign-up verification is the obvious defense. If your team runs many test accounts, expect more verification steps.

Read the terms before you train on model outputs. OpenAI says legitimate distillation is not its target. But using its outputs to build a competing model can still break its terms of service. If your team fine-tunes a smaller model on answers from a larger one, check which provider's terms apply.

Test your own model for the same trick. If you run a model service that hands users any protected or encrypted data, assume they can feed it back in. Check whether your model will decode, translate or summarize that data on request. This attack needed no stolen keys, only a model willing to help.

Watch for patterns across accounts, not just volume per account. The July spike came from more than 4,000 users sending similar requests over two days. Rate limits per account would not catch that. Look for the same prompt pattern spread across many new accounts.

Treat the attribution as a claim for now. OpenAI has named Moonshot but has not published its evidence. If your company uses Kimi models, follow how Moonshot and U.S. officials respond before changing plans.

Sources

  1. OpenAI reveals 'novel' encryption bypass used in distillation attack - CyberScoop
  2. OpenAI Disrupts AI Model Distillation Campaign Linked to Chinese Startup - SL Guardian
  3. OpenAI Links Moonshot Users to 16,000-Request Distillation Push - AI Weekly

Related articles

The daily brief

Three to five stories a day, and what each one means for the people who build software. Free, no spam.