OpenAI Says Attackers Got Its Model to Decrypt Its Own Protected Reasoning
OpenAI disrupted a distillation campaign that used a novel trick: asking the model to decrypt reasoning copied from another conversation. It points at Moonshot AI.

OpenAI says it shut down a coordinated campaign to copy the reasoning abilities of its models, and that people working for Chinese AI company Moonshot AI were behind the core of it. OpenAI calls the method novel, and says other AI models have the same weakness.
What happened
According to OpenAI's blog post, reported by CyberScoop:
- Low-level activity started on July 1 and built up to July 24–25, when OpenAI saw 16,000 prompts from 4,000 users following the same extraction pattern.
- By July 28 the number of suspicious users had reached 15,000, and OpenAI says it fully disrupted the operation that day.
- The attackers copied encrypted reasoning data out of one conversation, then asked the model in a separate conversation to decrypt it and write it out in plain text.
OpenAI stressed that nobody broke its encryption, breached a database or reached stored user conversations. The model itself was talked into reproducing protected reasoning in a form the requester could read. Outside researchers reported a similar flaw to OpenAI in August.
Attribution, and its limits
OpenAI named Moonshot AI, maker of the Kimi models, as responsible for a "core cluster" of the activity, though it said it is not clear all the activity is linked. Its post gives no technical evidence for the attribution, and OpenAI told CyberScoop it would not share more "for security reasons."
Distillation accusations against Chinese labs are not new. US companies and the US government have made them before, and security researchers say these operations rely on accounts bought in bulk on grey markets, used to send millions of prompts that help copy a model's capabilities.
OpenAI says it banned the accounts, tightened signup and infrastructure controls, expanded network monitoring and fixed the bug that let encrypted data from one conversation be decrypted in another. It has shared details with the Frontier Model Forum.
My read
The attribution will get the headlines. What stood out to me is that the model ended up decrypting its own protected output on request. The encryption held. The design around it let the model undo it.
Enterprises building on LLMs have the same exposure. If your agent stores anything "protected" that the same model can later be asked to read back, test that path on purpose. Access control belongs outside the model, in code, rather than in an instruction you hope it follows.
Source: CyberScoop — OpenAI reveals 'novel' encryption bypass used in distillation attack