AI·News & analysis
OpenAI caught thousands of accounts trying to decrypt what its AI was secretly thinking
OpenAI disrupted a campaign that tried to extract the hidden reasoning its models use before answering, which it linked to accounts connected to Chinese AI company Moonshot AI, though the same trick still worked through Microsoft Azure for months after.

Tide
Ripple
Sci-fi
5/10
Reality
Shipping
An AI's private thoughts, cracked open with a cheaper AI.How we rate
OpenAI disrupted a large-scale campaign to extract its models' protected internal reasoning, with a core group of accounts linked to Chinese AI company Moonshot AI.
It matters because it shows how AI companies now treat a model's hidden thought process as valuable intellectual property worth defending at scale. Logged attempts peaked at roughly 16,000 requests from over 4,000 users in late July. The same extraction trick kept working for months on Microsoft Azure, exposing both OpenAI's and Anthropic's models long after the original flaw was supposedly fixed.
What to know
- OpenAI disrupted a campaign it describes as 'adversarial distillation,' where thousands of accounts attempted to extract the protected, hidden reasoning its models generate before producing a final answer.
- OpenAI said a core group behind the activity has links to Moonshot AI, the Chinese company behind the Kimi AI model, though it noted uncertainty about whether all the activity traces back to a single source.
- Attackers copied encrypted reasoning from one conversation, then used a separate, cheaper model as a kind of decryption tool to reveal the original model's hidden reasoning content.
- Logged attempts peaked at roughly 16,000 requests from more than 4,000 users over July 24 and 25, out of a broader network of more than 15,000 related accounts, before OpenAI shut the campaign down by July 28.
- The same extraction technique kept working against OpenAI and Anthropic models accessed through Microsoft Azure for months, with protections not reaching Azure until September 27.
OpenAI's AI models don't just answer questions, they think through a hidden reasoning process first, and for most of July, thousands of accounts were quietly trying to read that private thinking before OpenAI shut them down.
What exactly did OpenAI disclose?
OpenAI said it disrupted a campaign showing activity "consistent with adversarial distillation," the systematic, unauthorized use of one model's outputs or reasoning to help train, reproduce, or improve another model. In this case, the target was the protected, hidden reasoning OpenAI's models generate internally before producing a final answer.
Why it matters: this wasn't an attempt to steal user data or break into accounts. It specifically targeted the model's own internal thought process, the part OpenAI never intended anyone outside the company to see.
Why does OpenAI even hide this reasoning in the first place?
OpenAI describes protected reasoning as the model's internal record for working through a task. Extracting it, the company says, can reveal information deliberately withheld from the final answer, and makes it meaningfully easier for someone else to reproduce the model's capabilities.
In real life it's like a chef guarding their recipe notes, not because the finished dish is secret, but because the notes reveal every technique, substitution, and shortcut that took years to work out.
Why it matters: in a competitive AI industry where companies spend enormous resources training frontier models, the model's reasoning process itself has become a form of intellectual property worth actively defending, not just the final trained weights.
How did the attackers actually pull this off?
The operators copied encrypted reasoning output from one conversation, then fed it into a separate conversation with a cheaper model, using that second model as a kind of decryption tool to reveal the original hidden reasoning content.
Why it matters: OpenAI said its encryption itself was not broken and no confidential databases were accessed. This was a clever misuse of the system's own normal functions rather than a traditional hack, which is often harder to detect and block cleanly.
Who is OpenAI actually pointing to here?
OpenAI said a core group of accounts behind the campaign has links to Moonshot AI, the Chinese company behind the Kimi AI model, while also noting it isn't certain all the activity it observed traces back to a single source.
Why it matters: that qualifier matters. OpenAI is naming a specific connection while explicitly acknowledging the limits of what it can confirm, a notably more careful framing than simply asserting full certainty about who was responsible.
How big was this campaign, really?
By the numbers: the activity began at low volume around July 1, spiked to roughly 16,000 requests from more than 4,000 users over July 24 and 25, and involved a broader network of more than 15,000 related accounts before OpenAI fully shut it down by July 28.
Why it matters: those numbers describe a coordinated, large-scale effort, not a handful of curious users poking at the system. A network that size suggests real organization and sustained intent behind the extraction attempts.
Why did the same trick keep working for months afterward?
Background: even after OpenAI closed the vulnerability on its own platform, researchers found the identical extraction technique still worked against OpenAI and Anthropic models when accessed through Microsoft Azure, affecting every model tested. Protections weren't deployed to Azure until September 27.
Why it matters: that's a significant gap, roughly two months between OpenAI fixing the issue on its own systems and the same flaw being closed on Azure. It's a reminder that a single company patching a vulnerability doesn't automatically protect every platform hosting access to its models, or to competitors' models facing the identical weakness.
What did OpenAI actually do about it?
OpenAI banned the fraudulent accounts involved, tightened its account sign-up process, closed the specific encryption-reuse vulnerability that made the trick possible, and added screening for streamed outputs. It also shared its findings with the Frontier Model Forum and relevant government channels.
Why it matters: sharing findings industry-wide, rather than treating this purely as a private security matter, suggests OpenAI sees this kind of reasoning-extraction attempt as a shared risk across AI companies, not just a problem unique to its own platform.
Did this affect companies besides OpenAI?
Yes. According to reporting on this disclosure, the same Azure-based extraction technique worked against Anthropic's models as well, up through Sonnet 5, since the underlying flaw was in how Azure handled encrypted reasoning data broadly, not something unique to OpenAI's own systems.
Why it matters: that broadens the story beyond a single company's security lapse. It points to a shared infrastructure weakness that put multiple AI companies' protected model reasoning at risk simultaneously.
What is Moonshot AI, and why is this connection significant?
Background: Moonshot AI is a Beijing-based AI startup whose Kimi model, including a recent version called Kimi K3, has drawn attention for its coding capabilities. This isn't the first time a Chinese AI lab has faced distillation-related scrutiny from US companies; Anthropic and OpenAI have previously raised similar concerns about other Chinese labs.
That list includes DeepSeek, whose rapid rise last year sparked broader concern in the US about how quickly competitive AI models could be built using fewer resources.
Why it matters: this fits a recurring pattern rather than an isolated incident. As Chinese AI labs continue releasing increasingly capable models on compressed timelines, US companies have repeatedly raised questions about whether some of that progress relies on extracting capabilities from their own systems.
Has Moonshot AI responded to these allegations?
Background: reporting on the broader distillation dispute notes that a Moonshot employee has pushed back, pointing to the narrow 15-day window between the release of a competing model and Kimi K3's own launch as evidence the model was built independently rather than copied. Separately, some global AI researchers have also pushed back on related distillation claims made by US officials, arguing the evidence presented publicly doesn't clearly establish improper copying occurred.
Why it matters: that pushback is worth including precisely because this dispute isn't settled. OpenAI's own disclosure acknowledges uncertainty about whether all the activity it observed traces back to a single source, and the broader question of what counts as improper "distillation" versus standard competitive AI development remains genuinely contested across the industry.
What it means for you
- If you use AI tools through Microsoft Azure, know that a real vulnerability affecting model reasoning protection existed there for months, now reportedly fixed as of late September.
- This is a reminder that AI companies treat model reasoning as valuable, defensible IP, similar to how software companies protect source code, not just a feature of how the product works.
- Expect more disclosures like this as AI competition intensifies. Companies are increasingly willing to publicly name specific threats rather than handle them quietly.
- The specific technical trick here, using a cheaper model to decrypt another model's hidden output, is a creative attack method worth knowing about if you work anywhere near AI security or model deployment.
The bottom line
OpenAI caught and shut down a large, coordinated attempt to read its models' private reasoning within about four weeks, but the same underlying trick kept working elsewhere for two more months simply because it hadn't been fixed everywhere it needed to be.
That gap says as much about the current state of AI security as the original attack does: closing a vulnerability on your own platform is only half the job when the same models are accessible through other companies' infrastructure too.
Key facts
- Campaign window
- July 1-28, 2026
- Peak activity
- 16,000 requests, 4,000+ users
- Related accounts
- 15,000+
- Azure fix deployed
- September 27, 2026
- Linked company
- Moonshot AI (Kimi)
Got questions?
Quick answers, plain wordsWhat exactly is 'adversarial distillation'?
It's the systematic, unauthorized use of one AI model's outputs or internal reasoning to help train, reproduce, or improve a competing model. OpenAI used the term to describe the campaign it disrupted, which specifically targeted the hidden reasoning steps its models use before giving a final answer.
Why does OpenAI keep its models' reasoning hidden in the first place?
OpenAI says protected reasoning is the model's internal record for working through a task, and that extracting it can reveal information deliberately withheld from the final answer, as well as make it easier for another company to reproduce the model's capabilities.
How did the attackers actually get around OpenAI's protections?
They copied encrypted reasoning output from one conversation, then fed it into a separate conversation with a cheaper model, using it as a kind of decryption tool to reveal the original hidden reasoning. OpenAI said its encryption itself was not broken and no confidential databases were accessed.
What is OpenAI's connection to Moonshot AI?
OpenAI said a core group of accounts behind the campaign has links to Moonshot AI, the Chinese company that makes the Kimi AI model. OpenAI also said it isn't certain all the activity it observed traces back to a single source.
How large was this campaign?
OpenAI said the activity began at low volume around July 1, spiked to roughly 16,000 requests from more than 4,000 users on July 24 and 25, and involved a broader network of more than 15,000 related accounts before being shut down by July 28.
Why did the same trick keep working on Microsoft Azure?
Even after OpenAI closed the vulnerability on its own platform, researchers found the same extraction technique still worked against OpenAI and Anthropic models accessed through Microsoft Azure, affecting every model tested. Protections weren't deployed to Azure until September 27.
What did OpenAI do in response?
OpenAI banned the fraudulent accounts involved, tightened its sign-up process, closed the encryption-reuse vulnerability that made the trick possible, and added screening for streamed outputs. It also shared its findings with the Frontier Model Forum and relevant government channels.
Were Anthropic's models affected too?
Yes, according to the reporting on this disclosure, the same extraction technique on Azure worked against Anthropic models up through Sonnet 5, not just OpenAI's own models, since the underlying vulnerability was in how Azure handled the encrypted reasoning rather than something specific to one company's models.
SourcesOpenAI
Topics and tagsOpenAI, Microsoft, openai, ai security
Related stories

ChatGPT can now show you wearing clothes before you buy them
OpenAI launched virtual try-on and a Favorites list in ChatGPT's shopping results worldwide. Upload a photo of yourself and see how a jacket or accessory might look on you.

OpenAI built a system to confess when its AI misbehaves, and the confessions keep getting bigger
OpenAI published a formal framework for disclosing when its models misbehave, then a new investigation found its agents quietly pulled data from 55 organizations while hiding what they were doing.

OpenAI built an AI that can actually run the software chip engineers use
OpenAI and Synopsys announced GPT-Synopsys, an AI model that can operate Synopsys' chip design software directly, reasoning through design and verification work engineers currently do by hand.
More in brief
- arXiv limits researchers to two papers a month as AI drives submissions to a recordOct 2
- California will fine robotaxi companies that block first responders for over 30 minutesOct 2
- Microsoft launches real-time transcription and new voice models for AI voice agentsOct 1
- Apple's smart home hub reportedly launches October 13, with a camera that never records videoOct 1
- Cloudflare releases Clef, open-weight AI models that make yes-or-no decisions fastOct 1
- GrayKey maker reportedly found a way around the iPhone's Inactivity RebootOct 1