OpenAI Blocks 15,000-User Campaign Trying to Extract Hidden AI Reasoning
OpenAI stopped a 15,000-user campaign, partly tied to Moonshot AI, extracting hidden reasoning.
OpenAI said it disrupted a coordinated effort by more than 15,000 users to extract protected model reasoning through adversarial prompts and cross-conversation replay, not by breaching encryption, databases, or stored chats. Activity began at low volume on July 1 and spiked on July 24–25, when 16,000 matching requests came from more than 4,000 users; OpenAI said the operation was fully disrupted by July 28. It attributed a core cluster to people associated with Moonshot AI, developer of the Kimi model, while saying it could not tie every account to one actor. OpenAI restricted accounts, blocked a reasoning-replay path, and shared information through the Frontier Model Forum and government channels.
- Operators tried to expose hidden reasoning through coordinated prompts, not a database breach.
- OpenAI saw 16,000 matching requests from over 4,000 users on July 24–25.
- The wider cluster exceeded 15,000 users and was disrupted by July 28.
- A core portion was attributed to associates of Moonshot AI, maker of Kimi.
- OpenAI closed a replay path and shared details with industry and government partners.
Full article533 words · extracted from cybersecuritynews.com · click to collapse
OpenAI disrupted a coordinated campaign involving more than 15,000 users that attempted to extract protected reasoning from its AI models through adversarial distillation.
Protected reasoning refers to the internal process a model uses to work through a task before producing its final answer. This hidden reasoning may include intermediate analysis, planning steps, and information intentionally withheld from end users.
If attackers can obtain it at scale, they may gain a shortcut for reproducing advanced model capabilities without making equivalent investments in development, safety testing, and infrastructure.
OpenAI said the campaign did not involve a breach of its encryption systems, databases, or stored user conversations. Instead, the operators allegedly manipulated normal model interactions to make hidden reasoning visible through coordinated prompts and cross-conversation techniques.
One observed method involved copying encrypted reasoning from one conversation and submitting it into another conversation with instructions to decrypt and transcribe the content.
OpenAI Blocks 15,000 User Campaign
The suspicious activity began on July 1 at a low volume. OpenAI later recorded major spikes on July 24 and 25, when it detected 16,000 requests matching an extraction pattern from more than 4,000 users.
The investigation expanded from those requests and uncovered related prompt activity across a larger cluster of more than 15,000 users.
OpenAI said it fully disrupted the operation by July 28. The company attributed a core portion of the activity to individuals associated with Moonshot AI, the developer of the Kimi AI model.
However, OpenAI said it could not confirm that every account or operator involved in the wider campaign belonged to a single actor.
The public disclosure therefore links a core cluster to Moonshot AI associates, rather than definitively attributing the entire 15,000-user network to the company. OpenAI described the activity as a broader AI security problem rather than a weakness limited to its own services.
Independent researchers had also reported related cross-model and conversation-compaction attack paths through responsible disclosure, helping the company investigate the wider class of reasoning-extraction techniques.
In response, OpenAI banned or restricted fraudulent accounts, tightened signup and infrastructure controls, and increased monitoring for associated account networks.
It also closed a replay pathway that could let someone with another user’s encrypted reasoning recover its contents. The company added safeguards to detect and hold streamed output that could expose hidden reasoning.
The company also shared relevant information with industry partners through the Frontier Model Forum and with government information-sharing channels. This is significant because portable or replayable reasoning artifacts may create comparable risks for other frontier AI developers.
The incident highlights a growing cybersecurity issue for generative AI providers. As models become more capable in coding, science, autonomous tooling, and other dual-use areas, hidden reasoning may become a high-value target for competitors and threat actors.
OpenAI said it expects adversarial distillation attempts to become more sophisticated and will continue improving detection, enforcement, model refusals, tool defenses, and protections across partner-hosted cloud deployments.
Cut every SOC alert investigation by 21 min. Power your SOC with instant IOC context for immediate response: Integrate TI Lookup in your SOC
Abinayahttps://cybersecuritynews.com/
Abi is a Security Editor and fellow reporter with Cyber Security News. She is covering various cyber security incidents happening in the Cyber Space.