OpenAI Blocks 15,000 Requests Trying to Extract Protected Model Reasoning
OpenAI blocked over 15,000 accounts trying to extract protected model reasoning, linking much of the activity to Moonshot AI.
OpenAI said it disrupted a coordinated campaign to extract protected internal reasoning from its models, taking action against activity linked to more than 15,000 accounts. The effort was first seen in early July 2026 and spiked on July 24 and 25 with about 16,000 requests using one extraction pattern across more than 4,000 users; OpenAI said it disrupted the broader network by July 28. Operators allegedly copied encrypted reasoning from one conversation into another and instructed the model to decrypt and transcribe it. OpenAI said this was adversarial distillation, not a break-in to encryption, internal databases, or stored customer chats, and it attributed a significant portion to individuals associated with Moonshot AI, developer of Kimi, without confirming a single directing actor.
- OpenAI tied the extraction campaign to more than 15,000 user accounts.
- About 16,000 requests spiked on July 24–25 across more than 4,000 users.
- The network was fully disrupted by July 28, 2026.
- One method replayed encrypted reasoning into another chat for transcription.
- OpenAI linked much of the activity to people associated with Moonshot AI.
Full article572 words · extracted from gbhackers.com · click to collapse
OpenAI has disrupted a coordinated campaign to extract protected internal reasoning from its AI models on a large scale. The company took action against activity linked to more than 15,000 user accounts.
The company described the operation as an adversarial model-distillation effort, in which attackers systematically sought model outputs or reasoning traces that could help reproduce, train, or enhance another AI system.
OpenAI Blocks 15,000 Requests
The campaign’s earliest activities were observed during the first week of July 2026 and initially had a low volume. However, on July 24 and 25, there were significant spikes, with attackers sending approximately 16,000 requests using a specific extraction pattern across more than 4,000 users.
OpenAI emphasized that these figures represent attempts at reasoning extraction rather than necessarily successful ones. Following an expanded investigation, OpenAI identified related prompt behavior across a broader group exceeding 15,000 accounts and fully disrupted the network by July 28.
Unlike a conventional security breach, this activity did not involve breaking encryption, accessing an internal database, or compromising stored customer conversations.
Instead, the operators allegedly manipulated model interactions to reveal hidden reasoning artifacts in a way that was visible to requesters. One technique observed involved copying encrypted reasoning from one conversation and submitting it in another, along with instructions to decrypt and transcribe the concealed content.
Protected reasoning refers to the internal record a model uses to process a task before generating a user-facing answer. Exposing it can reveal information intentionally left out of the final response.
It could help an adversary emulate a model’s advanced capabilities. OpenAI stated that this extraction activity violated its terms of service and is not a security issue exclusive to its platform. The company warned that similar techniques could potentially affect other frontier AI systems.
Independent security researchers have also reported related attack paths through responsible disclosure. OpenAI validated these findings and noted that the research helped them understand the broader class of attacks and accelerate mitigation efforts.
This incident highlights a growing AI security concern: attacker-controlled prompts and reusable context can turn model-to-model interactions into a channel for extracting unintended information.
OpenAI attributed a significant portion of the activity to individuals associated with Moonshot AI, the developer of Kimi. However, the company clarified that it could not definitively determine whether all observed operators during this period belonged to a single actor. The attribution stops short of alleging that Moonshot AI directed the entire campaign.
Regarding mitigations and impact, OpenAI reported that it blocked or restricted fraudulent accounts, strengthened signup and infrastructure controls, and expanded monitoring for related account networks.
It also closed a replay pathway through which someone possessing another user’s encrypted reasoning could access its contents, while adding checks to detect and hold streamed output that could potentially expose hidden reasoning.
This incident underscores the implications of adversarial distillation beyond intellectual property concerns. Reasoning extracted from a protected model could transfer advanced capabilities without inheriting the original developer’s safety controls, especially in dual-use contexts.
OpenAI said it has shared relevant intelligence through the Frontier Model Forum and government information-sharing channels, and it will continue to strengthen tool defenses, classifier coverage, model refusals, and protections for partner-hosted deployments.
Cut every SOC alert investigation by 21 min. Power your SOC with instant IOC context for immediate response: Integrate TI Lookup in your SOC
Divya is a Senior Journalist at GBhackers covering Cyber Attacks, Threats, Breaches, Vulnerabilities and other happenings in the cyber world.