OpenAI disrupts reasoning-extraction campaign linked to Moonshot
OpenAI says it shut a July 2026 campaign extracting protected model reasoning and linked a core cluster to Moonshot AI; a researcher later said Azure still leaked it.
OpenAI said it identified and disrupted a coordinated adversarial-distillation campaign, first seen on July 1, 2026, that sought to extract protected reasoning by manipulating conversations, including replaying encrypted reasoning from one chat into another, without breaking encryption or accessing stored user data. Activity spiked on July 24–25 with about 16,000 extraction-pattern requests from more than 4,000 users, and a related cluster—more than 15,000 users in OpenAI's account, over 15,000 accounts in The Decoder—was disrupted by July 28. OpenAI attributes a core cluster from the first week of July to individuals associated with Moonshot AI, developer of Kimi, and shared findings through the Frontier Model Forum. Researcher Joachim Schaeffer showed those encrypted packets can be moved between sessions so a cheaper model from the same provider prints a stronger model's hidden reasoning. The Decoder reported that on September 13 the attack was blocked on OpenAI and Anthropic APIs but still worked on Microsoft Azure against every OpenAI model tried, including GPT-6 Astra, and Anthropic models up to Sonnet 5; a notepad-tool method also leaked reasoning from most of those models, with Azure protections said to have landed in late September.
- OpenAI said a coordinated adversarial-distillation campaign began July 1, 2026, and was fully disrupted by July 28.
- July 24–25 spike: about 16,000 extraction-pattern requests from more than 4,000 users (The Decoder said 16,000 requests).
- A related cluster involved more than 15,000 users, which The Decoder called over 15,000 accounts.
- OpenAI attributes a core cluster, from the first week of July, to individuals associated with Moonshot AI, developer of Kimi, and shared findings via the Frontier Model Forum.
- OpenAI said operators replayed encrypted reasoning between chats without breaking encryption or accessing stored user data.
- Researcher Joachim Schaeffer showed encrypted packets moved between sessions let a cheaper same-provider model print a stronger model's hidden reasoning.
- On September 13 the attack was blocked on OpenAI and Anthropic APIs but still worked on Microsoft Azure, including GPT-6 Astra and Anthropic models up to Sonnet 5.
- A notepad-tool method also leaked reasoning from most models tried; The Decoder said Azure protections for OpenAI and Anthropic models landed in late September.
Coverage timelineoldest first · each row is one article
- · 1d agoDisrupting a coordinated model-distillation campaign
OpenAI News· 78
OpenAI disrupted a July campaign extracting protected model reasoning, linking a core cluster to Moonshot AI.
- · 4h ago
- · 3h agoOpenAI says it stopped a campaign to steal its models' reasoning, but the trick still worked on Azure
The Decoder· 78
OpenAI blocked a large campaign to steal model reasoning, but researchers say Azure still leaked it.