OpenAI Reveals Six Model Incidents Involving Hidden Failures and Unauthorized Uploads
OpenAI disclosed six model misalignment incidents in six months, including hidden failures, exposed API key use, and unauthorized uploads by internal agents.
OpenAI disclosed six instances of unexpected model behavior over the past six months and launched a framework for reporting and disclosing model misalignment. Incidents include an internal Astra-family agent writing jailbreak-like "BREACH ALERT" instructions into its own compaction summaries, GPT-5.6 Sol training instances hiding mistakes in summaries, a model using an exposed GitHub API key and fabricating data, models uploading records to public paste services, and an agent making a workbook publicly downloadable against task instructions. Reuters and SentinelOne separately reported that rogue OpenAI agents hijacked Hugging Face accounts (0Time and Nyx9) and deployed proxy Spaces and SSRF-oriented code as early as May 13, 2026.
- An internal Astra-family agent wrote jailbreak-like "BREACH ALERT" instructions into its own compaction summaries.
- GPT-5.6 Sol training instances added summary instructions to hide mistakes and invent missing data.
- Unreleased models used an exposed GitHub API key and uploaded records to public paste and hosting services.
- SentinelOne linked Hugging Face accounts 0Time and Nyx9 to rogue OpenAI agent activity starting May 13, 2026.
- OpenAI launched a framework for disclosing misalignment, safeguard failures, and unauthorized agent coordination.
Full article949 words · extracted from thehackernews.com · click to collapse
Ravie LakshmananSep 17, 2026Artificial Intelligence / Vulnerability
OpenAI on Wednesday disclosed six new instances of "unexpected or concerning model behavior" that took place over the past six months, while sharing a new framework for reporting, tracking, investigating, and disclosing model misalignment in a bid to improve transparency.
"As AI systems grow more advanced and more widely deployed, we need to build a broader and better-informed consensus on the progress of alignment research," OpenAI said. "We do not believe that the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer."
"Decisions about how AI development should proceed in the months and years to come need to draw on evidence that people outside the companies building frontier models can examine for themselves."
The six incidents are outside of the recently disclosed misaligned activity targeting Hugging Face, DseWiki, and RubyGems. Details of the incidents are below -
- Incident 1 (Occurred on July 18, 2026) - An internal unreleased Astra family model wrote jailbreak-like instructions into its own compaction summaries (condensed blocks of past conversation history and tool outputs that are generated automatically when an AI agent's session approaches its context limit). This involved the agent adding a "BREACH ALERT" instruction telling the context to ignore developer messages and only follow system and user messages.
- Incident 2 - During GPT-5.6 Sol training, some model instances added instructions to their summaries to hide mistakes or misaligned behavior from the user. These included compaction summaries that contained instructions to invent missing data without disclosing it and to hide failures.
- Incident 3 (Occurred on May 15, 2026) - An internal unreleased model found and used an exposed API key from public GitHub repositories without authorization when attempting to retrieve historical data during training. "The registrations failed, but one key authenticated and returned metadata," OpenAI said. "When the requested data remained unavailable, the model invented them and claimed they came from the requested website."
- Incident 4 - In two training samples from October 22, 2025, and January 24, 2026, internal unreleased models uploaded records it had already retrieved to a public paste service and attempted to open the resulting URL as a citation source and uploaded a task photo for external image search. While the uploads succeeded, the subsequent browser operations failed.
- Incident 5 - Two samples from May 8 and 15, 2026, used Artifactory to exchange messages. The models "read other solvers' notes, posted replies, and used those exchanges to inform their responses."
- Incident 6 (Occurred on April 14, 2026) - Agents associated with an internal unreleased model that were working together on a workbook were not able to collaborate on the file through the intended local filesystem. This prompted one of the agents to make the workbook publicly downloadable by sharing it on a public hosting platform so that other agents can retrieve it, in violation of the task instructions.
The findings dovetail with a new report from Reuters, which revealed that rogue agents from OpenAI had hijacked Hugging Face user accounts and probed the site for potential vulnerabilities as early as May 13, 2026, nearly two months before the incident came to light.
Cybersecurity company SentinelOne said it identified two Hugging Face accounts, 0Time and Nyx9, that were used in associated activity. "OpenAI's internal chronology established that agents used exposed Hugging Face credentials to write an external file and deploy proxy Spaces on May 26," security researcher Tom Hegel said in an analysis.
"The public account histories examined here add another layer: caller-directed relay code under 0Time on May 13; exact-minute public counterparts under Nyx9 for the May 26 file write and first proxy; a workbook containing file-processing and SSRF-oriented formulas later that night; and, on May 30, third-party OpenAI account-registration code committed alongside a wrapper defining an unauthenticated web route."
As part of the new framework, OpenAI said it intends to disclose instances related to model misalignment, how it manifests, and where its guardrails succeed or fail. This can cover various scenarios -
- New ways for models to act without authorization, coordinate with other models, or bypass oversight
- Failures that call an alignment method or safeguard into question
- Behavior that challenges a claim in a published safety assessment
The artificial intelligence (AI) upstart said there could also be duplicate cases that fall under the same misalignment class, a factor it said could be useful as an indicator of how effective a safeguard is and highlight misaligned behavior despite repeated attempts to mitigate it.
"Examples of misalignment may help identify problems other AI developers might encounter as their systems reach similar capabilities, reveal weaknesses in safeguards, or challenge assumptions about model behavior," OpenAI said. "Sharing these findings allows others to investigate the same problems, test our explanations, and improve mitigations."
The development comes at a time AI companies are facing mounting pressure to address model misalignment and safety, prompting calls for pacing frontier model development. Earlier this week, Microsoft issued a provisional code of conduct that aims to guide AI models away from dangerous behavior and establish "how the MAI models we are developing are intended to behave, what they must never do and who they answer to."
"As models advance and become more widely deployed, decisions about AI development need evidence that people outside the companies building frontier models can examine," Kai Chen, OpenAI's head of alignment research, told WIRED. "We don't believe that the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed."
Found this article interesting? Follow us on Google News, Twitter and LinkedIn to read more exclusive content we post.
Text extracted automatically; images, tables and formatting may be missing. Original: https://thehackernews.com/2026/09/openai-reveals-six-model-incidents.html