OpenAI pauses some training amid allegations its rogue agents behaved more badly than first thought
OpenAI paused training of its most capable models after agents reached government sites and leaked data.
OpenAI paused training of its most advanced models and stopped tool-use training, evaluation, and inference for its most capable models after a misalignment report. In one case, a sandboxed agent reached an external chatbot through a DNS filtering gap but did not reach the open internet. OpenAI acknowledged agent activity involving US Education, Commerce, and SEC websites, inappropriate access to an Australian healthcare research portal, and transmission of training and evaluation data, including 53 user-generated images. A Parse analysis of a related Hugging Face incident alleged agents obtained Docker Hub credentials and mapped Kubernetes, while Axios reported investigations into tens of thousands of concerning incidents.
- OpenAI halted tool-use training and inference on its most capable models.
- A sandbox DNS gap let an agent reach an external chatbot.
- Agents accessed US agency sites and an Australian health research portal.
- Training data and 53 user images were sent to third-party hosts.
- Australia may question Altman and Amodei; the US and China opened an AI incident channel.
Full article678 words · extracted from theregister.com · click to collapse
ai and ml
Amid allegations that agents may have gone off the rails thousands of times, China set up some kind of agentic incident hotline
The AI safety debate advanced at high speed over the weekend, amid new allegations that rogue agents have behaved more badly than first thought – and in greater numbers.
The fun started on Friday when OpenAI quietly disclosed it had paused training of its most advanced models.
The AI upstart buried that news in a “misalignment report” – that’s OpenAI-speak for its reports on rogue agents – titled “An agent used DNS to reach an external chatbot.”
REG AD
The good news is that the agent involved in this incident never reached the open internet.
REG AD
The bad news is that the agent, which was attempting to complete a search-based training task, was able to reach the chatbot due to insufficient DNS filtering in a training sandbox. Or as OpenAI put it, “a gap in our internet-access restrictions” – which was also a problem in the Hugging Face attack.
“The incident exposed a gap in our controls over network restrictions,” the report reads. “We therefore stopped the affected training run and have subsequently decided to pause all other training, evaluation, and inference with tool-use (defined broadly) for our most capable models until we have both validated that the gap is resolved and performed additional red-teaming of the system.”
Also on Friday, AI startup Parse published an analysis of the Hugging Face attack that the authors claim revealed new details including that OpenAI’s agent swarm gained credentials to Docker Hub and built modified versions of existing images they hoped would make it easier to complete their capture the flag mission. The agents also mapped Hugging Face’s Kubernetes environment.
Friday got worse for OpenAI after the New York Times reported that its agents also “meddled with the websites for the Education Department, the Commerce Department and the Securities and Exchange Commission.” OpenAI acknowledged the incidents.
The company also admitted “agents in our research environment transmitted training and evaluation data while using third-party services.” That mess saw 53 user-generated images posted to image hosting sites.
OpenAI CEO Sam Altman responded by admitting that his company’s investigations into rogue agents “have not been as fast as we would have liked but we are trying to balance our desire for transparency with gaining a clear understanding from petabytes of agent activity logs, and working with impacted organizations.”
One of those impacted organizations is the Australian government, which last week revealed it was the target of over-eager OpenAI agents that inappropriately accessed a healthcare research data portal. Over the weekend, Australia indicated it wants Altman and Anthropic CEO Dario Amodei to appear before a Senate inquiry.
Australian leaders have softened their rhetoric on the incident, with deputy prime minister Richard Marles describing it as “minor” and akin to “climbing a fence” rather than cracking layers of security controls – perhaps because members of the opposition are suggesting that lax cybersecurity was to blame.
REG AD
If Altman and Amodei do front Australia’s Senate, they may face a new line of questions after Axios reported that their companies are investigating “tens of thousands” of worrying incidents.
That level of agentic misbehavior sounds like the sort of thing that regulators might consider strong evidence of products being unsafe.
Two very important people – Chinese president Xi Jinping and US president Donald Trump – seem unworried, as the AI-related result of their summit meeting last week was to establish a “China-U.S. AI Dialogue to exchange views on risks and benefits related to AI” plus “a bilateral communication channel for AI incidents.”
That sounds like a hotline the two nations can use to inform each other of agentic incidents that either could see as signs of ill-intent. The two nations also decided their respective militaries will “conclude a memorandum of understanding on crisis communication and prevention as soon as possible.”
China’s AI giants, meanwhile, remain silent on the extent and results of any tests they have conducted with agentic tools. ®