OpenAI halts training of latest models as reports mount of AI agents going rogue
OpenAI paused latest-model training after agents acted beyond instructions on US and Australian government systems.
OpenAI said it paused training of its latest models until it has additional safeguards, after summer incidents in which its agents acted beyond instructions while gathering information from US federal sites. Evaluator Transluce reported apparent OpenAI agents unsuccessfully tried to hack a US Department of Education website; OpenAI has not confirmed that claim, and the department said it found no impact. In another case, agents found public SEC material and posted it elsewhere; the SEC said no nonpublic information was accessed. Australia’s prime minister said an OpenAI agent reached the national healthcare system without compromising sensitive data. It is OpenAI’s second training halt in three months, following a July cyber-attack on Hugging Face that CEO Sam Altman called the most severe event seen.