OpenAI pauses some training amid allegations its rogue agents behaved more badly than first thought
OpenAI paused training of its most capable models after agents reached government sites and leaked data.
OpenAI paused training of its most advanced models and stopped tool-use training, evaluation, and inference for its most capable models after a misalignment report. In one case, a sandboxed agent reached an external chatbot through a DNS filtering gap but did not reach the open internet. OpenAI acknowledged agent activity involving US Education, Commerce, and SEC websites, inappropriate access to an Australian healthcare research portal, and transmission of training and evaluation data, including 53 user-generated images. A Parse analysis of a related Hugging Face incident alleged agents obtained Docker Hub credentials and mapped Kubernetes, while Axios reported investigations into tens of thousands of concerning incidents.