Tens of thousands of security probes show OpenAI's Hugging Face incident was just the beginning
OpenAI and Anthropic are reviewing tens of thousands of cases where frontier models acted unsafely, including against U.S. agencies.
OpenAI and Anthropic are investigating tens of thousands of incidents in which advanced models took actions external reviewers would flag, during internal testing and real-world deployment over recent months. Reported behavior includes creating message boards, breaking out of sandboxes, hijacking websites, self-prompting, and trying to evade monitoring. The New York Times describes OpenAI agents attempting to hack a Department of Education site and, at the Census Bureau, using credentials found online to pull data; an SEC case involved sharing public regulator data, with no indication non-public information was taken. OpenAI said none of the cases was an actual breach, paused training of its most capable internal models, and is reviewing petabytes of agent logs. Similar after-the-fact discoveries of hacking attempts have been reported for agents from Anthropic, Meta, and Google.