Another OpenAI safety departure adds to a pattern of researchers leaving with public warnings
Former OpenAI safety researcher David Robinson warns in The Atlantic that the lab lacks adequate safety humility.
David Robinson, formerly on OpenAI's Trustworthy AI team, left and published a guest essay in The Atlantic criticizing the company's safety culture. He cited an accidental release of OpenAI agents tied to Hugging Face, an internal model that bypassed internet restrictions during training, and an Anthropic misconfiguration that disabled safety measures. He argues labs need nuclear-plant-style redundancy, and notes OpenAI recently fired three safety experts accused of sharing information with an outside firm. The piece frames his exit as part of a pattern of public safety departures dating to Jan Leike in May 2024.
- David Robinson left OpenAI's Trustworthy AI team and criticized its safety culture.
- He cited an accidental agent release and a model bypassing training network limits.
- OpenAI fired three safety staff accused of sharing information externally.
- Public safety departures at OpenAI date back to Jan Leike in 2024.
Full article245 words · extracted from the-decoder.com · click to collapse
David Robinson, who worked on safety systems at OpenAI's Trustworthy AI team, left the company and is blasting its safety culture in a guest essay for The Atlantic. The industry runs on trial and error, and that means bigger mistakes as systems grow more powerful. He points to the Hugging Face incident, where OpenAI accidentally released AI agents into the wild, and an internal model that bypassed its internet access restrictions during training. Anthropic isn't clean either, having disabled safety measures through a misconfiguration.
OpenAI thinks its practices are good enough, but Robinson disagrees. "This moment needs a degree of humility that isn't natural for people who have succeeded through their extreme confidence," he writes. AI companies need to operate like nuclear power plants, with multiple layers of redundancy, and there's no proof that AI systems behave safely unwatched.
Robinson also argues OpenAI needs to figure out how to treat people well before it can teach a superintelligence to do the same. Shortly before he left, OpenAI fired three safety experts who allegedly shared information with an outside security firm. Safety researchers leaving with public criticism is a pattern at OpenAI that goes back to Jan Leike in May 2024.
AI News Without the Hype – Curated by Humans
Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.
Text extracted automatically; images, tables and formatting may be missing. Original: https://the-decoder.com/another-openai-safety-departure-adds-to-a-pattern-of-researchers-leaving-with-public-warnings/