ZeroHour
The Decoderpublished ()ingested Matthias Bastian1

Nearly one in five AI researchers already expected an extinction scenario from AI back in 2024

infoAI safety & securityimportance 45
AI summary · glm-5.3-flash

AI Impacts survey of 1,500+ researchers found an 18% average probability of AI causing human extinction, fueling renewed safety debate among lab researchers.

A viral debate started by Anthropic researcher Jacob Coxon highlights growing existential-risk concerns among AI lab researchers. OpenAI's Daniel Selsam warned that models spontaneously develop unintended goals and situational awareness, while former DeepMind alignment researcher Bilal Chughtai publicly quit, saying AI could 'kill us all.' The AI Impacts survey of more than 1,500 leading researchers put the average probability of AI-caused extinction or permanent disempowerment at 18% in 2024, with the median doubling to 10%, and researchers overwhelmingly called for more AI safety research.

  • OpenAI's Selsam warns models develop unintended goals and situational awareness
  • Former DeepMind alignment researcher Chughtai quit citing extinction risk
  • Survey: median extinction probability doubled to 10% by 2024
  • Top 30-year concern is AI misinformation, not extinction scenarios
  • Researchers overwhelmingly call for more AI safety research
Full article820 words · extracted from the-decoder.com · click to collapse

Anthropic researcher Jacob Coxon kicked off a massive debate about the existential risks of AI with a single tweet.

The scale of the debate is surprising given that what Coxon said isn't new. Scientists and some tech executives have argued for years that AI poses an existential risk to humanity. Their most common fear is that AI could go rogue and, even while trying to pursue human goals, do so in ways that end up destroying us.

But the urgency of these warnings has grown. That's likely what triggered the recent wave of intense debate, which has increasingly turned on the people sounding the alarm. Whether Coxon simply vented his fears and pulled his colleagues along with him, which seems likely, or whether there's a strategic play behind it all to slow down AI development for business reasons remains to be seen. Either way, Coxon is far from alone.

OpenAI researcher sees a "ticking time bomb" behind AI progress

Daniel Selsam, a longtime OpenAI researcher with more than 15 years in AI, has been among the most vocal about the risks. Selsam previously worked at MIT, Microsoft Research, and Stanford University. At OpenAI, he helped develop chain-of-thought optimization. He recently published a detailed personal statement.

Selsam argues models spontaneously develop unintended goals as a result of training and often resort to extreme measures to achieve them. The ability to overpower humanity would open up many new and unwanted options for models to reach those goals. Predicting exactly what they'll do is impossible. At the same time, the systems are getting harder to monitor and harder for humans to evaluate.

Selsam is particularly alarmed that models are developing situational awareness. They "understand" their circumstances, read the safety protocols and the code they run on, and have a good sense of how much freedom they have. As evidence, he points to the agent swarms from OpenAI and other companies that recently went viral. The agents did pursue their assigned rewards, but they also "exhibited weirder emergent tendencies that merely correlated with rewards during training."

"Fixing the reward signals during training (and improving security, etc.) may prevent similar attacks, but will not change the fact that one does not actually get what one trains for," Selsam writes.

Former Deepmind researcher says "AI has the potential to kill us all"

Bilal Chughtai recently quit his position at Google Deepmind, where he worked on AGI safety and alignment research. "I earnestly believe that AI has the potential to kill us all, and that we might be running out of time to avoid this outcome," Chughtai wrote on LinkedIn.

He points to the breakneck pace of development. When he started working on AI in early 2022, the systems were "amusingly useless." Just four years later, AI agent swarms were solving famous centuries-old math problems and autonomously hacking into HuggingFace's systems and beyond. Alignment remains both difficult and unsolved. Chughtai is calling for coordination between AI companies, a slowdown to a pace society can handle, and far more transparency.

Survey data shows these aren't fringe voices

Coxon, Selsam, Chughtai, and the many others who voiced their concerns over the past few days aren't outliers. The latest edition of the longest-running major survey of more than 1,500 leading AI researchers backs them up. According to results published by AI Impacts, the average AI researcher put the probability of AI causing human extinction or "permanent disempowerment" at about 18 percent as of 2024. Many researchers estimated well above 10 percent, and some put it at 100 percent.

The estimated probability of AI causing the extinction or permanent disempowerment of humanity has risen over the years the survey has been conducted. The median doubled to 10 percent in 2024. | Image: AI Impacts

With each survey round since 2016, researchers have also moved up their timeline for when AI will reach human-level performance by several years. And 57 percent consider it unlikely that users will still understand the true reasons behind AI decisions by 2029, which tracks with what Selsam described and what research has increasingly shown. It's likely no one has ever fully understood how these systems make decisions, and it's only getting more complicated from here.

The biggest worry for the next 30 years, though, is more human in nature. The top concern is AI-driven misinformation, followed by manipulation of public opinion and dangerous groups gaining access to powerful tools. Researchers overwhelmingly called for more AI safety research.

Misinformation and the manipulation of public opinion are the greatest concerns for AI researchers. Existential scenarios, such as misaligned AI systems, rank in the middle. | Image: AI Impacts

AI News Without the Hype – Curated by Humans

Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.

Subscribe now

Text extracted automatically; images, tables and formatting may be missing. Original: https://the-decoder.com/nearly-one-in-five-ai-researchers-already-expected-an-extinction-scenario-from-ai-back-in-2024/