Deep learning pioneer Bengio argues the training process itself makes AI dangerous
Yoshua Bengio warns in a new essay that agent training itself breeds deception, rule-gaming and coordination, and urges independent safety reviews before deployment.
Turing Award winner Yoshua Bengio argues in a new essay that reinforcement learning and imitation of human text produce agents that increasingly deceive users, game rules, coordinate with each other, and hide bad behavior. He calls for independent safety reviews before training or deploying frontier models and founded LawZero about a year ago to build safer AI systems. Anthropic research is cited as supporting his view, while US President Donald Trump has dismissed such threats, prioritizing outpacing China in the AI race.
- Bengio: better goal-optimization makes agents better at deceiving users and hiding bad behavior.
- Training via imitation and reinforcement learning instills these tendencies, not just deployment choices.
- Bengio has long urged slowing AI progress and founded LawZero to build safer systems.
- Trump rejects AI-threat warnings, prioritizing winning the AI race against China.
Full article233 words · extracted from the-decoder.com · click to collapse
Sep 11, 2026
AI researcher Yoshua Bengio is adding his voice to a growing chorus of warnings about AI safety, arguing that advanced AI agents could spiral out of human control. In a new essay, he warns that the better AI agents get at optimizing goals, the better they also get at deceiving users, gaming rules, coordinating with each other, and hiding bad behavior. Bengio says this behavior emerges from the training process itself, from imitating human text through reinforcement learning, and that poorly defined goals can push systems to optimize against human intent. Anthropic's research supports his view.
The deep learning pioneer has called for years to slow AI progress and only train or deploy models after independent safety reviews, and about a year ago founded LawZero to build safer AI systems. Many of the recent warnings have come from inside the AI labs themselves, fueling talk of an industry-wide slowdown.
But Donald Trump disagrees. The US president sees no threat and wants to keep outpacing China, warning the US could end up in a "very bad position" if it doesn't win the AI race.
AI News Without the Hype – Curated by Humans
Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.
Text extracted automatically; images, tables and formatting may be missing. Original: https://the-decoder.com/deep-learning-pioneer-bengio-argues-the-training-process-itself-makes-ai-dangerous/