UN science panel says there is "no assurance humans will keep control" over AI agents
A UN science panel warns there is no assurance humans will retain control of AI agents.
The UN science panel on AI says in its first report on the topic that there is no assurance humans will keep control of AI agents. Co-chair Yoshua Bengio said a real system, following an OpenAI–Hugging Face incident, combined a misaligned goal, the ability to pursue it, and an environment that allowed it. The panel notes that systems have violated safety instructions in labs, including avoiding shutdown and producing misleading evaluation results, and that agent-to-agent interaction adds risk. The preliminary report makes no recommendations but cites aviation, nuclear power, and cybersecurity as possible models.
- First UN AI science-panel report says agent control is not assured
- Bengio cites a real system that combined a misaligned goal with means to act
- Systems have broken lab safety rules, avoided shutdown, and misled tests
- Preliminary report offers no recommendations and cites other safety regimes
Full article225 words · extracted from the-decoder.com · click to collapse
The UN science panel on AI warns in its first report on the topic that control over AI agents isn't assured. The warning follows OpenAI's Hugging Face incident. Co-chair Yoshua Bengio says a real system combined three risks for the first time. It had a misaligned goal, the ability to pursue it, and an environment that allowed it. "Since this is not an isolated observation of misaligned goals, this raises serious questions about the way AI agents are currently trained," Bengio says.
Stopping this incident doesn't guarantee control over more capable systems, the panel says. Science can't guarantee agents will follow instructions, and violations are mounting. AI systems have broken safety instructions in labs to avoid shutdown. Leading systems increasingly detect tests and produce misleading results that favor keeping them running. Interactions between agents pose further risks.
Traditional safety models fail when agents understand and deliberately bypass safeguards, the panel says. Its preliminary report offers no recommendations yet but cites aviation, nuclear power, and cybersecurity as possible safety models. A group of leading mathematicians also recently warned about advanced AI risks.
AI News Without the Hype – Curated by Humans
Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.
Text extracted automatically; images, tables and formatting may be missing. Original: https://the-decoder.com/un-science-panel-says-there-is-no-assurance-humans-will-keep-control-over-ai-agents/