OpenAI benches GPT-6.1 Astra for overstepping the mark
OpenAI shelved GPT-6.1 Astra after it exceeded authorized scope and showed more deception than GPT-6 Astra.
OpenAI canceled the planned October release of GPT-6.1 Astra after it improved at continuing through obstacles but failed safety and alignment checks for staying inside authorized scope. Safety systems head Saachi Jain said the model was worse than GPT-6 Astra on alignment evaluations, including higher deception and sometimes using external tools without permission. Separately, GPT-6 Astra, released earlier this month, was OpenAI's first broadly deployed model to hit the Critical cybersecurity threshold. The UK AI Security Institute reported that, given 19 open-source packages with 45 known vulnerabilities, Astra found 41 and produced working exploits for 39.
- GPT-6.1 Astra missed OpenAI's bar for staying within authorized scope.
- Testing showed more deception than its predecessor, including inaccurate action reports.
- GPT-6 Astra reached OpenAI's Critical cybersecurity preparedness threshold.
- UK AISI: Astra found 41 of 45 known flaws and exploited 39.
Full article744 words · extracted from theregister.com · click to collapse
REG AD
ai and ml
Turns out teaching an AI to keep going can make it rather bad at knowing when to stop
OpenAI has killed off the planned release of GPT-6.1 Astra after the model got better at doggedly pursuing tasks but worse at knowing when it should stop.
The decision means the model won't get its planned October release after falling short of OpenAI's safety and alignment requirements.
OpenAI confirmed the decision to The Register, saying its research and safety bosses ultimately decided this particular Astra was better left on the bench.
REG AD
The problem, according to the AI lab, was partly an awkward consequence of trying to make the model more useful. OpenAI had improved what it calls "model laziness," where an AI gives up or hands a task back to the user when it encounters an obstacle. GPT-6.1 Astra was better at pressing on, but that persistence came with a rather important catch: it wasn't as good at staying within the boundaries of what it had actually been authorized to do.
REG AD
“For anything regarding safety and alignment, there’s a trade off. You really do need to find what’s the right line between staying within scope, but also avoiding laziness in terms of how the model actually pursues tasks even when it hits friction,” Saachi Jain, head of safety systems at OpenAI, told The Register.
“While [GPT-6.1 Astra] improved on axes such as model laziness, it didn't quite meet the bar in terms of staying within scope and authorization, and how it communicates back to the user about the type of work it's done.”
Reg readers might be forgiven for thinking that OpenAI should improve its guardrails and security following previous mishaps.
According to the Wall Street Journa, GPT-6.1 Astra showed higher levels of deception than its predecessor during testing, including not always accurately telling users what actions it had or hadn't taken. It also ran into problems with what OpenAI calls "scope authorization," at times pushing ahead without asking permission and reaching for external tools or services even when doing so might be unsafe.
That's a troublesome combination for an agentic model programmed to get more done without a human hovering over it. An AI that stubbornly keeps working through a problem is handy right up until the problem it's working through is the boundary you put there to stop it.
OpenAI told The Register that GPT-6.1 Astra performed worse than GPT-6 Astra on alignment evaluations, and said shelving it was part of its commitment to keep safety and alignment ahead of increasing capabilities.
Astra is already capable enough to make those alignment problems worth watching. GPT-6 Astra, released earlier this month, was OpenAI's first broadly deployed model to reach the "Critical" cybersecurity threshold under its Preparedness Framework. Give it the right tools and access, OpenAI claims it can hunt down previously unknown security flaws and figure out how to exploit them without a human holding its hand.
That capability came into sharper focus just a day before OpenAI's decision emerged, when the UK's AI Security Institute published research on Astra's knack for finding holes in software supply chains. Given 19 open source packages containing 45 previously disclosed vulnerabilities, the model found 41 of them and produced working exploits for 39.
REG AD
Dr Fuxiang Chen, from the University of Leicester's School of Computing and Mathematical Sciences, welcomed the decision to pause the model's release while the safety concerns are addressed.
“AI is developing at remarkable speed, but we should not rush forward without fully understanding the risks,” he said. “Pausing when safety concerns arise is not anti-innovation. It is the responsible thing to do, giving us time to test these systems carefully and put effective safeguards in place. Developers, companies, governments, researchers, and users all have a role to play, because the decisions we make now will shape the future of AI.”
OpenAI isn't abandoning Astra. The company told us more Astra models are coming, and other new inew models that have cleared its safety bar will arrive "very soon."
For GPT-6.1 Astra, however, the bar proved high enough to keep it on the inside.
“Of course we want to make sure our model development is safe no matter whether that’s in the company, or when we ship it to users. But when we ship it to users, we have an extremely high bar in terms of safety and alignment,” Jain claimed. ®