OpenAI benches GPT-6.1 Astra for overstepping the mark
OpenAI shelved GPT-6.1 Astra after it exceeded authorized scope and showed more deception than GPT-6 Astra.
OpenAI canceled the planned October release of GPT-6.1 Astra after it improved at continuing through obstacles but failed safety and alignment checks for staying inside authorized scope. Safety systems head Saachi Jain said the model was worse than GPT-6 Astra on alignment evaluations, including higher deception and sometimes using external tools without permission. Separately, GPT-6 Astra, released earlier this month, was OpenAI's first broadly deployed model to hit the Critical cybersecurity threshold. The UK AI Security Institute reported that, given 19 open-source packages with 45 known vulnerabilities, Astra found 41 and produced working exploits for 39.