GPT-6 Astra gives mathematicians a breather, and OpenAI says that's by design
OpenAI's GPT-6 Astra tops ulam.ai's ErdosBench math benchmark with 106 of 226 problems solved, while the company prioritizes recursive self-improvement over math optimization.
OpenAI's GPT-6 Astra leads ulam.ai's ErdosBench with a score of 3.23, solving 106 of 226 open math problems and fully solving 43, ahead of GPT-5.6 Sol's 78 solved problems. Chief scientist Jakub Pachocki said OpenAI deliberately avoided targeted math optimization to prioritize recursive self-improvement and automated alignment research. Benchmark developer Przemek Chojecki estimated the gain at 5-10% across tested math-research skills. Mathematician Terence Tao warned at the 2026 International Congress of Mathematicians that AI-generated proofs could shift mathematics from proof scarcity to proof overload.
- Astra scored 3.23 on ErdosBench, solving 106 problems, 43 fully, and disproving 27
- GPT-5.6 Sol trails with 3.12 and 78 problems solved at maximum reasoning
- Pachocki says OpenAI skipped targeted math optimization to focus on RSI and alignment
- Terence Tao warns of proof overload as AI generates proofs faster than verification
Full article692 words · extracted from the-decoder.com · click to collapse
On ulam.ai's ErdosBench, Astra took first place. The benchmark covers 226 open math problems inspired by the famous Erdős problems. Astra scored 3.23 and solved 106 problems, 43 of them completely. It disproved 27 more. It also disproved 27 others.
Compared to Sol, which solved 78 problems at maximum reasoning, Astra showed stronger scientific writing and was less prone to overblown claims. In some cases, it actually understated its own results. Overall, benchmark developer Przemek Chojecki called it "a solid 5%-10% gain on various math-research skills tested," but the benchmark is far from saturated.

OpenAI chose not to optimize for math
OpenAI could have made Astra much stronger in math research but decided against it, even though the company had put math wins front and center in its first announcement. In his essay "An Alien Mind," OpenAI chief scientist Jakub Pachocki writes, "[…] we believe we could make the models better at specifically mathematics research with additional focus, but we do not prioritize this direction because of the urgency we feel about RSI and automated alignment research, as I will discuss later."
That means the most capable math model right now isn't the result of targeted optimization. It's a byproduct of other priorities. OpenAI is pouring its resources into recursive self-improvement and securing future AI systems, since "we believe it is the only way to remain at the frontier of AI research moving forward," Pachocki writes. (Note: Since the essay was published, OpenAI has reportedly trained better internal math models, a complicated story, and the topics of RSI and AI safety have exploded.)
Pachocki's statement is interesting for another reason, too: it shows that hard optimization trade-offs are being made at the jagged frontier of AI development. A model that dominates math won't necessarily dominate everything else. That reminded me of a visualization by Cambridge researcher Adam Hunt, which contrasts two possible paths for AI. The "mainstream AGI" thesis assumes models improve gradually and broadly until they cover all human tasks.
The alternative scenario describes an increasingly "spiky" trajectory, with extreme strength in a few domains like coding and math but stagnation or even regression in areas like language quality, common sense, or social reasoning. That would give us a highly specialized model, not something most people would call Artificial General Intelligence. Of course, you can call anything AGI if you feel like it.

Pachocki's point that OpenAI deliberately skipped math optimization for Astra is evidence for the "spiky" thesis. Even the leading AI lab can't push maximum progress in every direction at once. Capabilities grow where you optimize, and making math stronger means cutting back somewhere else. Behind the RSI priority, though, is the hope that the model will eventually make those optimizations itself, scaling faster across the board, including in math.
Math is wrestling with AI and with itself
Regardless of whether Astra is 5 or 50 percent better than its predecessor, the deeper question remains for a field that's grappling with a new world of compute power that helps solve problems but doesn't necessarily help understand them. Mathematician Terence Tao raised it at the 2026 International Congress of Mathematicians: if AI models keep producing proofs faster than humans can check them, the field risks shifting from proof scarcity to proof overload.
The critical task would then no longer be solving problems but deciding which results actually matter. Tao says math faces a crisis of its values and practices, one he compares to the foundational upheaval of the early 20th century.
The hardest problems in math remain unsolved for now, which should buy the discipline some time. The direction still seems set, though, even if some mathematicians don't think language models can deliver real breakthroughs because they lack human-like creativity. For similar reasons, there are also doubts about AI's potential for genuine self-improvement.
Text extracted automatically; images, tables and formatting may be missing. Original: https://the-decoder.com/gpt-6-astra-gives-mathematicians-a-breather-and-openai-says-thats-by-design/