OpenAI math results draw disputes over counts and standards
OpenAI published math results from an unreleased model, many in Lean; sources later split on 372 versus 719 items and on scholarly standards.
On Oct. 6, 2026, OpenAI said an unnamed internal frontier model produced mathematical results published on GitHub with Lean formalizations of many proofs, ten reasoning summaries, ChatGPT Pro compute estimates, and statistics on attempted problems. It said the average result used roughly three hours of ChatGPT Pro thinking, that it consulted the Institute for Advanced Study's Advisory Group on Mathematics and Artificial Intelligence, that it will fund workshops and conferences, and that the model itself was not released while it works toward a responsible release. On Oct. 7, The Decoder described 372 model-generated proofs with revision logs and citations, claiming each solves an open problem or makes substantial progress, including algorithm improvements, Riemann-hypothesis-related work, and a Navier-Stokes result still under formal review; it said nearly every result came from one prompt and prompts were not shared. That outlet's summary says Fields medalists are split, while one of its bullets says twenty-five warned that mass-produced proofs may undermine mathematical understanding. On Oct. 8, TechCrunch instead reported 719 claimed solutions and said the advisory group, which it places at Princeton's Institute for Advanced Study, had urged labs not to test proprietary models on advanced problems, differing from OpenAI's account of following the group's advice. TechCrunch added that only ten manuscripts included chain-of-thought, 42 percent of proofs were not formalized, critics say the release misses standards for human understanding, and a University of Cambridge and King's College London paper found at least two natural-language versus Lean discrepancies on a Navier-Stokes-derived problem.
- On Oct. 6, 2026, OpenAI said an unnamed, unreleased internal frontier model produced math results posted on GitHub, with Lean formalizations of many proofs, ten reasoning summaries, compute estimates, and statistics on attempted problems.
- OpenAI and The Decoder agree the average result used about three hours of ChatGPT Pro thinking; The Decoder adds that nearly every result came from one prompt and that prompts were not shared.
- Counts conflict: The Decoder (Oct. 7) reports 372 GitHub proofs with revision logs and citations; TechCrunch (Oct. 8) reports 719 claimed solutions to open problems.
- Cited work includes algorithm improvements, results related to the Riemann hypothesis, and a Navier-Stokes result The Decoder said remained under formal review.
- TechCrunch says only ten manuscripts included chain-of-thought and that 42 percent of the proofs had not been formalized.
- A University of Cambridge and King's College London paper found at least two mismatches between OpenAI's natural-language argument and its Lean code for a Navier-Stokes-derived problem.
- OpenAI says it followed the Institute for Advanced Study advisory group and will fund workshops while working toward a responsible model release; TechCrunch says that group, hosted at Princeton's IAS, had urged labs not to test proprietary…
- The Decoder's summary says Fields medalists are split, while its own bullet says twenty-five warned mass-produced proofs may undermine understanding; TechCrunch says critics find the release short of standards for human understanding.
Coverage timelineoldest first · each row is one article
- · 2d agoSharing AI progress in mathematics
OpenAI News· 58
OpenAI shared math results from an internal frontier model, with Lean proofs and compute details.
- · 1d agoOpenAI dumps 372 AI-generated math proofs on GitHub, telling the academic world to keep up
The Decoder· 67
OpenAI published 372 AI-generated math proofs on GitHub, including Lean formalizations and claimed progress on open problems.
- · 5h agoOpenAI’s math solutions aren’t meeting the field’s standards yet
TechCrunch · AI· 62
Mathematicians say OpenAI's hundreds of claimed hard-problem proofs miss standards for human understanding.