Sharing AI progress in mathematics
OpenAI shared math results from an internal frontier model, with Lean proofs and compute details.
OpenAI published a set of mathematical results produced by an internal frontier model, following advice from the Institute for Advanced Study's Advisory Group on Mathematics and Artificial Intelligence. The release includes a GitHub repository with Lean formalizations of many proofs, ten reasoning summaries, ChatGPT Pro compute estimates, and statistics on attempted problems. The average result used roughly three hours of ChatGPT Pro thinking. OpenAI says it will fund workshops and conferences on AI-produced results and is working toward a responsible release of the model.
- Results come from an unnamed internal OpenAI frontier model.
- Many proofs are formalized in Lean and published on GitHub.
- Average result used about three hours of ChatGPT Pro thinking compute.
- OpenAI consulted the IAS Advisory Group on Mathematics and AI.
- The model itself is not released; OpenAI says it is working toward a responsible release.
Full article365 words · extracted from openai.com · click to collapse
We’re releasing a broad range of new mathematical results produced by an internal frontier model.
As we look to improve how we share results with the math community, we’ve been consulting with the independent Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study(opens in a new window) to develop best practices, and we have drawn on their advice and public recommendations(opens in a new window) to inform how we release these results.
For this release, we’re publishing the results in a GitHub repository, with protocols for paper revisions and citations. We’re continuing to explore other community-hosted alternatives for this release which meet the committee’s guidelines. For future releases, we are committed to further improving the quality of the papers via the citations, mathematical exposition, and presentation of the results for better understanding.
As part of our GitHub repository, we are sharing formalizations of many of the proofs in Lean, a programming language that allows mathematical proofs to be checked by a computer. We will update the repository with more formalizations as we obtain them.
To promote scientific transparency and openness, we are also publishing additional details about how we obtained the results in the repository. These include 10 summaries of the model’s reasoning, estimations of compute spent in terms of Pro usage on ChatGPT, and statistics about the number of attempted problems. The average result used the equivalent compute of roughly three hours of ChatGPT Pro thinking.
We want this progress to push the frontier of human knowledge and enable further progress in mathematics. We will be funding a series of workshops, conferences, and special programs around the understanding of major results produced by AI—we will share more on this in the near future.
We want to directly empower scientists with state-of-the-art capabilities and are working to responsibly release the model that produced these results. This is why it is important to continue to evaluate our internal frontier models on mathematics and other sciences, so we can accelerate developing the tools to advance those fields. We will continue to act on feedback from the community and update our standards for future disclosures of major scientific advancements.
Text extracted automatically; images, tables and formatting may be missing. Original: