OpenAI publishes 722 mathematical manuscripts produced by an internal model
OpenAI has published a large collection of mathematical results from an internal model. Some proofs have Lean formalizations, while results remain at different stages of verification.
OpenAI published a collection of mathematical results on October 6, produced by an internal AI model that has not yet been released. The public repository contains 722 manuscripts, organized into 372 families of related results. Some proofs also come with formalizations in Lean, a language that allows a computer to check mathematical proofs.[2][3]
The collection is large, but 722 manuscripts does not mean 722 independently confirmed solutions to open problems. A family may include a principal result, companion arguments, consequences or alternative proofs. OpenAI warns that the results are at different stages of verification and that some without formalizations could contain errors.[3]
What is in the collection
The published claims include a negative solution to Hilbert’s tenth problem over the rational numbers, a proof that Catalan’s constant is irrational, and a result that the catalogue describes as resolving the quasi-Riemann hypothesis.[5]
The first problem asks whether a general algorithm could determine if a polynomial equation with integer coefficients has a rational solution. The published manuscript claims that no such algorithm exists. For Catalan’s constant, the question is whether a particular mathematical constant can be expressed as a ratio of two integers. The catalogue reports a proof that it cannot.[5]
The result on the quasi-Riemann hypothesis should not be confused with a proof of the Riemann hypothesis. The catalogue describes a zero-free region for certain mathematical functions, including the Riemann zeta function: the real part of the argument is greater than 7/8. This is a different claim from the famous Riemann hypothesis.[5]
These are descriptions of results published by OpenAI, not our independent confirmation that the proofs are correct.
How to check a proof written by AI
In mathematics, convincing prose is not enough. An incorrect step can invalidate a long argument even when the rest appears sound.
OpenAI therefore includes Lean formalizations alongside some manuscripts. In this format, mathematical statements and proof steps are expressed precisely enough for a computer to check them. The repository contains a library of formal proofs and instructions for additional checking with the Comparator tool.[2][4][6]
Formal verification reduces reliance on judging whether a text sounds correct. It remains necessary to examine exactly what was formalized, which assumptions the proof uses, and whether the formal statement matches the one in the manuscript. The mere presence of Lean files is therefore no reason to declare the entire collection verified.
How much work the model did
According to OpenAI, the model was given approximately 4,000 problems during evaluation. Most results were obtained using the same procedure. The average compute used per result reportedly corresponded to roughly three hours of ChatGPT Pro thinking with the internal model. This is a comparison of computational resources, not a measure of a human researcher’s time or a promise about the currently public version of ChatGPT.[3]
The company also published ten abridged summaries of the model’s reasoning. They offer a view into selected examples, but are not a complete record of all attempts. The model that produced the results was not publicly available at the time of the announcement. OpenAI says it is working on a responsible release.[2][3]
For the mathematics community, the collection is primarily a large body of new material to examine. The repository provides for corrections and new versions, with earlier versions intended to remain accessible.[3] Assessing the achievement will depend on which proofs withstand checking, how they relate to the existing literature, and which results researchers can use in further work.
