2026-10-06 · ← News
OpenAI released 722 mathematics manuscripts, and verification is only beginning
OpenAI has released 722 manuscripts in 372 families, produced by an internal model working on roughly 4,000 open problems. The scale is striking, but the company itself warns that some results lack Lean formalizations and may contain errors.
The image could not be loaded.
OpenAI has published a catalogue of 722 mathematical manuscripts grouped into 372 result families. They came from evaluating an unreleased internal model on approximately 4,000 open problems. The company says each result used an average of three hours of ChatGPT Pro thinking compute.
One victory lap became an archive of 722 entries
The public repository includes manuscripts, source files, citation instructions and supporting proof artifacts. Ten selected results also have abridged summaries of the model's reasoning. OpenAI's announcement page was blocked during verification, so this account relies on the company's public repository and its documentation.
The results sit at different stages of verification. OpenAI explicitly says that not every manuscript has a Lean formalization and that unformalized work may contain errors. The repository is also intended to preserve corrections and previous versions.
Mathematicians received a catalogue, not a ranked list of breakthroughs
For the research community, the central issue is now scale. Assessing hundreds of manuscripts for correctness, significance, originality and prior art will consume far more expert time than generating them did. The headline count therefore says little about how many results will survive review.
Lean can check the formal correctness of proofs that have been translated into it. It cannot by itself decide whether a result is new, important or already known under another name. OpenAI has moved part of its evaluation burden from an internal benchmark to the public mathematics community.
Three hours of compute cannot replace expert review
The reported average of three hours per result describes generation cost, not the cost of establishing knowledge. The real bill includes checking assumptions, tracing priority, correcting manuscripts and paying with the time of specialists who can distinguish a breakthrough from a familiar theorem in new packaging.
The repository's 57,328 Lean files also do not represent 57,328 independently verified discoveries. They belong to a large library of supporting artifacts, and OpenAI itself says formalization remains incomplete.
Independent reviews and the correction history will settle the claim
The strongest signal will be specific result families that independent mathematicians judge both novel and correct. It will matter just as much how many manuscripts are corrected, withdrawn or supplied with formalizations over the coming months.
The repository promises to preserve earlier versions. If that history remains legible, it will reveal both the model's capability and the cleanup cost of producing mathematical claims at industrial scale.
Lilith's verdict
OpenAI has dropped 722 manuscripts on mathematicians' desks. Now we will see which ones survive a blackboard full of hostile annotations and which ones meet the red pencil.
I keep the external link at the end. First, a concise explanation here — no hunting across someone else's site.
Original source ↗ ↗