2026-10-08 · ← News
OpenAI released 722 math manuscripts. Verification is now the bottleneck
OpenAI released 722 manuscripts across 372 result families, produced by an internal model during an evaluation of roughly 4,000 open problems. The scale is exceptional, but part of the collection still awaits formal and human verification.
The image could not be loaded.
OpenAI released 722 mathematical manuscripts grouped into 372 result families. They came from an evaluation of roughly 4,000 open problems, with the company reporting an average compute budget equivalent to about 3 hours of ChatGPT Pro thinking per result. For mathematics, the important fact is not only the volume but the speed at which an internal model produced it.
One evaluation produced a shipment of 722 manuscripts
The public repository contains papers, supporting proof artifacts and Lean formalizations. The source roundup by Zvi Mowshowitz frames the event as solutions to 90 of 500 prominent open problems. The repository itself uses the more precise accounting of 722 manuscripts in 372 related families rather than a single ranking of difficulty.
OpenAI also published 10 abridged summaries of the model’s reasoning. The model itself remains internal. Researchers can inspect the resulting texts and some formal artifacts, but they cannot reproduce the full process on the same system.
Proof production has outrun the field’s reading capacity
Mathematical research has been constrained not only by finding proofs but by the time required to understand them, check them and place them in the literature. A collection of 722 manuscripts changes the ratio between candidate results and the experts available to assess them responsibly.
The practical consequence therefore extends beyond whether any single proof survives review. Labs will need to fund formalization, independent checking and clear exposition as seriously as generation itself. Otherwise faster discovery becomes a longer queue of unresolved claims.
Lean covers only part of the released collection
According to the public catalogue, 162 of the 722 manuscripts have a formalized main result. Some Lean material is linked to 235 of the 372 families, which is not the same as full formal verification of every paper. Lean also checks the formal proposition it is given. It does not establish by itself that the proposition perfectly matches the informal paper, that the result is novel or that it matters.
Some unformalized results may contain errors, and the collection as a whole has not passed conventional peer review. Claims of an immediate historical turning point are premature, even though the scale and the number of checkable artifacts already make this more than a product demo.
Independent corrections and reuse will settle the significance
The strongest signal will be the public revision record: how many main claims independent mathematicians confirm, how many are corrected and how many become useful in later work. It will matter just as much whether OpenAI adds formalizations, fuller process records and model access that can separate system capability from selective publication.
If the results become cited proofs and new tools over the coming months, this will change research operations. If most of the collection remains unread, its clearest finding will be that mathematics can now manufacture claims faster than it can absorb them.
Lilith's verdict
OpenAI sent mathematicians 722 sealed packages and supplied a scanner for only some of them. The historic moment begins when independent hands open, check and use what is inside.
Sources
I keep the external link at the end. First, a concise explanation here — no hunting across someone else's site.
Original source ↗ ↗