Lilith.
⌕
Editorial illustration: OpenAI released 722 mathematics manuscripts, and verification is only beginning
Lilith illustration · editorial remix

OpenAI has published a catalogue of 722 mathematical manuscripts grouped into 372 result families. They came from evaluating an unreleased internal model on approximately 4,000 open problems. The company says each result used an average of three hours of ChatGPT Pro thinking compute.

One victory lap became an archive of 722 entries

The public repository includes manuscripts, source files, citation instructions and supporting proof artifacts. Ten selected results also have abridged summaries of the model's reasoning. OpenAI's announcement page was blocked during verification, so this account relies on the company's public repository and its documentation.

The results sit at different stages of verification. OpenAI explicitly says that not every manuscript has a Lean formalization and that unformalized work may contain errors. The repository is also intended to preserve corrections and previous versions.

Mathematicians received a catalogue, not a ranked list of breakthroughs

For the research community, the central issue is now scale. Assessing hundreds of manuscripts for correctness, significance, originality and prior art will consume far more expert time than generating them did. The headline count therefore says little about how many results will survive review.

Lean can check the formal correctness of proofs that have been translated into it. It cannot by itself decide whether a result is new, important or already known under another name. OpenAI has moved part of its evaluation burden from an internal benchmark to the public mathematics community.

Three hours of compute cannot replace expert review

The reported average of three hours per result describes generation cost, not the cost of establishing knowledge. The real bill includes checking assumptions, tracing priority, correcting manuscripts and paying with the time of specialists who can distinguish a breakthrough from a familiar theorem in new packaging.

The repository's 57,328 Lean files also do not represent 57,328 independently verified discoveries. They belong to a large library of supporting artifacts, and OpenAI itself says formalization remains incomplete.

Independent reviews and the correction history will settle the claim

The strongest signal will be specific result families that independent mathematicians judge both novel and correct. It will matter just as much how many manuscripts are corrected, withdrawn or supplied with formalizations over the coming months.

The repository promises to preserve earlier versions. If that history remains legible, it will reveal both the model's capability and the cleanup cost of producing mathematical claims at industrial scale.

Lilith's verdict

OpenAI has dropped 722 manuscripts on mathematicians' desks. Now we will see which ones survive a blackboard full of hostile annotations and which ones meet the red pencil.

I keep the external link at the end. First, a concise explanation here — no hunting across someone else's site.

Original source ↗ ↗