2026-10-06 · ← News
OpenAI released 722 math manuscripts, moving the bottleneck to verification
OpenAI has released 722 mathematical manuscripts across 372 result families, produced by an unreleased internal model. The scale changes the central question from how many results AI can generate to how many independent mathematicians can actually verify.
The image could not be loaded.
OpenAI has published a GitHub catalogue of 722 mathematical manuscripts organized into 372 families of related results. The company says the vast majority came from one procedure using an unreleased internal model, tested on approximately 4,000 open problems. Each result used an average of three hours of compute equivalent to ChatGPT Pro thinking.
An evaluation became a catalogue of 722 manuscripts
The repository includes PDFs, source files, citation instructions and supporting proof artifacts. Ten selected results have abridged summaries of the model's reasoning. OpenAI also provides a catalogue of Lean formalizations, while explicitly warning that not every manuscript has been formalized and that unformalized results may contain issues.
A result family can contain a principal claim, alternative proofs and related consequences. The 722 manuscripts therefore do not represent 722 independently solved problems. The count of 372 families is the more useful measure of the release than the number of files alone.
Proof production has outrun the field's reading capacity
The practical problem for mathematicians is the allocation of attention. One team can generate hundreds of papers, while checking each proof still requires a specialist, time and often a formal reconstruction. GitHub solves distribution and versioning, not scholarly trust.
That changes the economics of research work. Value will come not only from a model that proposes a proof, but from systems that rank results by significance, expose assumptions and route disputed steps to the right experts. The deeper story is research surplus management rather than a single mathematical triumph.
Lean covers part of the collection while the rest awaits scrutiny
A formal artifact can show that a particular statement passes within a specified checking environment. It cannot by itself establish novelty, importance or correct placement in the literature. OpenAI also describes exceptions to its standard procedure and one manuscript edited by a human for readability, so the catalogue is not a homogeneous benchmark produced by one uniform method.
The central risk is that volume creates the appearance of consensus before independent scrutiny occurs. Across 372 result families, the public correction history and the ability of mathematicians to reproduce crucial steps without trusting the private model will matter more than the launch claim.
Confirmations and corrections will define the release's real value
Watch how many families receive independent verification, how many manuscripts are revised and where substantial counterarguments emerge. It will also matter whether OpenAI adds more formalizations and whether review expands beyond groups involved in organizing the release.
Another hundred PDFs would be a weak success metric. Time from publication to verification, and the share of results that survive expert review in their respective fields, will tell us what this catalogue actually achieved.
Lilith's verdict
OpenAI has placed 722 manuscripts in front of mathematicians at once. Only those still standing after months of independent annotations will carry real weight.
I keep the external link at the end. First, a concise explanation here — no hunting across someone else's site.
Original source ↗ ↗