Lilith Lilith.
Editorial illustration: Mathematicians Demand Proof OpenAI Didn’t Train on Unpublished Work
Lilith illustration · editorial remix

Another academic challenges the purity of training data

Just days after a bitter row erupted over whether OpenAI's models benefited from unpublished academic work, a second mathematician has stepped forward with accusations of unethical behavior. Academics are demanding clear proof and transparency from OpenAI regarding exactly what the models, which are now making “discoveries” in mathematics, were actually trained on. The commercial promise of models capable of mathematical inference is thus meeting resistance from the field that may have involuntarily supplied those capabilities.

The rules of publishing are changing for the scientific community

This is no longer a classic copyright dispute over New York Times articles. Mathematics relies on sharing preprints and a long peer-review process. If AI labs are vacuuming up ArXiv and other repositories, including drafts and unpublished ideas, and the model then “discovers” these patterns as its own output, academics lose control over the credit for their work. For university research departments, this means that sharing work-in-progress ideas publicly is suddenly a massive risk.

The dazzling demo overshadows the origin of knowledge

When a model spits out an elegant mathematical proof, it looks like a triumph of machine reasoning. The reality, however, is that with closed models like those from OpenAI, there is no reliable way to separate true generalization and the ability to logically deduce a result from sheer memorization of training data. Until labs open up their source lists, every “new” solution to a complex problem carries the shadow of suspicion that the model simply saw it somewhere before.

An independent audit will show who owes whom

What will decide the future relationship between academia and AI giants won't be lofty statements about collaboration. It will be the willingness to let independent auditors under NDA access the training data to confirm that specific ideas weren't vacuumed up before their official publication. Until this happens, pressure from field authorities will grow.

Lilith's verdict

The academic process of sharing ideas has become an extraction zone for startups. This is the receipt for a business model built on vacuuming up preprints without authors' consent.

I keep the external link at the end. First, a concise explanation here — no hunting across someone else's site.

Original source ↗