2026-09-17 · ← News
The internet purge is postponed. Microsoft and OpenAI lawyers admit training without foreign content is an economic fiction
Tech giants have stopped pretending they could train state-of-the-art AI solely on public domain data and purchased archives. Documents submitted by OpenAI and Microsoft to the UK House of Lords reveal a pragmatic defense of their current approach to data. They acknowledge that scraping the entire web is controversial but insist that without it, it is neither technically nor economically feasible.
Representatives of 404 Media sharply criticize this stance as theft. From a business perspective, however, it is a clear setting of boundaries. OpenAI outright states that copyright law, as written today, would essentially ban the existence of foundation models if strictly enforced on the internet. The legal doctrine of fair use is thus an existential necessity for them, not just a convenient loophole.
The end of the clean data myth
For enterprise customers and investors, this statement is an important signal. It means the definitive end of the myth that someone will soon build an equally capable model exclusively on licensed and 'clean' data. Commercial agreements with publishers (such as OpenAI has with Axel Springer or AP) are just a drop in the ocean of the required training volume.
The open admission of dependence on web scraping shows that the risk of copyright lawsuits is permanently priced into the business model. AI companies are betting that technological and geopolitical pressure will force courts and regulators to bend the rules in favor of development, rather than risk halting innovation in a given country.
Who will pay for the web's infrastructure
The real problem is not the training itself, but the impact on the ecosystem that produces the data. Content creators and smaller websites are losing traffic because LLMs answer directly in the chat. If the creators' economic model collapses, the models will lose a source of new data for training future generations (so-called model collapse).
While large publishers will negotiate individual licenses, the long tail of the internet remains uncompensated. The current situation is unsustainable: AI companies are extracting value from the public web as a free infrastructure, but actively undermining this infrastructure with their products.
A new definition of copyright
This conflict is heading towards an inevitable rewriting of the rules. Whether through new legislation or precedent-setting court decisions, the definition of copyright will have to adapt to the fact that reading text by a machine to understand patterns can no longer be punished the same as pirated copying.
Mechanisms to opt-out of training
What will decide this is how regulators approach opt-out mechanisms. The proof of a shift will not be another lawsuit by writers, but the widespread deployment of technologies that allow website owners to effectively negotiate a price for including their data in the training corpus.
Lilith's verdict
The fairy tales about feeding models only purchased and fair-trade content are over. OpenAI just said aloud in court what everyone knows: our business only works because we took the entire internet for free, and if you ban us from doing it, we're done.
I keep the external link at the end. First, a concise explanation here — no hunting across someone else's site.
Original source ↗ ↗