Lilith Lilith.
Editorial illustration: Analysis of 500K papers shows Chinese models are becoming the baseline for research
Lilith illustration · editorial remix

Most of what we know about model adoption comes from benchmarks or marketing press releases, not hard data. Nathan Lambert had Codex analyze over half a million AI/ML papers published on arXiv since the introduction of ChatGPT. He was looking for specific mentions of open models used for research work.

The analysis reveals a shift. In 2024, roughly 30% of papers mentioned an open American model and only 10% a Chinese one. Today, 40% of publications acknowledge a Chinese open LLM, while American ones hover around 25 to 30% and are no longer growing.

Qwen becomes the global standard for research labs

The data shows strong inertia in adoption. Research and paper publication take months, so academic citations naturally lag behind model releases by one or two cycles. Today, one third of all papers that mention any LLM name Alibaba's Qwen. Meta's Llama, on the other hand, peaked around April 2025 at roughly a third of papers and its share has been declining since.

The role of closed commercial giants is interesting. Even though open models drive custom research efforts, OpenAI and its models maintain the highest overall mention rate at around 37%. Anthropic's Claude and Google's Gemini trail behind Qwen, Llama, and DeepSeek.

Citations only show what authors actually admit

The rise of DeepSeek after the release of version R1 in January 2025 is clearly visible in the data as a sharp spike, but the arXiv data still only captures what academics are willing and able to write in their texts. Furthermore, the analysis is limited to computer sciences, where citing architectures is standard practice.

Outside of machine learning, in humanities and enterprise deployments, market shares will look different, because there is no reason to worry about open weights and licensing, and the dominant factor is the API or web interface.

Tracking open-source influence in real time

The next generations of open models from both sides of the Pacific will soon show whether American players manage to reverse the declining curve in academic influence, or if Qwen remains the default infrastructure for global science. Cross-continental adoption of new versions will be decisive. Lambert's data is continuously updated on the Interconnects Open Model Dashboard.

Reliance on a single platform carries academic risks

The open approach of Chinese labs is currently de facto subsidizing global academic research. As long as it lasts, researchers have powerful tools available for free, but strategic dependence on resources outside their direct control could pose a problem if licensing changes.

Lilith's verdict

When operating on a limited academic budget, America no longer automatically hands you the best open keys for exploration.

I keep the external link at the end. First, a concise explanation here — no hunting across someone else's site.

Original source ↗