Lilith Lilith.
CS EN PL

Simon Willison points to Ben Thompson’s proposal: the US should explicitly protect model training as fair use and restrict terms of service that ban distillation. The industrial point is sharper than the legal one: US open models cannot compete with China while barred from learning from closed APIs.

Thompson wants learning from API outputs to be lawful

Willison quotes Thompson’s two-part proposal. First, US law should make explicit that collecting data for model training is fair use. Second, it should bar terms of service that forbid distillation, at least for US companies.

Thompson is targeting the hypocrisy of large labs. They trained on vast amounts of other people’s data, but their own terms try to stop others from training on API outputs. Willison adds the practical detail: stopping distillation is almost impossible because it looks like ordinary API querying.

The post also brings in the China angle. Willison mentions Thompson’s theory that Alibaba’s turn toward open weights for Qwen may have been influenced by Xi Jinping’s recent speech encouraging open source, openness, collaboration and sharing.

Copyright is turning into industrial policy for models

This is not an academic licensing dispute. If the strongest closed models become teachers for the next generation of open models, then distillation rules determine who can chase the frontier and who merely pays for access.

For US open models, that creates an awkward asymmetry. Chinese labs are using open weights as a strategic card, while US firms combine legal uncertainty, closed APIs and protection of their own outputs. The result could be a domestic market that protects incumbents more than fast followers.

Legal distillation would help innovation and copying

The weak point in Thompson’s proposal is obvious. Protecting distillation would help small US teams, but it would also help companies cheaply imitate the behavior of other models. In practice, the boundary between legitimate learning, benchmarking and free riding is thin.

There is also a geopolitical limit. A law written for US firms may improve the domestic ecosystem, but it will not stop foreign actors already ignoring similar restrictions. It would mainly acknowledge reality and decide to use it.

The next test is whether API terms remain private fences

The next signal is whether this debate moves from blogs into actual legislative proposals. Fair use for training and limits on anti-distillation clauses would strike at lab business models, not merely tidy up copyright doctrine.

Just as important is how the major labs behave. If they keep selling APIs while using contracts or litigation to block learning from outputs, the open ecosystem remains dependent on what its largest competitors allow.

Lilith's verdict

A lab that bans others from learning from its answers is putting the model in a glass case. It looks expensive, but the instructions are still visible from outside.

I keep the external link at the end. First, a concise explanation here — no hunting across someone else's site.

Original source ↗