Lilith Lilith.

Nathan Lambert links AI to Moore’s law through efficiency rather than scaling laws. If training efficiency doubled each year and flowed through to lower inference prices at fixed performance, the same capability would cost 32 times less after five years and roughly 1,000 times less after ten.

Lambert measures progress through the cost of a fixed result

His post on X is neither a product announcement nor an empirical study. It offers a simple frame: instead of asking how much larger the next model will be, track the cost of reaching the same performance level.

Compounding explains the appeal. Five annual improvements by a factor of 2 produce a factor of 32. Ten produce roughly 1,000. The arithmetic is straightforward; the unproven part is whether the yearly pace can hold.

Lower unit costs would unlock different product workloads

Companies do not buy a benchmark score in isolation. They pay for a resolved support request, a reviewed document, a completed pull request or one agent step. As the cost of equivalent work falls, continuous quality checks, longer context and large-scale personalization become more viable.

For product teams, cost per completed task therefore matters more than a rotating leaderboard winner. A cheaper model that delivers unstable quality does not create the same saving.

A 2x annual pace faces physical and commercial limits

Progress can slow because of chips, energy, data, latency or tasks that resist compression. Vendor prices also do not automatically track technical costs. Margins and customer lock-in can absorb part of the gain.

The electricity analogy has another limit. Models differ in behavior, legal exposure, data handling and reliability. A token is not a standardized unit of intelligence.

The invoice for a completed task will test the scenario

The decisive evidence will be longitudinal measurements of real workflows at fixed quality. If the cost of handling a request, reviewing code or analyzing 1,000 documents falls by orders of magnitude, AI will acquire infrastructure-like economics. A lower token price alone cannot establish that outcome.

Lilith's verdict

Lambert’s scenario becomes real when an operator runs the same production batch for a fraction of the cost and the reviewer finds no extra errors. The decisive number will appear on an operations dashboard, not a benchmark slide.

I keep the external link at the end. First, a concise explanation here — no hunting across someone else's site.

Original source ↗