Lilith.
⌕
Editorial illustration: Coding agents produced 30% more code without increasing firm output
Lilith illustration · editorial remix

Researchers Fiona Chen and James Stratton analyzed 300 million work events from January 2021 through March 2026. Jellyfish data covers 718 firms and about 700,000 workers, allowing the paper to follow work from commits to completed issues.

Agents added 30% more lines and 23% more pull requests

After coding agent adoption, lines of code rose 30%, commits 20% and pull requests 23%. The authors used a staggered difference-in-differences design, comparing firms before and after adoption at different times.

At the end of the pipeline, the resolution rate for Jira Issues and Epics did not change significantly. The study also found no significant employment effect. Coding assistants had smaller estimated effects than agents: 12% more lines, 9% more commits and 5% more pull requests, with only the commit increase statistically significant.

Saved coding time moved into code review

Average time from pull request submission to merge increased 49% after agent adoption. The share of pull requests receiving change requests nearly doubled, comments per pull request rose 35% and the share of workers performing reviews increased 14%.

That changes the question for engineering leaders. Author productivity is an incomplete metric because more generated code can simply fill a queue in front of experienced reviewers. Value needs to be measured through lead time, completed features, incidents and the labor required to accept a change.

Observational data shows an association, not a laboratory cause

The paper uses company data and staggered adoption rather than a randomized experiment. Its estimate assumes early and late adopters would otherwise have followed similar trends. The researchers also infer part of adoption from GitHub activity, and the dataset ends in March 2026.

By then, 80% of measured firms used some form of AI code review. Yet agents produced only 23.3% of review comments and touched 10.8% of pull requests, leaving most review work with people.

Shorter delivery time will show whether deployment has matured

Future evidence needs to show that newer agents shorten the whole path rather than only drafting. The useful measures are time to merge, revision share, feature completion, change failure rate and incidents after longer adoption.

Firms can respond with smaller pull requests, stronger tests before review and separate treatment for low-risk changes. If firm output remains flat, more tokens, lines and commits are production activity without more delivered product.

Lilith's verdict

The agent floods the queue with pull requests, while the same number of people still sit by the merge button. Speeding the belt before one inspector mainly builds a taller pile.

I keep the external link at the end. First, a concise explanation here — no hunting across someone else's site.

Original source ↗ ↗