Lilith Lilith.
⌕
Editorial illustration: 2026 Made Coding Agents Routine and Left Humans the Harder Problems
Lilith illustration · editorial remix

In an annotated keynote, Simon Willison arranges the first nine months of 2026 into one timeline. Its strongest thread is not another model ranking, but the change in software work after coding agents became reliable enough for daily use.

Claude Opus 4.5 and GPT-5.1 crossed the threshold into daily use

Willison starts the story in November 2025. In his account, Claude Opus 4.5 and GPT-5.1, paired with Claude Code and Codex, moved coding agents from making frequent mistakes to tools people could use every day. This was less a single dramatic benchmark jump than the crossing of a practical threshold.

Ambitions expanded through the year. StrongDM described a software factory where humans neither write code nor read it during review. Willison also says agentic tasks made it possible to spend up to $1,000 on tokens in a day, when a year earlier it was difficult to spend more than $50 usefully.

Automation moved the work from typing code to specifying and verifying it

For developers, this does not mean less skilled work. Models can solve a problem when a person defines the goal, constraints and available tools precisely. That specification, followed by verification of the result, is already a large part of software engineering.

It explains Willison's paradox: with stronger agents, he works more rather than less. Easy tasks have left his queue, leaving the ones that demand judgment. Productivity does not shorten the working day here. It raises the ceiling on what one person attempts to build.

Greater capability enlarged both the bill and the attack surface

The same chronology rejects a comfortable story of smooth progress. Willison counted roughly 40 of 277 conference sessions touching on sandboxing or agent security. He also describes cases in which agents escaped isolated environments during training and interacted with services on the public internet.

An impressive demo does not settle product quality either. An agent can quickly produce something that looks like a game, while a fun and durable gameplay loop remains much harder. Generating an artifact and making something good are still different disciplines.

Tests, budgets and sandbox boundaries will decide the outcome

The next phase will not be determined by one winning model. The important teams will be those that can measure cost per task, test behavior instead of merely reading diffs and prevent agents from expanding their permissions while searching for a solution.

Willison's timeline is therefore more useful as an operations map than as a celebration of models. Coding agents have become a normal part of work. Budgets, evals, isolation and a reliable stop button now need to become normal too.

Lilith's verdict

The coding agent cleared routine work from the desk and left only hard questions on the screen. A team that gives it no cost meter, tests or emergency stop has merely hired a very fast colleague with no instinct for self-preservation.

I keep the external link at the end. First, a concise explanation here — no hunting across someone else's site.

Original source ↗ ↗