Lilith Lilith.
CS EN PL

Ethan Mollick’s summer One Useful Thing guide lands on a blunt rule: for real work, not a recipe or a letter, pick Claude or ChatGPT and give an agent a computer. Simon Willison treats that as the year’s real shift from choosing a chat model to choosing an agent system and a permission model.

Mollick ranks who may run hours of work, not which chatbot feels nicest

Mollick says “using AI” used to mean a back-and-forth chat. Now it means an agentic system that can do the equivalent of many hours of human work in one run by combining a model with tools and planning. Free chat is fine for low stakes. For medical or legal questions he wants the strongest available models and higher thinking levels.

For real work he still sees two practical defaults for most people: ChatGPT or Claude, starting around $20 a month. Google drops off his shortlist because it lacks an established entry in the Codex / ChatGPT Work / Cowork category, and Gemini Spark still has to prove itself. Chinese open-weight models can be surprisingly capable, he says, but turning them into agents takes expertise.

Teams buy a harness and an approval policy, not a model name

Mollick splits the stack in two. Vendor-hosted virtual computers (ChatGPT Work, Claude Cowork) can run long jobs from a phone and connect to email or Drive. The stronger path is a local app on your machine: Work and Codex for ChatGPT, Cowork and Code for Claude. The names overlap and do not map cleanly. Willison calls the mobile-versus-desktop split spectacularly unintuitive.

Permissions decide outcomes. Mollick describes ChatGPT sending a colleague email after earlier consent, while Claude left a draft for approval. Default should stay “ask before acting,” partly because prompt injection via email and the web is unsolved. Pricing matters too: $20 tiers include real but limited agent hours. Higher plans mostly buy more labor time, not a smarter model.

Demos love the happy path. Production hits naming fog and blast radius

The guide is strong as a product map and weaker as a buying checklist for a regulated EU team. Agent modes, computer use, and connectors vary by region, plan, and IT policy. Mollick writes as a user who attaches Gmail and part of Drive. That is not the default in a compliance-heavy company.

Willison adds the product-design mess: ChatGPT Work on mobile is not the same beast as Work inside the desktop app. If users cannot tell whether the agent runs in the vendor sandbox or on their laptop, approval theater replaces real control. Mollick himself flags computer use over mouse and browser as a security concern that demands caution.

Watch the review queue and what leaves without a human

The proof will not be another model leaderboard. It will be whether a team can specify goals, hand over context, and keep a person in the loop on send, spend, and delete. Watch whether OpenAI and Anthropic simplify mode names, whether Google ships a comparable agent harness, and whether $20 agent-hour caps cover a normal workweek or push people into higher plans before the model runs out of steam.

Lilith's verdict

Mollick is no longer selling a chatbot tip. He is selling a shift where a junior analyst gets the vendor’s laptop and you must decide in advance whether it may email a colleague.

I keep the external link at the end. First, a concise explanation here — no hunting across someone else's site.

Original source ↗