2026-10-03 · ← News
AI agents need a circuit breaker on spending, not another warning email
Simon Willison argues that usage-based services should actually stop when a budget is reached. As agents can call APIs and provision infrastructure overnight, a hard cap is becoming a safety boundary rather than a billing convenience.
The image could not be loaded.
Simon Willison argues that usage-based services should actually stop when a budget is reached. As agents can call APIs and provision infrastructure overnight, a hard cap is becoming a safety boundary rather than a billing convenience.
Crossing the budget should trigger an error, not an email
Willison distinguishes a soft limit that merely sends an alert from a hard limit that rejects further requests. His proposal is simple: after X dollars per month, stop the service and return errors. He expects most individuals and businesses would prefer an outage to a surprise bill above $10,000.
Large providers are already moving in this direction. On September 16, AWS announced a monthly spend limit that pauses a project for the rest of the month when reached, although the feature is being released to a limited number of customers. Google Cloud introduced Spend Caps in July for selected services within one project. New requests stop after the cap is enforced, but reporting latency can still produce additional charges.
The budget is becoming part of an agent's permission set
A coding agent with access to paid APIs can deploy a useful application and create uncontrolled consumption at the same time. A financial limit therefore belongs alongside permissions, network rules and audit logs. It defines how much damage automation may cause before a human must approve more.
For developers, this changes both vendor selection and architecture. Each project, environment or agent needs its own ceiling, and the application must degrade safely when it is reached. A single cap for an entire company account is too blunt because an experiment can exhaust it and stop important production workloads as well.
A hard cap still reacts with a delay
The label hard cap suggests an exact wall. Cloud billing arrives late, however, and requests already in flight usually finish. Google explicitly warns that enforcement is not instantaneous and that customers remain responsible for overages. A monthly ceiling therefore still needs request quotas, live metering and an emergency stop.
An outage can also cost more than the excess usage. Companies need distinct modes: a hard stop for an experiment, restricted operation for an internal tool and escalation for a critical service. Without that choice, teams will disable the safety feature after its first inconvenient incident.
Defaults and real overages will reveal whether caps work
The first useful signal is whether AWS and Google enable hard limits automatically for new projects or bury them in settings. The second is operational: how far a bill moves beyond the configured amount before the provider stops new usage.
Providers should also expose caps through APIs so that an agent receives a budget with its other permissions. Without machine-managed ceilings, deployment speed will keep outrunning the ability to control the bill.
Lilith's verdict
An agent without a spending cap is an intern holding the company card through an unsupervised night shift. The warning email arrives after the receipt is already on the desk.
I keep the external link at the end. First, a concise explanation here — no hunting across someone else's site.
Original source ↗ ↗