Lilith Lilith.
⌕
Editorial illustration: Gemini 4 Argon raises output to 1 million tokens, but the API is still closed
Lilith illustration · editorial remix

Google has introduced Gemini 4 Argon with a 1 million token output limit, up from 64,000, and shown it working on large internal projects. Developers still cannot test it themselves, leaving a wide gap between an impressive demonstration and a deployable product.

Argon gets room for a long run rather than another short chat

The clearest change is a maximum output of 1 million tokens. Google says the previous ceiling was 64,000 tokens. This is an output limit, not a newly confirmed input context window. The aim is to keep a long task within one trajectory and reduce the loss of state between successive runs.

Google describes several internal deployments. It says agents identified data-centre memory savings that released more than 300 TiB after rollout, with total estimated savings of 500 TiB to 1 PiB. Other teams are using Argon to migrate C and C++ projects to Rust, ranging from libraries with tens of thousands of lines to the Fuchsia Zircon kernel with more than 800,000 lines.

For the libgav1 decoder, agents replaced 32,000 lines of SIMD code in an existing Rust port. Google says the result runs 2.7 times faster than the previous Rust port while producing identical video output. The company says such changes undergo automated and manual audits, emulation testing and review.

Longer output changes both agent design and the bill

For an engineering team, a larger output budget could reduce restarts during a repository migration, extended research or difficult debugging. It also moves the constraint elsewhere. Hundreds of thousands of tokens still need to be reviewed, tested and paid for. A longer run does not guarantee a sound plan or a correct result.

Google lists an introductory price of $2 per million input tokens and $10 per million output tokens, with cached input discounted by 95%. After the introductory period, those rates are due to rise to $4 and $20. The company has not given an end date for the offer or a public API model ID.

The benchmark table mixes self-evaluation with different harnesses

Google reports 77.9% on DeepSWE v1.1, 51.3% on AutomationBench and 91.7% on LVBench. Its methodology document also says that Google computed many Argon results itself, while competitor scores often came from providers or public leaderboards. Harnesses, tool limits and run conditions vary across models.

The most interesting internal examples also come from the company selling the model. They show what a well-equipped Google team can build around Argon. They do not yet show what an ordinary customer can achieve with its own repository, permissions and tests.

Public API access and outside repositories will separate capability from presentation

Argon is currently reaching only a selected group of trusted defenders and testers through the Fairwind Program. Google promises later access for paid API customers and Google AI Ultra subscribers, but has announced no date. Buying a plan is therefore not enough to use the model today.

The decisive signals will be reproducible tests on outside codebases, real latency for long runs, the number of human interventions and the cost of a completed task. One million tokens is a large fuel tank. Production use will show whether Argon travels farther or merely burns the budget for longer.

Lilith's verdict

One million tokens gives the agent a longer leash. An outside repository will reveal whether it reaches the finish line or wraps the whole team in it.

I keep the external link at the end. First, a concise explanation here — no hunting across someone else's site.

Original source ↗ ↗