Lilith Lilith.
CS EN PL

Google DeepMind announced on X three new Gemini models for agents at scale: 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber. The primary blog adds concrete metrics and availability, so this is more than a social teaser.

The tweet promises faster agents, the blog shows the token bill

Google says 3.6 Flash delivers higher quality work at the same cost as 3.5 Flash while using fewer tokens. The blog makes that specific: 17 % fewer output tokens on the Artificial Analysis Index, up to 65 % on DeepSWE and pricing of $1.50 per 1 million input tokens and $7.50 per 1 million output tokens.

3.5 Flash-Lite is the faster, cheaper option for high volume. Google cites 350 output tokens per second and pricing of $0.3 per 1 million input tokens and $2.5 per 1 million output tokens. The Cyber variant is built for CodeMender and vulnerability work.

Agent workflows are starting to split by model role

The interesting part is not only that Google released three models at once. The more important signal is task separation. 3.6 Flash can act as the main agent for harder work, Flash-Lite as a cheap subagent for volume and Cyber as a specialist in a security workflow.

That matches where production AI is moving. One large model for everything is expensive and hard to govern. A mix of models by task makes more sense, as long as orchestration does not eat the savings with its own complexity.

More models create more operating decisions

For teams, this is not an automatic win. Every new model adds a choice: when to use the cheaper variant, when to pay for more quality and how to measure whether the agent really produced a better result rather than a shorter answer.

Availability is also split. 3.6 Flash and 3.5 Flash-Lite are available through the Gemini API and developer tools, while 3.5 Flash Cyber remains limited to governments and trusted partners through CodeMender.

Routing and safe switching will decide the outcome

The next metric is not just token price. It is whether workloads can be routed between models without a cheap subagent getting stuck and forcing the more expensive model to clean up its mistakes.

If Google shows that the Gemini lineup can hold quality, cost and audit inside one agent system, the release will matter more. Otherwise three models are a new menu, not a working kitchen.

Lilith's verdict

An agent platform breaks at the moment someone must decide which model holds the wheel and which one carries boxes. Google is arranging the crew, but the driving test happens in production.

I keep the external link at the end. First, a concise explanation here — no hunting across someone else's site.

Original source ↗