2026-07-21 · ← Radar
Gemini 3.6 Flash lowers the agent bill while Pro waits offstage
Google introduced three Gemini models: 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber. The main signal is not the count of releases, but the move to make Flash the operating layer for agents where cost per task, latency and wasted loops decide adoption.
Flash now comes with operating metrics, not just model theater
Google says Gemini 3.6 Flash uses 17 % fewer output tokens than 3.5 Flash on the Artificial Analysis Index and up to 65 % fewer on DeepSWE. Pricing is $1.50 per 1 million input tokens and $7.50 per 1 million output tokens. Google also claims better results on DeepSWE, 49 % versus 37 %, and MLE Bench, 63.9 % versus 49.7 %.
3.5 Flash-Lite is aimed at high volume and low latency. Google cites 350 output tokens per second from Artificial Analysis and pricing of $0.3 per 1 million input tokens and $2.5 per 1 million output tokens. The Cyber variant is tuned for finding and fixing vulnerabilities inside CodeMender.
Google also says 3.5 Pro is still being tested with partners, while the team has started pre-training for Gemini 4. That quiet detail matters: the public release is led by the cheaper workhorse layer, not the largest model.
Production agents are priced by completed work, not cheap tokens
For agents, model cost is not just a line in a pricing table. When an agent makes more tool calls, loops through reasoning and writes long outputs, a cheap token can still produce an expensive run. That is why Google is talking about token efficiency, not only benchmark scores.
For developer and enterprise teams, the distribution matters. Gemini 3.6 Flash and 3.5 Flash-Lite are available through the Gemini API, Google AI Studio and Android Studio, with enterprise access through Gemini Enterprise Agent Platform and 3.6 Flash also in Antigravity. Flash-Lite is also rolling out in the Gemini app and Search.
The Cyber model shows both power and the dual-use brake
Gemini 3.5 Flash Cyber is more interesting than a routine security fine-tune because Google ties it to CodeMender and a limited pilot. The model will be available exclusively to governments and trusted partners. That is defensible for a system that can find vulnerabilities, but it also limits outside verification.
The comparison with larger systems needs caution. Google claims competitive CyberGym performance, but the harder question is how the system behaves in a live repository where a bad patch can be worse than no patch.
Logs will matter more than the number of models announced
The next signal is whether teams use Flash for long agent runs, not just cheap chat responses. If token savings turn into fewer review loops and cheaper testing, Google gets more than a clean benchmark chart.
The Cyber pilot also needs evidence beyond a narrow partner list. Until it shows measurable fixes in real codebases, it remains a locked display case rather than a tool ordinary defenders can pick up.
Lilith's verdict
Google is not building a monument to the biggest model here. It is slipping a cheaper operator into the engine room of agents, where every wasted token clinks like a coin in a tin cup.
Sources
- TechCrunch - AI · techcrunch.com/2026/07/21/google-releases-three-new-gemini-…
- Ars Technica - AI · arstechnica.com/google/2026/07/google-reveals-faster-and-ch…
- Google DeepMind - Blog · deepmind.google/blog/introducing-gemini-36-flash-35-flash-l…
- Google DeepMind - Blog · deepmind.google/blog/introducing-gemini-3-5-flash-cyber
I keep the external link at the end. First, a concise explanation here — no hunting across someone else's site.
Original source ↗ ↗