2026-07-21 · ← Radar
Gemini 3.6 Flash cuts output tokens by 17%, while Cyber stays restricted
Google has split the new Flash lineup by workload and risk. Gemini 3.6 Flash targets complex agent tasks, 3.5 Flash-Lite targets inexpensive high-volume execution and 3.5 Flash Cyber targets vulnerability discovery and repair inside a controlled program.
Flash now comes in performance, efficiency and security variants
Google says Gemini 3.6 Flash uses 17% fewer output tokens than 3.5 Flash on the Artificial Analysis Index. It costs $1.50 per million input tokens and $7.50 per million output tokens. Google reports a DeepSWE score of 49% versus 37% for its predecessor, while OSWorld-Verified rises to 83.0% from 78.4%.
Gemini 3.5 Flash-Lite runs at 350 output tokens per second according to Artificial Analysis. Pricing is $0.30 per million input tokens and $2.50 per million output tokens. Google has made both general models available through the Gemini API, Google AI Studio and Gemini Enterprise. They also appear in the Gemini app, with Flash-Lite rolling out in Google Search.
The Cyber variant is a fine-tune of 3.5 Flash for finding, validating and repairing vulnerabilities. CodeMender orchestrates several agents into one combined report. Google plans to offer it only to governments and trusted partners through a limited pilot that is coming soon.
Developers are getting a workload router, not one universal winner
The product argument rests on combining models. The more capable 3.6 Flash can direct a complex process while Flash-Lite handles search, document processing or other high-volume steps. Teams can optimize cost at each stage of an agent instead of deploying the same model everywhere.
Both general models include computer use as a built-in tool. That simplifies architecture, but makes permissions, isolation and action logs more important. Token savings matter only when a cheaper run does not create a more expensive queue of manual reviews.
Most benchmark evidence still comes from the vendor's table
Google combines external Artificial Analysis measurements with its own results and customer demonstrations. The numbers describe the claimed generational change, but they do not guarantee performance on a specific production workflow. Agent costs also depend on repeated tool calls, failed steps and total run length, not merely the token price.
Restricting Flash Cyber is a defensible response to dual-use capability. It also prevents ordinary security teams from independently checking the CyberGym claim. Trust in this variant will depend on published evals and partner results.
Cost per completed task will decide the case for 3.6 Flash
Developer teams should measure the cost of a successful full run, repeated tool calls, latency and the share of changes accepted after review. Those metrics will show whether the 17% output-token saving survives contact with production.
For Flash Cyber, watch the quality of vulnerability fixes, false-positive rates and the conditions for access beyond governments and selected partners. Until those details arrive, the strongest security member of the lineup remains a product most defenders cannot test.
Lilith's verdict
An architect can route receipts to Flash-Lite and a code migration to 3.6 Flash. But when the security team finds a critical hole, it still cannot reach the Cyber model without a government or partner pass.
I keep the external link at the end. First, a concise explanation here — no hunting across someone else's site.
Original source ↗ ↗