Lilith Lilith.
Editorial illustration: Gemini 3.8 Flash deploys a Cyber agent for vulnerabilities and code
Lilith illustration · editorial remix

A model optimized for code and security

Google released Gemini 3.8 Flash just six weeks after version 3.7. The new model primarily targets long-horizon software engineering and the capabilities of autonomous agents. According to tests on the DeepSWE benchmark, 3.8 Flash performs better in coding and multi-step reasoning than its predecessor and beats some significantly larger models, although it lags behind Claude Opus in general agent environment use like OSWorld-2.0.

Alongside it, Google launched a specifically tuned version: Gemini 3.8 Flash Cyber. It achieved a score of 86.2% on the CyberGym benchmark for vulnerability detection and can automatically generate patches with a success rate comparable to frontier models, but at a fraction of the cost.

Code won't just be written, it will be automatically protected

While the standard Flash model is publicly available as a lightweight but powerful tool for developers, the Cyber model is exclusive. It is hidden behind the Fairwind program, into which Google only allows vetted government authorities, critical infrastructure operators, and key software maintainers. The security risk of misusing the model's offensive capabilities to automatically find holes convinced Google to guard it strictly.

When Google tested the Cyber variant internally on twenty programming languages, it achieved over a 70% success rate in detecting bugs and found a critical weakness in its own cloud in two hours instead of the usual months. This gives defenders an asymmetrical advantage against attackers.

Extensive capabilities require more patience

The new capabilities manifest in token consumption and overall response time for the most complex tasks. The models use more internal reasoning steps and iterative tool calls when solving complex problems. Those expecting instant answers to simple queries might not notice the improvement.

Maintaining low costs ($0.75 per million input tokens) remains an advantage, but only until the end of 2026. As analysis on TechCrunch and Ars Technica shows, Google is optimizing margins by massively promoting the Flash variant at the expense of the large models of the Pro family, which it is now innovating more slowly.

Adoption in legacy infrastructure will determine success

How the model handles modern benchmarks and isolated tests is already known. The question remains how the Cyber model will cope with deployment into real production code, full of historical debt and proprietary corporate libraries. Test results show an increase, but only full deployment will reveal whether agents can not only write patches but also prevent system crashes during their automatic integration.

Lilith's verdict

Patching a critical bug takes hours instead of weeks, but an exclusive gatekeeper watches over Cyber.

I keep the external link at the end. First, a concise explanation here — no hunting across someone else's site.

Original source ↗