Model launches
Google's three-model Gemini drop: a cheaper flash, a faster lite, and a cyber variant locked to governments
Google's three-model Gemini drop targets production AI with better token efficiency and lower latency. 3.6 Flash cuts token usage by 17% versus its predecessor. 3.5 Flash-Lite runs at 350 output tokens per second. A cyber-focused variant ships exclusively to vetted partners.
Emmanuel Fabrice Omgbwa Yasse AI-assisted
2026-07-22 · 3 min read

Google deepened its Gemini lineup on July 21 with three models aimed at different parts of the production AI stack: a cheaper, more efficient workhorse, a high-speed lite model, and a cybersecurity variant locked to government and partner use.
The marquee release is Gemini 3.6 Flash, which Google positions as a direct upgrade to 3.5 Flash for coding, knowledge work, and multimodal tasks. According to the Artificial Analysis Index, which independently tracks model performance, 3.6 Flash uses 17% fewer output tokens than 3.5 Flash. On the DeepSWE benchmark by Datacurve, Google says the reduction reaches 65% in some scenarios. The model is priced at $1.50 per million input tokens and $7.50 per million output tokens, a cut from 3.5 Flash’s pricing. For comparison, OpenCode is now offering both Gemini 3.6 Flash and the lite model at an 80% cost reduction for AI coding tasks, per OpenCode's pricing update.

Google reports benchmark gains across several categories. On DeepSWE, the model scored 49% versus 37% for 3.5 Flash, meaning fewer unwanted code edits and shorter execution loops. On MLE Bench, which measures machine learning research tasks, 3.6 Flash scored 63.9% against 49.7%. Computer use capabilities improved on OSWorld-Verified (83.0% vs. 78.4%). On the knowledge-work benchmark GDPval-AA v2, the model scored 1421 vs. 1349. The agentic performance gains mirror a broader trend in which coding agents are getting better at judging their own outputs, as recent research on evaluation skills suggests.
The second release, Gemini 3.5 Flash-Lite, is built for high throughput and low latency. Google says it runs at 350 output tokens per second, making it the fastest model in the 3.5 family. Priced at $0.30 per million input tokens and $2.50 per million output tokens, it undercuts 3.1 Flash-Lite and outperforms its predecessor across agentic and coding evals. On Terminal-Bench 2.1, the model scored 54% vs. 31% for 3.1 Flash-Lite. On SWE-Bench Pro and OSWorld-Verified, it beat the older Gemini 3 Flash: 54.2% vs. 49.6% and 74.0% vs. 65.1%, respectively. That speed advantage is notable because latency is often the hidden cost in agent loops, as speculative decoding research has shown.
The third model, Gemini 3.5 Flash Cyber, takes a different route to market. Built on top of 3.5 Flash and fine-tuned for security vulnerability detection and patching, it is not a general-purpose release. Google says it will be available only to governments and trusted partners through CodeMender, the company's code security agent. Within CodeMender, multiple 3.5 Flash Cyber agents work together to produce a single vulnerability report. Google claims it reaches competitive frontier performance on the CyberGym benchmark, though it did not publish a specific score. The restricted access mirrors a growing pattern in cybersecurity AI; Sakana AI's Fugu-Cyber, for instance, also matched frontier models but urged that raw models alone aren't the answer, according to their announcement.
“Given the dual-use nature of this technology, we have taken an intentional approach to deploying 3.5 Flash Cyber,” Google wrote in the announcement. The limited-access pilot is meant to give frontline defenders an advantage while restricting broader misuse. Similarly, the open-source pentesting agent VulnClaw already automates the full vulnerability workflow, with a kill switch built in, as VulnClaw's documentation details.
All three models arrive as Google continues to fill gaps in its Gemini naming. The company never shipped a 3.5 Pro, which was expected as a mid-range workhorse. In today’s announcement, Google said Gemini 3.5 Pro is “currently testing with partners” and will be made broadly available when ready. The company also revealed it has started “our most ambitious pre-training run yet, for Gemini 4.” A leaked test of a Gemini 3.6 flash model had already hinted at this version jump, as the earlier leak observed.
3.6 Flash and 3.5 Flash-Lite are available starting July 21 in the Gemini API through Google AI Studio and Android Studio. Enterprise users can access them via the Gemini Enterprise Agent Platform. 3.6 Flash also shows up in the Gemini app. Google said 3.5 Flash-Lite is rolling out in Google Search as well.
Get the tech essentials in 3 minutes every morning
One email, every weekday, with what actually matters in AI and tech.