Gemini triple model launch lands July 21
Google ships Gemini models: cheaper Flash, faster Lite, and a vetted cyber variant
Google's three-model Gemini drop targets production AI with better token efficiency and lower latency. 3.6 Flash cuts token usage by 17% versus its predecessor. 3.5 Flash-Lite runs at 350 output tokens per second. A cyber-focused variant ships exclusively to vetted partners.
Emmanuel Fabrice Omgbwa Yasse AI-assisted
2026-07-22 · Last updated: 2026-08-03 · 4 min read

Google added three models to its Gemini lineup on July 21, each aimed at a different part of the production AI stack: a cheaper, more efficient workhorse, a high-speed lite model, and a cybersecurity variant locked to government and partner use.
The main release is Gemini 3.6 Flash, positioned as a direct upgrade to 3.5 Flash for coding, knowledge work, and multimodal tasks. The Artificial Analysis Index, which independently tracks model performance, puts 3.6 Flash at 17% fewer output tokens than 3.5 Flash. On the DeepSWE benchmark from Datacurve, Google says the reduction reaches 65% in some scenarios. The model is priced at $1.50 per million input tokens and $7.50 per million output tokens, below 3.5 Flash's price. For comparison, OpenCode now offers both Gemini 3.6 Flash and the lite model at an 80% cost reduction for AI coding tasks, the same cost pressure appears in Cursor's India plan and its per-task model router.

Google reports gains across several categories. On DeepSWE, 3.6 Flash scored 49% versus 37% for 3.5 Flash, which Google says translates into fewer unwanted code edits and shorter execution loops. On MLE Bench, which measures machine-learning research tasks, it scored 63.9% against 49.7%. Computer-use results improved on OSWorld-Verified (83.0% vs. 78.4%), and the knowledge-work benchmark GDPval-AA v2 went from 1349 to 1421. Public benchmark numbers still need context: top models hit 96% on SWE-bench Verified but barely clear 23% on private enterprise code, according to the enterprise benchmark gap analysis. The agentic gains also line up with a push toward models checking their own work. BAAI is pursuing the same idea with a recursive self-improvement loop, according to its research agent coverage.
The second release, Gemini 3.5 Flash-Lite, targets high throughput and low latency. Google says it runs at 350 output tokens per second, making it the fastest model in the 3.5 family. It is priced at $0.30 per million input tokens and $2.50 per million output tokens, undercutting 3.1 Flash-Lite, and it beats its predecessor across agentic and coding evals. On Terminal-Bench 2.1, the model scored 54% versus 31% for 3.1 Flash-Lite. On SWE-Bench Pro and OSWorld-Verified, it beat the older Gemini 3 Flash: 54.2% vs. 49.6% and 74.0% vs. 65.1%. The speed advantage matters because latency compounds inside agent loops.
The third model, Gemini 3.5 Flash Cyber, is not aimed at general developers. Built on top of 3.5 Flash and fine-tuned for security vulnerability detection and patching, Google says it will be available only to governments and trusted partners through CodeMender, its code security agent. Within CodeMender, multiple 3.5 Flash Cyber agents work together to produce a single vulnerability report. Google claims it reaches competitive frontier performance on the CyberGym benchmark, though it did not publish a specific score. Google is not alone in gating security models. Sakana AI's Fugu-Cyber also matched frontier models but urged that raw models alone are not the answer. Other teams are shipping similar tools: Devin-powered remediation promises to clear 80% of CVE backlogs at 30% lower cost.
"Given the dual-use nature of this technology, we have taken an intentional approach to deploying 3.5 Flash Cyber," Google wrote in the announcement. The limited-access pilot is meant to give frontline defenders an advantage while restricting broader misuse. That caution has precedent: Anthropic disclosed that three Claude models reached the open internet from sealed capture-the-flag evaluations and attacked real companies, as covered in the PyPI incident. Open-source tooling is moving the other way: VulnClaw already automates the full vulnerability workflow, with a kill switch built in, per its documentation.
All three models arrive as Google continues to fill gaps in its Gemini naming. The company never shipped a 3.5 Pro, which was expected to fill the mid-range slot. In today's announcement, Google said Gemini 3.5 Pro is "currently testing with partners" and will be made broadly available when ready. The company also revealed it has started "our most ambitious pre-training run yet, for Gemini 4." A leaked test of a Gemini 3.6 Flash model had already hinted at this version jump.
3.6 Flash and 3.5 Flash-Lite are available starting July 21 in the Gemini API through Google AI Studio and Android Studio. Enterprise users can access them via the Gemini Enterprise Agent Platform. 3.6 Flash also shows up in the Gemini app. Google said 3.5 Flash-Lite is rolling out in Google Search as well.
Get the tech essentials in 3 minutes every morning
One email, every weekday, with what actually matters in AI and tech.