Skip to content

Google Drops Three Gemini Models, Teases Gemini 4

5 min read

Google Drops Three Gemini Models, Teases Gemini 4
Photo by Google DeepMind on Pexels

Three Models at Once, One Clear Winner

On Tuesday, July 21, Google shipped three new Gemini models simultaneously: Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber. For most developers, the headline is Gemini 3.6 Flash — it scores 49% on the DeepSWE coding benchmark (up from 37% for 3.5 Flash) and reduces output token usage by 17%, which translates directly to lower inference costs at scale.

The release confirms where Google is concentrating its near-term engineering effort: smaller, faster models optimized for agentic workflows, not the frontier flagship that’s been promised for months. Gemini 3.5 Pro is still absent, having missed its window multiple times, while a trio of mid-tier updates ship instead. That’s a deliberate prioritization, and for production workloads, it’s arguably the right one.

Gemini 3.6 Flash: Better Code, Fewer Tokens

Gemini 3.6 Flash is the main event. On the DeepSWE benchmark — which measures how reliably a model resolves real-world GitHub issues — it scores 49%, up from 37% for its predecessor. On MLE-Bench, which tests machine learning engineering tasks, it climbs from 49.7% to 63.9%. On OSWorld-Verified, the computer use accuracy benchmark, it hits 83.0%, up from 78.4%.

The token efficiency improvement is the bigger deal for anyone running production agents. Google reports 17% fewer output tokens versus 3.5 Flash on average, with some agentic workloads seeing up to 65% reduction on long-horizon engineering tasks. The model “takes fewer reasoning steps and tool calls to accomplish multi-step workflows” — this is less about compression and more about the model having learned to stop padding its responses with unnecessary elaboration.

Pricing reflects the efficiency gain: output drops from $9.00 to $7.50 per million tokens, with input held at $1.50/MTok. The model ships with a 1-million-token input context window and 64K max output, with thinking controls and native computer use available via the Gemini API and Gemini Enterprise.

One practical improvement worth flagging: the knowledge cutoff advances from January 2025 to March 2026. That 14-month jump matters for any application where recent events affect responses, and it’s a gap that had become noticeable in deployments using 3.5 Flash against current events or recent documentation.

Gemini 3.5 Flash-Lite: The Volume Play

Gemini 3.5 Flash-Lite targets high-throughput, latency-sensitive workloads — the use cases where you’re generating millions of short completions and cannot absorb 3.6 Flash pricing. Artificial Analysis benchmarks it at 350 output tokens per second, making it the fastest Gemini model currently in production.

Despite the “Lite” label, the benchmarks are more than capable for most tasks. It scores 54.2% on SWE-Bench Pro (versus 49.6% for the older 3 Flash), 74.0% on OSWorld-Verified (up from 65.1%), and 72.2% on GDM-MRCR v2, a long-context retrieval and reasoning test. On Terminal-Bench 2.1, which measures the ability to complete real terminal-level tasks autonomously, it jumps from 31% to 54%.

Flash-Lite shares the 1M token context window and 64K output limit with 3.6 Flash, and ships with thinking controls and built-in computer use. For high-volume agentic pipelines where cost-per-action matters more than raw reasoning depth, it’s the more interesting launch of the two.

Gemini 3.5 Flash Cyber: Vulnerability Hunting Under Restriction

The most unusual part of Tuesday’s announcement is Gemini 3.5 Flash Cyber — a security-tuned model integrated into Google’s CodeMender agent, purpose-built to find and patch code vulnerabilities at scale.

The benchmark numbers are concrete. When tested against the V8 JavaScript engine — used in Chrome, Node.js, and Deno — Gemini 3.5 Flash Cyber found 55 unique confirmed vulnerabilities. The base 3.5 Flash found 47. Claude Opus 4.6 found 36. The cyber model identified 10 issues that no other model caught at all.

Google is not making Flash Cyber generally available. Access is restricted to governments and trusted partners through a limited-access pilot via CodeMender, with human approval required before any patches ship. The reasoning is explicit: a model that finds vulnerabilities faster than defenders can patch them is also a model that could accelerate attacks in the wrong hands.

That tradeoff is worth examining carefully. Google chose to build and restrict rather than withhold — and they chose to publish the benchmark numbers. It signals that security-specialized AI has reached a capability level where release policy is no longer a theoretical question. Whether competitors follow a similar restricted-access model, or opt for wider release, will be an interesting policy divergence to watch over the next year.

What’s Still Missing — and What’s Coming

The conspicuous absence from Tuesday’s release is Gemini 3.5 Pro. Google’s flagship model in the 3.5 line has now missed its announced window several times. Google confirmed it remains in development, with no release date given. That matters competitively: both GPT-5.6 Sol and Claude Sonnet 5 are in the market and outperform the current Gemini Flash tier on sustained multi-step reasoning benchmarks.

The Flash tier’s strengths — throughput, cost, computer use, and agentic task completion — are real and meaningfully improved with this release. But the gap at the frontier remains, and developers who need peak reasoning performance still have better options elsewhere.

On the forward-looking side, Google confirmed on July 21 that pre-training for Gemini 4 has begun. No capabilities were announced and no date was given — pre-training runs take months before the model is ready for evaluation, let alone release. But the signal confirms the roadmap is progressing. As we noted when Gemini 3.5 Flash launched, Google’s explicit bet is that agentic throughput and cost-per-call matter more for most production use cases than raw benchmark scores. The 3.6 Flash and Flash-Lite numbers suggest that bet continues to compound — even if the frontier Pro model is taking longer than expected to arrive.

Further Reading

Don’t miss on Ai tips!

We don’t spam! We are not selling your data. Read our privacy policy for more info.

Don’t miss on Ai tips!

We don’t spam! We are not selling your data. Read our privacy policy for more info.

Enjoyed this? Get one AI insight per day.

Join engineers and decision-makers who start their morning with vortx.ch. No fluff, no hype — just what matters in AI.