Chuyển đến nội dung chính

Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber: Google's New Models

Everything you need to know about Google's newest Gemini models — release dates, pricing, benchmarks, and how 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber each fit into your workflow.
21 thg 7, 2026  · 9 phút đọc

Khám phá với AI

Mở trong ChatGPTMở trong ClaudeMở trong Perplexity

Google released three Gemini models today with the same underlying goal: do more with fewer tokens/

  • 3.6 Flash
  • 3.5 Flash-Lite, and
  • 3.5 Flash Cyber. 

3.6 Flash is the general-purpose upgrade. Flash-Lite trades quality for speed and price. Cyber is a narrow security tool that most people won't get access to.

Below: what each model actually does, how the benchmarks hold up, and where Gemini 3.5 Pro fits into all this. 

Choosing Between the New Gemini Models

Each one sits at a different point on the cost-to-capability curve. If you are choosing based on the benchmark results, skip ahead; the benchmarks are further down.

What is Gemini 3.6 Flash?

Gemini 3.6 Flash is the release's centerpiece. It is the model Google wants you to default to for coding, knowledge work, and anything multimodal. 

Per the Artificial Analysis Index, 3.6 Flash cuts output token usage by 17% versus 3.5 Flash, and Google reports reductions as high as 65% on individual evals like DeepSWE. Some of that is a smarter model saying less. The rest comes from reaching a conclusion in fewer reasoning steps and tool calls, which is a different kind of gain and the harder one to fake.

3.6 Flash takes text, image, video, audio, and PDF as input, runs a 1-million-token context window with a 64k output cap. Computer use is now built into both the Gemini API and Gemini Enterprise, alongside function calling, structured output, and search-as-a-tool. (Although this maybe is not unique to 3.6 Flash anymore; most frontier labs ship some version of computer use by now.) 

What about Gemini 3.6 Flash vs. Gemini 3.5 Flash? 

Against its own predecessor, the jump is real almost everywhere: DeepSWE nearly doubles, MLE-Bench climbs from 49.7% to 63.9%, long-context performance on GDM-MRCR v2 improves by double digits.

Gemini 3.6 Flash vs Gemini 3.5 Flash Benchmarks

What is Gemini 3.5 Flash-Lite?

3.5 Flash-Lite isn't trying to be the smartest model in the lineup. It's the fastest, at 350 output tokens per second, and the cheapest, at $0.30 per million input tokens and $2.50 per million output tokens.

If your workload is high-volume and latency-sensitive — think like agentic search, bulk document processing, anything running thousands of times a day — the quality gap other models offer probably isn't worth what it costs to close it. So that's the bet Flash-Lite is making.

Developers can dial the thinking level up or down by job: minimal or low for cheap, fast, high-volume execution, or higher for multi-step subagent work that actually needs to reason through something.

What about Gemini 3.5 Flash-Lite vs. Gemini 3.1 Flash-Lite? 

The comparison against its own predecessor, 3.1 Flash-Lite, isn't close: Terminal-Bench 2.1 goes from 31% to 54%, GDM-MRCR v2 from 60.1% to 72.2%, GDPVal-AA v2 from 642 to 1140.

Benchmark Gemini 3.5 Flash-Lite Gemini 3.1 Flash-Lite
Terminal-Bench 2.1 54% 31%
GDM-MRCR v2 (long context) 72.2% 60.1%
GDPVal-AA v2 (Elo) 1140 642

The more interesting comparison, honestly, is upward. On several agentic and coding evals, 3.5 Flash-Lite now beats the older, larger Gemini 3 Flash outright — SWE-Bench Pro (54.2% vs. 49.6%), OSWorld-Verified (74.0% vs. 65.1%). If that holds up on your own workload, it's a straightforward cost cut with no quality tax attached.

What is Gemini 3.5 Flash Cyber?

Finally, there's Cyber, which most people reading this will never get to touch. It's 3.5 Flash fine-tuned specifically to find and patch security vulnerabilities, deployed inside Google's CodeMender system, where several Cyber-model agents work together and hand back one combined report.

Google says it's competitive with frontier-scale models on CyberGym, a widely used cybersecurity benchmark, despite its smaller footprint — which, if it holds up, is arguably the more interesting efficiency story in this whole release.

Access is limited to governments and trusted partners through CodeMender's pilot program. No public pricing, no general release date. 

Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber Benchmarks

Here is a summary of the benchmark results. I included the pricing information also, which is important piece of the puzzle when choosing.

Benchmark Gemini 3.6 Flash Gemini 3.5 Flash Gemini 3.1 Pro GPT-5.6 Luna Grok 4.5 Claude Sonnet 5
Input price ($/1M) $1.50 $1.50 $2.00 $1.00 $2.00 $3.00
Output price ($/1M) $7.50 $9.00 $12.00 $6.00 $6.00 $15.00
SWE-Bench Pro 58.7% 55.1% 54.2% 62.7% 64.7% 63.2%
DeepSWE v1.1 49% 37% 12% 67% 54% 54%
Terminal-Bench 2.1 78.0% 76.2% 73.8% 84.7% 83.3% 80.4%
MLE-Bench 63.9% 49.7% 42.6% 47.6% 43.2% 66.9%
OSWorld-Verified 83.0% 78.4% 76.2% 72.6% 81.2%
GDPVal-AA v2 (Elo) 1421 1349 965 1584 1535 1607
CharXiv Reasoning (no tools) 85.2% 84.2% 83.3% 82.7% 81.6% 77.0%
GDM-MRCR v2, 128k 91.8% 77.3% 84.9% 74.8% 81.4% 71.6%

Against the outside competition it's a split decision. GPT-5.6 Luna and Grok 4.5 lead on the coding-heavy benchmarks — SWE-Bench Pro, DeepSWE — and that's a real gap if your workload is genuinely coding-heavy rather than coding-adjacent.

Claude Sonnet 5 leads on knowledge work and MLE-Bench.

Where 3.6 Flash actually wins is long-context retrieval and chart-based reasoning, by a wide enough margin that it's not close, and it undercuts every model on that list except GPT-5.6 Luna and Grok 4.5 on output price. 

Release Date, Availability, and Pricing

3.6 Flash and 3.5 Flash-Lite are live now — Gemini API, Google AI Studio, Android Studio, and for 3.6 Flash specifically, Google Antigravity too. Enterprise customers get both through the Gemini Enterprise Agent Platform, with 3.6 Flash also in the Gemini Enterprise app. On the consumer side, both are in the Gemini app, and Flash-Lite is starting to show up in Google Search's AI Mode.

Pricing, if you skipped ahead for it:

  • 3.6 Flash: $1.50/1M input tokens, $7.50/1M output tokens
  • 3.5 Flash-Lite: $0.30/1M input tokens, $2.50/1M output tokens
  • 3.5 Flash Cyber: not publicly priced — limited to the CodeMender pilot program

Gemini 3.6 Flash Safety: Frontier Safeguards and Jailbreak Resistance

3.6 Flash ships with what Google calls enhanced Frontier Safety safeguards around CBRN and cyber-offense misuse — the standard language for "we tried to make jailbreaking harder without making the model annoying to use," a balance every lab claims to have struck but commentary on X always tells a different story.

Cyber's restricted rollout is really the same instinct applied differently. Rather than try to safety-tune a vulnerability-hunting model for general release, Google just didn't release it generally.

What's Next for Gemini: 3.5 Pro and Gemini 4 Release Timeline

3.5 Pro is still in partner testing, with Google saying general availability is coming once it's ready.

More notably, Google confirmed it's begun pretraining Gemini 4, calling it its most ambitious pretraining run so far. To keep up-to-date with the latest AI news, subscribe to our newsletter, The Median, and we will keep you informed about important model releses such as this. 

Conclusion

None of these three models is trying to top every leaderboard, and that's probably the right call. 3.6 Flash gives up ground to GPT-5.6 Luna and Grok 4.5 on coding benchmarks, and to Claude Sonnet 5 on knowledge work, in exchange for a real efficiency jump over its own predecessor and a clear lead on long-context retrieval.

3.5 Flash-Lite doesn't even try to compete on quality. It's betting that for a huge share of agentic workloads, quality past a certain point stops mattering as much as speed and cost do. That's probably true. It's also the kind of bet that looks obviously correct right up until an agent makes a wrong call that a smarter model wouldn't have.

Cyber isn't on a leaderboard at all. It's a narrow tool built for one job, gated to the people Google trusts to use it responsibly — which says as much about how seriously Google is treating the vulnerability-discovery risk as anything in the safety section above.

Put together, this reads less like three separate launches and more like Google tuning the entire Flash tier for agentic workloads at scale, while it keeps the actually ambitious models — 3.5 Pro, Gemini 4 — in reserve for whenever they're ready.


Josef Waples's photo
Author
Josef Waples

I'm a data science writer and editor with contributions to research articles in scientific journals. I'm especially interested in linear algebra, statistics, R, and the like. I also play a fair amount of chess! 

FAQs

What are the three new models?

Gemini 3.6 Flash (general-purpose workhorse), Gemini 3.5 Flash-Lite (fast, cheap, high-throughput), and Gemini 3.5 Flash Cyber (cybersecurity specialist for Google's CodeMender system).

How much more efficient is 3.6 Flash than 3.5 Flash?

A 17% cut in output token usage on the Artificial Analysis Index, per Google, with individual benchmarks like DeepSWE showing reductions of up to 65%.

What's the context window on 3.6 Flash?

1 million input tokens, 64,000-token output cap. Text, image, video, audio, and PDF go in; text comes out.

How fast is 3.5 Flash-Lite?

350 output tokens per second, per Artificial Analysis — the fastest model in the 3.5 family.

Does 3.5 Flash-Lite outperform larger models?

On some benchmarks, yes. It beats the older, larger Gemini 3 Flash on SWE-Bench Pro and OSWorld-Verified specifically — that doesn't mean it wins everywhere.

Can I access 3.5 Flash Cyber?

Not broadly. It's limited to a pilot for governments and trusted partners through CodeMender, because a model built to find vulnerabilities is also one that could be misused to find them.

How does 3.6 Flash compare to GPT-5.6 Luna and Grok 4.5?

Both lead on coding-specific benchmarks like SWE-Bench Pro and DeepSWE. 3.6 Flash leads on long-context retrieval and chart-based reasoning instead.

Is Gemini 3.5 Pro out yet?

No. Still in partner testing, general availability promised once it's ready.

Chủ đề

Learn with DataCamp

Courses

Triển khai Giải pháp AI trong Doanh nghiệp

2 giờ
52.9K
Khám phá cách khai thác giá trị kinh doanh từ AI. Học cách xác định cơ hội cho AI, tạo POC, triển khai giải pháp và xây dựng chiến lược AI.
Xem chi tiếtRight Arrow
Bắt Đầu Khóa Học
Xem thêmRight Arrow
Có liên quan

blogs

Gemini 3.5 Flash: Google's Fastest Agentic Model

Google launched Gemini 3.5 Flash at I/O 2026, a model that outperforms Gemini 3.1 Pro on agentic and coding benchmarks while running four times faster than competitors.
Matt Crabtree's photo

Matt Crabtree

8 phút

blogs

Gemini 3.5 Flash vs Claude Opus 4.7: The Sprinter and the Surgeon

Google's speed-optimized Flash model takes on Anthropic's deep-coding flagship across agentic workflows, reasoning, multimodal tasks, and pricing.
Tom Farnschläder's photo

Tom Farnschläder

12 phút

blogs

Gemini 3.5 Flash vs GPT-5.5: The Multitool and the Sledgehammer

One model is built for versatile tool-calling at scale; the other brute-forces the hardest reasoning problems. Compare Google's Gemini 3.5 Flash and OpenAI's GPT-5.5 across coding, agentic workflows, multimodal tasks, and pricing.
Tom Farnschläder's photo

Tom Farnschläder

11 phút

blogs

Gemini 3.1: Features, Benchmarks, Hands-On Tests, and More

Learn about Gemini 3.1 Pro, Google's latest reasoning model. Explore its features, benchmarks, hands-on tests, and how it compares to Claude Opus 4.6, Claude Sonnet 4.6, and GPT-5.2.
Khalid Abdelaty's photo

Khalid Abdelaty

11 phút

gemini 2.5 pro with a large context

blogs

Gemini 2.5 Pro: Features, Tests, Access, Benchmarks, and More

Explore Google's Gemini 2.5 Pro, and learn about its impressive 1 million token context window, multimodal capabilities, hands-on test results, and how to access it.
Alex Olteanu's photo

Alex Olteanu

8 phút

blogs

Gemini 3: Google’s Most Powerful LLM

Learn about Gemini 3 Pro, Google’s latest and most powerful LLM, which is topping benchmarks across the board. Plus, discover Gemini 3 Deep Think mode and Google Antigravity.
Matt Crabtree's photo

Matt Crabtree

13 phút

Xem ThêmXem Thêm