
Google released three new Gemini models that run faster and cost less
Google released three new Gemini models at once, and all of them use fewer resources to do the same work. Let's break down how 3.6 Flash, 3.5 Flash-Lite, and the special cybersecurity model differ, and which one fits your tasks.
What are these new Gemini models and why three at once?
Google introduced three new models in the Gemini Flash lineup at once. The logic here is simple, anyone building AI agents for real production work needs models that aren't just smart, but also use fewer resources, respond faster, and behave more predictably. The Flash lineup was built exactly for that balance between quality and efficiency.
The first model, 3.6 Flash, is the main workhorse. It handles coding, knowledge work, and tasks involving images and documents better than before. According to the independent Artificial Analysis ranking, the model uses 17 percent fewer output tokens than the previous 3.5 Flash, and on some benchmarks the savings reach up to 65 percent.
The second model, 3.5 Flash-Lite, is the fastest and most affordable in its generation. It delivers 350 output tokens per second and clearly outperforms earlier models of its kind in agentic workflows. The third model, 3.5 Flash Cyber, is a specialized version for finding and fixing code vulnerabilities, more on that below.
How Gemini 3.6 Flash improves on the previous version
Gemini 3.6 Flash was built directly on feedback from developers working with the 3.5 Flash version. The model didn't just get stronger at coding and knowledge work, it does that while spending noticeably fewer tokens. It needs fewer reasoning steps and fewer calls to outside tools to finish a multi-step task.
The price also dropped compared to the previous version. Right now, 3.6 Flash costs 1.5 dollars per million input tokens and 7.5 dollars per million output tokens. Together, that means the same task gets cheaper for two reasons at once, the per-token price is lower and fewer tokens get used in the first place.
3.6 Flash shows better token efficiency and reduced verbosity than 3.5 Flash in an OSWorld verified task (API)
Here's what the specific improvements over 3.5 Flash look like,
In coding, the model makes fewer unnecessary code edits and gets stuck in loops less often, the DeepSWE score rose from 37 to 49 percent
In machine learning research tasks, the jump is even bigger, from 49.7 to 63.9 percent
In computer control, the score rose from 78.4 to 83 percent, and this feature is now built directly into the Gemini API and the enterprise platform
In knowledge and analytical work, the model is stronger too, 1421 points against the previous version's 1349
Customers like Hebbia and Harvey note the model is especially good at parsing documents, analyzing charts, and drafting reports. So 3.6 Flash suits not only developers, but anyone working through large volumes of documents and data every day.
3.6 Flash executes code migrations, using multi-agent orchestration on AGY, with lower latency and higher quality than 3.5 Flash (AGY)

Customers report 3.6 Flash is a step forward in both cost and quality, balancing token efficiency, accuracy, and speed across complex workflows and knowledge-based tasks
How safety works in Gemini 3.6 Flash
Alongside the model, Google shipped strengthened safeguards in two sensitive areas, chemical, biological, radiological, and nuclear topics, and cyber attacks. These safeguards make the model noticeably harder to trick into bypassing its limits.
At the same time, Google specifically points out that the model was trained to refuse legitimate, useful requests less often. In other words, the protection got stricter exactly where real risk exists, rather than everywhere across the board.
What Gemini 3.5 Flash-Lite does and who it suits
Alongside the main model, Google released Gemini 3.5 Flash-Lite. It's built for tasks where response speed and the ability to handle large volumes of requests matter most, like agentic search and processing documents in bulk.
This is the fastest model in its generation, delivering 350 output tokens per second. The price is 0.3 dollars per million input tokens and 2.5 dollars per million output tokens, with noticeably better quality than the previous model of its type. That combination makes it a strong-value option for anyone pushing a heavy stream of requests through an AI constantly.
3.5 Flash-Lite executes high volume tasks at a lower latency than 3.5 Flash
The model has a handy setting for how deeply it thinks. You can set it to minimal, and it responds as fast and cheaply as possible on repetitive, high-volume tasks. Or you can turn up the level when a task needs several steps and additional AI subagents. Computer control is built directly into the model too.
Here's how the quality improved over the previous 3.1 Flash-Lite,
Command line work rose from 31 to 54 percent
Long-text handling rose from 60.1 to 72.2 percent
Real-world task execution rose from 642 to 1140 points

On many agentic and coding evals, 3.5 Flash-Lite even outperforms 3 Flash, including on SWE-Bench Pro (54.2% vs. 49.6%) and OSWorld-Verified (74.0% vs. 65.1%), making it a faster and more capable option for workloads on both 2.5 and 3 Flash

That creates an interesting situation, the cheaper model beats older, higher-tier models from the previous generation on a number of tasks. For anyone still running on 2.5 or 3 Flash, that's a direct reason to reconsider what they're using and what they're paying for.
What the 3.5 Flash Cyber model is for
AI models have gotten better at finding software vulnerabilities faster than existing systems can patch them. That's exactly why Google released a separate model called 3.5 Flash Cyber, built on top of the standard 3.5 Flash but specifically fine-tuned for finding and fixing code vulnerabilities.
The model works inside a tool called CodeMender, where several of these AI agents run in parallel and combine their findings into a single report. According to the widely used CyberGym benchmark, this setup performs at the level of the strongest solutions on the market while costing less per token than larger models.

Given that this kind of technology can be misused, Google took a cautious approach to releasing it. The model will be available only to governments and trusted partners through CodeMender as part of a limited pilot program. The point is to give defenders a head start in finding and patching critical vulnerabilities before attackers can exploit them.
Where to try the new Gemini models right now
The 3.6 Flash and 3.5 Flash-Lite models are available starting today. Developers can work with them through the Gemini API in Google AI Studio and Android Studio, and 3.6 Flash is also available in Google Antigravity. Enterprise customers get access through the Gemini Enterprise Agent Platform, and 3.6 Flash also runs in the Gemini Enterprise app.
Regular users can try the new models directly in the Gemini app, and 3.5 Flash-Lite is gradually rolling out in Google Search too. Separately, Google announced that Gemini 3.5 Pro is currently in testing with partners and will launch broadly as soon as it's ready.
And the most interesting part for the future, Google's team has already begun its most ambitious pre-training run ever, for the next generation, Gemini 4.
Three Models, One Takeaway
The main takeaway from this launch is simple, the AI race has shifted from "who's smarter" to "who does the same work cheaper and faster." Gemini 3.6 Flash uses fewer tokens and costs less than its predecessor, while Flash-Lite beats older, higher-tier models on a range of tasks at a noticeably lower price.
For an everyday user, that means choosing a model should now come down to the specific task rather than the loudest name. Some jobs need speed and volume, others need accuracy in parsing documents, others need solid coding. The same task can cost several times less simply by picking the right model instead of the most hyped one.
The problem is, testing all these models separately, dealing with Google AI Studio access, separate subscriptions, and payments, takes far too long. It's much more practical when every current AI model sits in one place and you can compare them directly on your own real task.
That's exactly what unitool.ai is for. Get one subscription and unlock access to the newest models at once, including Google's latest releases, no foreign cards and no separate sign-up on each platform. See for yourself which model turns out more accurate and more cost-effective for your specific tasks.