Google has released Gemini 3.6 Flash, alongside two sibling models, expanding its lineup of efficiency-focused AI aimed at developers running agentic workloads at scale.

The update, announced July 21 by Google product lead Tulsee Doshi, adds three models to the Gemini family: Gemini 3.6 Flash, a faster general-purpose model; Gemini 3.5 Flash-Lite, a stripped-down option for high-volume tasks; and Gemini 3.5 Flash Cyber, a security-focused variant restricted to government use.

Cheaper tokens, higher throughput

Google says Gemini 3.6 Flash cuts output-token usage by 17% compared with its predecessor, according to the Artificial Analysis Index, with reductions of up to 65% on some benchmarks. It scored 83.0% on OSWorld-Verified, a computer-use benchmark, and 63.9% on MLE-Bench. Pricing is set at $1.50 per million input tokens and $7.50 per million output tokens, and the model is live now across the Gemini API, Google AI Studio, Android Studio, Google Antigravity, Gemini Enterprise, and the Gemini app.

Gemini 3.5 Flash-Lite targets even higher-volume, lower-cost workloads, processing 350 output tokens per second and scoring 54% on Terminal-Bench 2.1. It is rolling out through the same developer surfaces and into Google Search.

A gated model for finding bugs

The third model, Gemini 3.5 Flash Cyber, is tuned specifically to detect and patch software vulnerabilities at a lower token cost. Unlike the other two, it won’t reach the public: Google says it will be available only to governments and “trusted partners” through a limited CodeMender pilot program, echoing OpenAI’s own campaign to patch open-source vulnerabilities with AI.

The launch leaves Gemini 3.5 Pro, Google’s flagship model that has already slipped past its original release window, still in partner testing. Google also disclosed it has begun what it called its “most ambitious pre-training run yet” for Gemini 4, without giving a timeline.