Published: July 22, 2026
Google released Gemini 3.6 Flash on July 21, 2026, the new default workhorse model in the Gemini family and the successor to Gemini 3.5 Flash. In the same announcement Google DeepMind also launched Gemini 3.5 Flash-Lite and Gemini 3.5 Flash Cyber. Gemini 3.6 Flash keeps the 1 million-token context window, moves the knowledge cutoff forward to March 2026, and uses about 17 percent fewer output tokens than 3.5 Flash while scoring higher on coding, long-context and computer-use benchmarks.
What did Google release today?
Google released three models at once on July 21, 2026. Gemini 3.6 Flash is the headline release: the new default model in the Gemini family, positioned as the everyday workhorse for coding, knowledge work and multimodal tasks. Alongside it, Google shipped Gemini 3.5 Flash-Lite, the most cost-effective model in the Flash class, and Gemini 3.5 Flash Cyber, a model fine-tuned to find and fix cybersecurity vulnerabilities that is limited to governments and trusted partners under a pilot access program.
Google did not release Gemini 3.5 Pro in this announcement. The company said the Pro model fell short of internal expectations on coding and complex reasoning, so its broader release was delayed. Google also teased Gemini 4 as its next major model.
What can Gemini 3.6 Flash do?
Gemini 3.6 Flash accepts text, image, video, audio and PDF as input, and runs a 1 million-token context window with a 64,000-token output cap. Its knowledge cutoff is March 2026, up from January 2025 on Gemini 3.5 Flash. Google reports that the model takes fewer reasoning steps and tool calls to complete multi-step workflows, which is a large part of why it uses roughly 17 percent fewer output tokens than its predecessor.
The model is built for agentic and tool-use work as well as everyday coding and knowledge tasks. It runs at about 280 tokens per second, which makes it one of the faster models in its class for interactive use. Gemini 3.6 Flash became available in GitHub Copilot on the same day as the announcement.
Gemini 3.6 Flash benchmarks and technical specs
Gemini 3.6 Flash scores 50 on the Artificial Analysis Intelligence Index, a composite benchmark covering reasoning, knowledge, mathematics and coding. On SWE-Bench Pro it reaches 58.7 percent, up from 55.1 percent for Gemini 3.5 Flash, and on Terminal-Bench 2.1 it scores 78.0 percent versus 76.2 percent. On the GDM-MRCR v2 long-context retrieval test at the full 1 million-token depth it reaches 54.0 percent, roughly double both predecessor models.
For computer use, Gemini 3.6 Flash scores 83 percent on OSWorld Verified, compared with 78.4 percent for Gemini 3.5 Flash. Google reports token-usage reductions as high as 65 percent on individual evaluations such as DeepSWE, while the overall output-token reduction is about 17 percent across the Artificial Analysis Index. Taken together, the numbers show Gemini 3.6 Flash improving on coding, long-context retrieval and computer use while doing more with fewer tokens.
How much does Gemini 3.6 Flash cost?
Gemini 3.6 Flash is priced at $1.50 per 1 million input tokens and $7.50 per 1 million output tokens through the Gemini API. Output pricing is down from $9.00 per 1 million tokens on Gemini 3.5 Flash. Combined with the lower token usage, the effective cost of typical workloads falls further than the headline output price suggests.
Gemini 3.5 Flash-Lite sits below Gemini 3.6 Flash as the most cost-effective option in the Flash class for high-volume, latency-sensitive tasks. Gemini 3.5 Flash Cyber is not generally available and is offered only to governments and trusted partners under a limited pilot. Google positioned the pricing as part of an intensifying price competition among frontier model providers.
How does Gemini 3.6 Flash compare to Gemini 3.5 Flash?
Gemini 3.6 Flash improves on Gemini 3.5 Flash across the board while lowering cost. It gains on coding (58.7 percent versus 55.1 percent on SWE-Bench Pro), long-context retrieval (54.0 percent versus roughly half that at the same depth), and computer use (83 percent versus 78.4 percent on OSWorld Verified). At the same time it uses about 17 percent fewer output tokens and drops the output price from $9.00 to $7.50 per 1 million tokens.
The knowledge cutoff also moves forward more than a year, from January 2025 to March 2026, which reduces how often the model needs external retrieval for recent facts. For teams already using Gemini 3.5 Flash, Gemini 3.6 Flash is a same-context, lower-cost upgrade rather than a new tier.
Full details are available through the official Google and Google DeepMind announcement.