Google Launches Gemini 3.8 Flash & Gemini 3.8 Flash Cyber

03-09-2026

Google has made Gemini 3.8 Flash generally available with a 1 million token context window and introductory pricing equal to Gemini 3.7 Flash. The upgrade targets coding agents and professional workflows, but Google warns that it may use more reasoning tokens.

Written by:

Jorick van Weelie

Marketing Lead at DataNorth | AI Enthusiast & Tech Storyteller

google launches gemini 3.8 flash & gemini 3.8 flash cyber
Sign up for our Newsletter

Publication date: September 3, 2026

Google released Gemini 3.8 Flash and the restricted Gemini 3.8 Flash Cyber on September 2. Gemini 3.8 Flash is generally available through the Gemini API and Google AI Studio, with a 1 million token context window and introductory pricing that matches Gemini 3.7 Flash.

The practical change is not just a benchmark refresh. Google is positioning 3.8 Flash as a stronger default for agentic coding and professional workflows, while warning that it may use more reasoning tokens than 3.7 Flash.

What is Gemini 3.8 Flash and what can it do?

Gemini 3.8 Flash is Google’s latest production Flash model. It accepts text, images, video, audio and PDFs as input, and returns text. Google lists a 1,048,576 token input limit and a 65,536 token output limit.

For software teams, the main focus is agentic work: planning, tool use, coding and longer tasks that need multiple reasoning steps. Google says 3.8 Flash substantially improves over 3.7 Flash and can approach more expensive frontier models on selected evaluations. Those are vendor-reported results, not independent testing from DataNorth.

Google also released Gemini 3.8 Flash Cyber. That variant is limited to trusted defenders through the Fairwind Program. It targets vulnerability discovery and patching rather than general developer access.

Gemini 3.8 Flash pricing, access and key specifications

ItemGemini 3.8 Flash
ItemGemini 3.8 Flash
StatusGenerally available
Model IDgemini-3.8-flash
Context window1,048,576 input tokens
Maximum output65,536 tokens
Intro input price$0.75 per 1M tokens
Intro output price$3.75 per 1M tokens
Thinking levelsLow, medium and high

The introductory price applies through December 31, 2026. Google says standard pricing starts January 1, 2027 at $1.50 per million input tokens and $7.50 per million output tokens.

Developers can use 3.8 Flash in Google AI Studio, the Gemini API, Android Studio, Antigravity and Stitch. Google also lists Gemini Enterprise access. The model supports function calling, code execution, structured outputs, search grounding and context caching. It does not support the Live API, image generation or audio generation.

Gemini 3.8 Flash benchmarks and pricing

The pattern is consistent. Gemini 3.8 Flash closes on Claude Opus 5 in coding and professional analysis, then loses badly on harder agent work.

What is measuredGemini 3.8 FlashClaude Opus 5
DeepSWE v1.1 (long coding jobs)73.7%74.0%
Terminal-Bench 2.1 (terminal coding)89.4%89.1%
Vals Finance Agent v2 (financial analysis)61.4%58.6%
HLE-Verified (expert reasoning)54.9%54.4%
Terminal-Bench 4.0 (harder agent work)19.1%51.8%
OSWorld 2.0 (operating a computer)59.0%75.4%

These are Google’s own figures, published with the launch. The two bottom rows come from the same evaluation set but did not make the blog post. Independent testing by Artificial Analysis puts Gemini 3.8 Flash at 59 on its intelligence index, against 63 for Opus 5.

What can Gemini 3.8 Flash Cyber do?

The Cyber version is trained to find software flaws and write the fix. Google prioritised patching over attacking, and says the same security training is part of why the main model improved at coding. Access runs through the Fairwind Program, which is limited to government bodies, critical infrastructure operators and software maintainers.

  • CWE-Bench, an external patching test: 47.2%, just behind a leading rival at 47.8%.
  • Google’s own vulnerability test across 20 programming languages: above 70% success.
  • Chrome security team: 2.6 times more correct patches than much larger commercial models.
  • Security vendor Wiz: higher recall on its penetration testing set, at a lower cost.
  • Google Cloud research: one critical flaw found in under two hours, work that usually takes months.
  • Availability: application only, with no published price and no general release date.

What is missing from the announcement?

Google did not publish an exact launch timestamp, only the September 2 calendar date. The date overlaps the rolling 24-hour research window, and Google says the model is available now, so it qualifies under the date-only policy.

Migration is also not completely drop-in for API users. Google’s guide says developers should remove temperature, top_p and top_k settings when moving to Gemini 3.8. It replaces thinking_budget with thinking_level, does not support the minimal thinking level, and removes candidate_count.

Google has not provided independent evidence that the reported benchmark gains translate to lower total cost per completed task. That matters because a model with the same per-token price can still cost more when it reasons for longer.

What this means

Product and engineering teams already using Gemini 3.7 Flash should test 3.8 Flash first on coding agents, document-heavy workflows and tool-using tasks. Track total tokens, latency and task completion, not only per-token price.

Verdict: Gemini 3.8 Flash is worth testing now. The model is production-ready and starts at familiar pricing, but the higher reasoning-token appetite makes workload-level measurement important before a broad migration.

For more information, visit the official announcement of Gemini 3.8 Flash on the Google website.

Add DataNorth AI to your Google favorites