Published: 1 October 2026
Google DeepMind released Gemini 4 Argon on 30 September 2026, a frontier model that can write up to 1 million tokens in a single response, against 64,000 on earlier Gemini models. It is not generally available. For now, access is limited to cyber defenders in Google’s Fairwind Program while the model goes through a US government pre-release review.
What is Gemini 4 Argon built for?
Argon is aimed at long jobs rather than single questions. Google names three: real-world software engineering, knowledge work in law and finance, and cyber defence. The 1 million token output follows from that. A model that can write out a full refactor or a complete contract review in one pass does not need you to restart it halfway and stitch the pieces back together.
That ceiling matters more than it sounds. Earlier Gemini models stopped at 64,000 output tokens, so an agent working through a large codebase had to break up its own work and carry state between calls. Every restart is a point where context gets dropped and the agent repeats itself. Argon removes that particular failure mode, at least on paper.
Gemini 4 Argon benchmarks against GPT-6 Astra and Claude Opus 5.5
Google published 18 benchmarks and leads or ties on 13 of them. These six show where Argon pulls ahead and where it does not.
| What is measured | Gemini 4 Argon | Best named rival |
|---|---|---|
| DeepSWE v1.1 (real-world software engineering) | 77.9% | Claude Opus 5.5: 74.2% |
| Vals Index (finance, legal, tax and coding work) | 68.9% | Claude Opus 5.5: 67.0% |
| AutomationBench (running business tasks end to end) | 51.3% | Claude Opus 5.5: 42.5% |
| Harvey Legal Agent Benchmark (legal agent work) | 19.6% | GPT-6 Astra: 5.4% |
| FrontierSWE v2 (the hardest coding tasks) | 55.0% | GPT-6 Astra: 65.5% |
| Terminal-Bench 4.0 (command line work) | 57.4% | Claude Opus 5.5: 66.4% |
Every figure here is Google’s own, taken from the launch post. Google does not say where the GPT-6 Astra and Claude Opus 5.5 numbers came from, so read the rival column as Google’s framing rather than an independent rerun. The last two rows are the ones Google did not put in its headline. On the hardest coding set and on command line work, Argon loses by around 10 points.
One number is worth more than the coding scores for security teams. On the Gray Swan indirect prompt injection benchmark, which tests whether hidden instructions in a document can hijack the model, Argon shows a 0.7% attack success rate. That is the figure a defender would actually check before pointing a model at untrusted input.
What does Gemini 4 Argon cost and who can use it?
- Introductory price: $2 per million input tokens and $10 per million output tokens
- Standard price afterwards: $4 per million input and $20 per million output
- Cached input gets 95 percent off, which works out at $0.10 per million tokens
- Available today only to trusted cyber defenders in Google’s Fairwind Program, with the cyber safety guardrails removed
- Next in line: paid API customers and Google AI Ultra subscribers, with no date given
- There is no public API model ID yet, so you cannot write code against it
At the introductory rate Argon costs roughly one fifth of GPT-6 Astra per token and about half of Claude Opus 5.5. That gap narrows sharply once the standard price applies, and Google has not said when the introductory period ends. A cost model built on $2 and $10 is a cost model with a hole in it.
What Google is not saying
Three gaps stand out. There is no general availability date, no published rate limits or data governance terms, and no input context window figure anywhere in the launch post. That last one is strange for a model sold on long-horizon work. Google published the output ceiling in the headline and left the input side out.
Google also says Argon found a critical vulnerability in healthcare software that earlier frontier models missed. No CVE, no product named and no method described. It is a good anecdote and it is not evidence. CWE-bench v1, where Argon ties for first at 68%, is the measurable version of that claim and it sits well below a score you would trust unsupervised.
What this means
For almost everyone this is worth watching, not worth testing, because you cannot test it. There is no model ID, no availability date and no rate limits. If you run an agent platform and you are planning capacity for the next quarter, the useful signal is the price: Google is willing to open at $2 and $10 against GPT-6 Astra’s far higher rate, which tells you where frontier pricing is heading more clearly than any benchmark here does.
The exception is a security team already inside the Fairwind Program. For them this is worth testing now, and the thing to test first is not the coding scores but the 0.7% prompt injection figure, on your own documents and your own tooling. For an engineering team choosing a coding agent today, Argon changes nothing yet: GPT-6 Astra still wins the hardest coding benchmark by 10 points and you can actually buy it. Revisit when the model ID appears.
For more information, visit the official announcement of Gemini 4 Argon on the Google blog.