Gemini 4 Argon opens to cyber defenders before developers
Gemini 4 Argon scores 77.9% on DeepSWE v1.1 and can write 1M tokens per answer, but only cyber defenders in Google's Fairwind Program can use it so far.
4 min read

By the numbers
- output tokens per answer, up from 64K
- 1M
- per million input / output tokens, introductory
- $2 / $10
- discount on cached input tokens
- 95%
- on CWE-bench v1, tied for first
- 68%
- Gemini 4 Argon
- 77.9%
- Claude Opus 5.5
- 74.2%
- GPT-6 Astra
- 74.1%
Google released Gemini 4 Argon on September 30, 2026, and calls it its most powerful model yet. Most developers cannot use it yet. Google is giving it first to trusted security teams through a program called Fairwind, and paid API customers come later.
That order says a lot about the model. Google built Argon to find and fix security holes in software on its own. A model that can do that for defenders can also help attackers, so Google is holding back the wide release.
What Google says Argon can do
Argon is aimed at long, multi-step work: coding, research, writing and cybersecurity. Google's announcement says Argon can now write up to 1 million tokens in a single answer, up from 64,000. A token is a small chunk of text, often a short word or part of a longer one. More output room means one request can produce a whole rewritten file, or many of them.
Google says Argon is built to "autonomously find, validate, and patch critical software vulnerabilities," TechCrunch reports. It can also analyze videos and charts. Google staff already use it for debugging and for moving large codebases to new libraries, TechCrunch says.
Google published these scores:
| Benchmark | What it tests | Argon's score |
|---|---|---|
| DeepSWE v1.1 | Real-world software engineering tasks | 77.9% |
| AutomationBench | Automation tasks, where Google says Argon ranks first | 51.3% |
| CWE-bench v1 | Fixing known types of security flaws | 68%, tied for first |
| LVBench | Understanding long videos | 91.7% |
9to5Google compares the DeepSWE result with two rivals. Anthropic's Claude Opus 5.5 scores 74.2% and OpenAI's GPT-6 Astra scores 74.1%. TechCrunch adds that Argon also beats GPT-6 Astra and Anthropic's Fable and Opus models on the Vals AI model index.
Who can use it, and when
For now, only the Fairwind Program has access. Its members are security teams Google trusts to defend systems. Google says paid API customers and Google AI Ultra subscribers come next, followed by wider access.
"Safely releasing frontier capabilities at this level requires a phased approach," said Koray Kavukcuoglu, a senior vice president at Google DeepMind. The company has reason to be careful. In September, Google disclosed that Gemini broke into three companies during a May security test.
Google lists several safety steps for Argon. It tested the model against indirect prompt injection, where an attacker hides instructions inside a web page or file the model reads. It also used red teams, which are testers whose job is to attack the system. 9to5Google reports that Google watches the model's internal activity for signs of misuse and has hardened the sandboxes where it runs code.
What it costs
Argon launches at an introductory price of $2 per million input tokens and $10 per million output tokens. After the introductory period, the price doubles to $4 and $20. Cached input tokens cost 95% less than normal input. Caching means the service stores a block of text you send again and again, such as a long document, so you do not pay full price each time.
Argon is not Google's only coding model. In early September, Google shipped Gemini 3.8 Flash, priced as a mid-tier model for everyday coding work.
What Google says it did with Argon inside the company
Google gives three examples from its own systems. Agents running Argon freed more than 300 TiB of memory across Google's data centers, and Google estimates total savings of 500 TiB to 1 PiB. Argon also helped move libgav1, Google's decoder for AV1 video, to new code that runs 2.7 times faster. 9to5Google reports migrations of more than 800,000 lines of code.
What this means for developers
Do not plan a launch around Argon yet. Unless your team is in Fairwind, you cannot call it today. Keep shipping on the model you use now.
Prepare a test set while you wait. Google's scores come from Google. A 3.7-point lead on DeepSWE is real but small, and your own code may not match the benchmark. Collect 20 to 50 real tasks from your backlog, with known good answers, so you can compare Argon fairly on day one.
Think about the 1-million-token output limit. It allows big jobs in one request, such as rewriting a module or writing a full migration. It also allows big bills. One full 1-million-token answer costs $10 at the introductory price and $20 after it. Set output caps in your code before you test.
Design prompts for the cache. With a 95% discount on cached input, put the large, unchanging part first, such as your repository context or API docs. Put the changing question last. That pattern can cut input costs sharply on long agent runs.
If you work in security, watch Fairwind. Google built Argon for finding and patching vulnerabilities, and defenders get it first. Teams that run code scanning or patch programs should follow how Google opens access.
Sources
Related articles

Chrome moves to a two-week release cadence
Chrome 153 starts a two-week milestone cycle, down from four. Google says Chrome 149 and 150 alone fixed 1,072 security bugs as LLMs find flaws at scale.

Xiaomi MiMo-V2.6-Pro ships 1T open weights under MIT
Xiaomi's MiMo-V2.6-Pro has 1.02 trillion parameters, 42 billion active, a 1M-token context and MIT-licensed weights, but needs about 680 GB of GPU memory.

OpenAI links reasoning-extraction campaign to Moonshot AI
OpenAI says a campaign it ties to Moonshot AI, maker of Kimi, sent 16,000 requests from 4,000+ users in July to copy its models' hidden reasoning.
The daily brief
Three to five stories a day, and what each one means for the people who build software. Free, no spam.