
Gemini 4 Argon is Google’s new frontier AI model, announced on September 30, 2026, and built for long-running coding, enterprise knowledge work, and autonomous cyber defense. Today only trusted security partners in Google’s Fairwind Program can use it. Google says broader access for developers, enterprises, and consumers is coming soon.
Every AI lab calls its newest release its best ever, and those claims tend to blur together. This launch deserves a closer look because Google attached concrete details: a 1 million token output limit, a reported state-of-the-art score on a real-world software engineering test, and a model that can reportedly find, validate, and patch serious software flaws without a human steering each step. This guide separates what Google has documented from what is still marketing, so you can judge what it means for your team.
Sources: Google’s official announcement by Koray Kavukcuoglu and TechCrunch’s launch coverage by Lucas Ropek. Benchmark figures are reported by Google unless noted otherwise.
At a Glance: Key Facts
- Announced: September 30, 2026, by Google DeepMind.
- Focus areas: Real-world software engineering, enterprise knowledge work such as legal and finance, and cybersecurity defense.
- Output limit: 1 million tokens, up from the previous 64,000.
- Access today: Trusted cyber defenders through the Fairwind Program, plus Google’s internal teams.
- Launch pricing: $2 per million input tokens and $10 per million output tokens, rising to $4 and $20 after the introductory period.
- Headline coding score: 77.9% on DeepSWE v1.1, which Google describes as a new state of the art.
- Wider release: Paid API customers and Google AI Ultra subscribers first, with no date announced.
What Is Gemini 4 Argon?
Definition: Gemini 4 Argon is a frontier large language model from Google DeepMind, designed to sustain deep reasoning across complex, long-horizon workflows. “Long-horizon” means tasks with many dependent steps, such as migrating a large codebase, researching a financial question across dozens of documents, or hunting for a vulnerability across millions of lines of code.
Expansion: Google positions Argon around three domains. The first is real-world software engineering. The second is enterprise knowledge work, including legal and finance tasks. The third is cybersecurity defense, where the model was trained specifically for defensive work. It also handles visual inputs, including charts and long videos. TechCrunch frames it as a workhorse for coding and cybersecurity, and Google says thousands of its own employees already use it for specialized coding, deeper research, and writing.
How Is Argon Different From Earlier Gemini Models?
Short answer: Three things stand out. The output limit jumps from 64,000 tokens to 1 million. The model is trained for defensive cyber work, and trusted defenders get a version without cyber guardrails. And the release is gated: it starts with vetted partners instead of the public.
| Feature | Previous Gemini models (per Google) | Argon |
|---|---|---|
| Maximum output tokens | 64,000 | 1,000,000 |
| Cyber benchmark | Gemini 3.8 Flash Cyber set frontier performance on CWE-bench v0 | Ties for first on CWE-bench v1 with 68% |
| Vulnerability discovery | Baseline set by 3.8 Flash Cyber | Described by Google as a leap, on both Google’s and Wiz’s internal tests |
| Release approach | Not covered in the announcement | Phased: trusted defenders first, then paid API and Ultra users |
The 64,000 figure is Google’s stated previous output limit. Google does not say which specific model it applied to.
Who Can Use It Right Now?
Is Gemini 4 Argon Available to the Public?
No. As of Google’s September 30 announcement, Argon is rolling out only to a set of trusted cyber defenders through the Fairwind Program. Google says it will expand access gradually while it strengthens safeguards, but it has not published a public release date.
What Is the Fairwind Program?
Definition: The Fairwind Program is Google’s security initiative for working with trusted cyber partners. TechCrunch describes it as the channel through which Argon is reaching a select group of the company’s cyber partners.
Expansion: Google says that for trusted defenders and its own internal teams, it will release Argon without cyber guardrails. The goal is to let them use the model’s full defensive capability. Wiz, a cloud security company, is one early user, which the cybersecurity section below covers.
When Will Wider Access Arrive?
Short answer: Google has not said. The company states it is engaged in the U.S. government’s voluntary process for pre-release model access while it expands availability gradually. When broader release comes, Google says it will start with paid API customers and Google AI Ultra subscribers.
| Stage | Audience | Status |
|---|---|---|
| 1 | Trusted cyber defenders (Fairwind Program) and Google internal teams | Active |
| 2 | Paid API customers and Google AI Ultra subscribers | Planned, no date |
| 3 | Remaining developers, enterprises, and consumers | Planned, no date |
Core Capabilities
What Does a 1 Million Token Output Limit Mean?
Short answer: The model can produce up to 1 million tokens in a single run, roughly 15 times the earlier 64,000 cap.
Why it matters: Google says that when the model has room to think deeply and generate hundreds of thousands of tokens in one trajectory, it can solve hard problems in one pass instead of being interrupted by length limits. For agent workflows like codebase refactors or long research tasks, that means fewer hand-offs and less stitching together of partial results. The trade-off is cost, since output tokens are the more expensive side of the pricing.
Coding and Software Engineering
Argon posts 77.9% on DeepSWE v1.1, a benchmark of real-world, long-horizon software engineering tasks. Google calls that a new state of the art. The company also shared internal examples:
- Memory optimization: A team of Argon agents analyzed fleet-wide profiling data and applied memory optimizations across Google’s data centers. Google says this frees over 300 TiB once rolled out, with total estimated savings of 500 TiB to 1 PiB.
- Codebase migration: Argon agents are working on C/C++ to Rust migrations, from tens of thousands of lines in libraries like re2 and libgav1 up to 800,000+ lines for the Fuchsia Zircon kernel. Google notes these rewrites face automated and manual auditing, emulation testing, and review before reaching production.
- Performance engineering: For the libgav1 video decoder, Argon replaced 32,000 lines of SIMD code with safe Rust. Google reports a memory-safe decoder that runs 2.7x faster than the earlier Rust port, with identical video output.
- Quantum research: Argon helped researchers optimize the resource cost of quantum subroutines, beating a published baseline by 40% in minutes, according to Google.
Enterprise Knowledge Work
Beyond code, Google says Argon leads the Vals Index, which measures economic impact across finance, coding, legal, and tax work, weighting each sector by its contribution to U.S. GDP. It also reports leading results on Vals Finance Agent v2 (multi-step financial research) and Harvey’s Legal Agent Benchmark (legal research and drafting). On Zapier’s AutomationBench, which tests end-to-end execution of business functions, Argon ranks first at 51.3%.
Visual and Video Understanding
Argon is built to drive professional chart analysis, pick out details from long videos, and act on a series of documents. On LVBench, a long-video understanding test, Google reports a state-of-the-art 91.7%.
Why Cybersecurity Is the Headline Feature
What Can This AI Cybersecurity Model Actually Do?
Short answer: Google says Argon can autonomously find, validate, and patch critical software vulnerabilities.
Expansion: That sentence describes three separate jobs. Finding means discovering a flaw in code or a running system. Validating means confirming it is real, which filters out the false alarms that bury human analysts. Patching means writing a fix. Chaining all three is the promise of AI vulnerability patching: shrinking the gap between a bug existing and a bug being closed.
Real-World Proof Point: Wiz and Healthcare Software
Wiz is already using Argon through its Scan for Good initiative, a free program for protecting critical public infrastructure. In an early demonstration, the model uncovered a critical vulnerability that exposed sensitive personal information in healthcare software used by hospitals worldwide. Google says earlier frontier models had missed it. This is a single example supplied by the vendor and its partner, but it is the kind of concrete result that separates a useful AI cybersecurity model from a benchmark trophy.
How Does Argon Perform on Cyber Benchmarks?
- CWE-bench v1: Argon ties for first place with 68%, evaluating how well a model remediates security vulnerabilities.
- Google’s internal vulnerability benchmark: The model uncovered a wide range of exposures across complex codebases in 20 programming languages.
- Wiz’s black-box penetration test benchmark: Without source code access, Argon outperformed Gemini 3.8 Flash Cyber at mapping the attack surface, identifying vulnerabilities, and producing proof-of-concept evidence.
Why Release It Without Cyber Guardrails?
Short answer: Defenders need the full capability, and the same capability is dangerous in the wrong hands.
Offensive and defensive security skills overlap heavily. A model that can find and exploit a flaw to prove it exists can also be misused. Google’s approach is to give vetted defenders and its own teams the unrestricted version, while keeping safeguards in place and strengthening them before any broad release. That is why the Fairwind Program gate exists, and it is a reasonable reading of why public access is delayed.
Benchmark Results at a Glance
| Benchmark | What it measures | Reported Argon result |
|---|---|---|
| DeepSWE v1.1 | Long-horizon, real-world software engineering | 77.9% (state of the art) |
| Vals Index | Economic impact across finance, coding, legal, and tax | Leading model |
| Vals Finance Agent v2 | Multi-step financial research | Leading |
| Harvey Legal Agent Benchmark | Legal research and drafting | Leading |
| AutomationBench (Zapier) | End-to-end business function execution | 51.3% (#1) |
| LVBench | Long video understanding | 91.7% (state of the art) |
| CWE-bench v1 | Remediating security vulnerabilities | 68% (tied for first) |
| Gray Swan IPI | Resistance to indirect prompt injection | Leading |
All figures are reported by Google. Treat them as strong claims awaiting independent replication.
Gemini 4 Argon vs GPT-6 Astra, Fable, and Opus
How does Argon compare with rival models? According to TechCrunch, Google’s announcement claims Argon scored significantly higher than OpenAI’s GPT-6 Astra and Anthropic’s Fable and Opus models across a range of benchmarks. Google also cites Vals, an AI benchmarking startup, to show Argon leading that company’s model index.
Two cautions apply. First, these are vendor-reported results, and every major lab publishes benchmarks that flatter its own model. Second, because access is limited to trusted partners, independent evaluators cannot yet run their own head-to-head tests. Until they can, the fairest reading is that Google has made a credible, specific, and unverified claim.
The competitive context explains the urgency. TechCrunch notes that Google was once considered behind in the AI race, but the Gemini app announced more than a billion monthly users in August, putting it in the same range as ChatGPT. Each lab is now racing to ship a flagship that outdoes the last.
Argon Pricing: What Will It Cost?
Short answer: Argon pricing starts at $2 per million input tokens and $10 per million output tokens during an introductory period, then rises to $4 and $20.
Gemini 4 Argon will launch at that introductory rate, with cached input tokens priced at 95% off the input price. That works out to roughly $0.10 per million cached input tokens at the introductory rate, which could matter for workflows that reuse a large codebase or document set.
| Pricing tier | Input (per 1M tokens) | Output (per 1M tokens) | Cached input |
|---|---|---|---|
| Introductory | $2 | $10 | 95% off input price |
| After introductory period | $4 | $20 | 95% off input price |
Worked example: A long agent run that reads 1 million input tokens and writes 100,000 output tokens would cost about 3attheintroductoryrate(2 for input, $1 for output). The same run would cost about $6 afterward. Because the output limit is now 1 million tokens, a run that fills it would cost $10 in output alone at introductory rates, so budget accordingly.
Safety Safeguards Google Says It Is Adding
Google lists four areas of work to strengthen before broad release:
- Defending against misuse. The model is designed to refuse harmful requests tied to cyber or chemical, biological, radiological, and nuclear (CBRN) attacks while preserving legitimate dual-use science. Google says it is improving monitoring of the model’s internal activations and testing the safeguards with internal and external red teams.
- Defending against prompt injection. Gemini 4 Argon is described as Google’s most resilient model yet against indirect prompt injection, where malicious instructions hidden in content hijack a model’s behavior. Google reports leading results on Gray Swan’s IPI benchmark.
- Monitoring for misalignment. Mitigations watch the model’s chain-of-thought and actions and halt execution when it steps beyond what the user intended.
- Hardening systems. Google is isolating and sealing its sandboxed environments before high-risk training or evaluations begin.
What Argon Means for Different Teams
Security Teams
If you are a defender, the practical question is access. The Fairwind Program is the current route, so teams responsible for critical infrastructure should look into how to engage with it. Longer term, AI vulnerability patching could change triage: instead of a queue of unverified alerts, analysts review validated findings with proposed fixes attached.
Engineering Leaders
The migration and optimization examples suggest where agentic models are heading: large, auditable refactors rather than snippet generation. Note Google’s own caveat that critical rewrites still go through rigorous automated and manual review. Plan for human verification, not blind trust.
Legal, Finance, and Operations Leaders
Strong results on finance, legal, and automation benchmarks hint at multi-step research and document workflows becoming more automatable. Benchmarks are not your workflows, so pilot on your own data once the model is available.
Limitations and Open Questions
- No public availability or date. Everything here is based on Google’s claims and early-partner use.
- Self-reported benchmarks. Independent verification is pending, especially for the head-to-head claims against competitors.
- Migrations still under audit. The Rust rewrites are not yet in production.
- Introductory pricing is temporary. Launch costs will double after the introductory period.
- Dual-use risk. A model strong enough at defense is also strong at offense, which is why the gating and safeguards matter.
Frequently Asked Questions
Is Gemini 4 Argon free to use?
No. Google has announced token-based pricing for API use, and says broader release will begin with paid API customers and Google AI Ultra subscribers. No free tier has been announced.
What is Gemini 4 Argon best at?
Google emphasizes three strengths: long-horizon software engineering, enterprise knowledge work in areas like finance and legal, and defensive cybersecurity. It also reports state-of-the-art long video understanding.
How much does Gemini 4 Argon cost?
At launch, $2 per million input tokens and $10 per million output tokens, rising to $4 and $20 after the introductory period. Cached input tokens are 95% cheaper than standard input.
Can I use Argon today?
Only if you are a trusted cyber defender in the Fairwind Program or part of Google’s internal teams. Everyone else is waiting on a phased rollout with no announced date.
Who announced the model?
Koray Kavukcuoglu, Google DeepMind’s SVP and Chief AI Architect, announced it on Google’s blog on September 30, 2026.
Bottom Line
Gemini 4 Argon matters less as another “most powerful model” headline and more as a signal of where frontier AI is going: very long outputs, autonomous multi-step work, and a split release strategy where the most capable cyber features reach vetted defenders before the public. The claims are specific and impressive, but they are Google’s own, and access is narrow. Watch for independent evaluations, the public rollout timeline, and real-world results from partners like Wiz. Those will tell you whether this is a step change or a very good benchmark run.