Overview
Google Gemini 4 Argon is a frontier model released by Google DeepMind on September 30, 2026, and is the first product in the Gemini 4 series, announced in a post by DeepMind SVP and Google Chief AI Architect Koray Kavukcuoglu. This is Google's first frontier proprietary model in nearly seven months to surpass the Flash tier, and the repeatedly delayed Gemini 3.5 Pro has been confirmed as canceled. Argon is aimed at deep reasoning for long-horizon complex workflows, with a focus on real-world software engineering, enterprise knowledge work such as legal and finance, cybersecurity defense, and creative writing. Its output token limit has increased from 64K in the previous Gemini generation to 1 million, allowing a single response to handle long-output tasks such as large-scale code migration. Against the backdrop of Gemini 3.5 Pro's failure, DeepMind leadership changes, and talent attrition, Argon is seen as Google's response to the narrative that it has fallen behind, and as a product that embodies its shift from emphasizing frontier capabilities to emphasizing cost-effectiveness.
Key Features
- Million-level output tokens: The output token limit has increased from 64K in the previous Gemini generation to 1M (1 million), allowing a single response to handle long-output tasks such as large-scale code migration.
- Long-horizon software engineering capability: It scored 77.9% on the DeepSWE v1.1 long-horizon software engineering benchmark, higher than GPT-6 Astra's 74.1%, Claude Opus 5.5's 74.2%, and Claude Fable 5.1's 67.4% in officially published comparisons.
- Enterprise knowledge work and business execution: It scored 51.3% on Zapier AutomationBench for end-to-end business execution, higher than Astra's 41.4%, Opus 5.5's 42.5%, and Fable 5.1's 31.4%; 19.6% on the Harvey legal agent benchmark, about three times the second place; and 65.4% on Vals Finance Agent v2, surpassing Claude Opus 5.5 and GPT-6 Astra.
- Cybersecurity defense and vulnerability remediation: It scored 68% on CWE-bench v1 vulnerability remediation, tied for first with GPT-6 Astra; it is first made available through the Fairwind program to trusted cyber defenders such as governments, critical infrastructure operators, and core technology platforms. That version does not include cybersecurity guardrails, while the public version retains refusals for cyber and CBRN attacks.
- Long-video understanding and creative writing: It scored 91.7% on LVBench for long-video understanding; the official materials list creative writing as one of its key scenarios.
- Third-party evaluation performance: On independent organization Artificial Analysis's Intelligence Index, Argon scored 53, tied with GPT-6 Astra and entering the top three intelligence tier; it topped the Arena.ai Text Arena leaderboard with 1,525 points; on Arena's Agent Arena, Argon (High) ranked 8th, with a net improvement of 7.92%.
Use Cases
- Long-horizon software engineering: used for debugging, migration of C/C++ to Rust codebases of up to more than 800,000 lines, and algorithm design.
- Enterprise knowledge work: scenarios such as legal agents, financial agents, and end-to-end business execution.
- Cybersecurity defense: made available through the Fairwind program to trusted cyber defenders, and Wiz has already used Argon in the Scan for Good project to scan public infrastructure.
- Long-video understanding: scored 91.7% on the LVBench long-video understanding benchmark.
- Creative writing: one of the key scenarios listed officially.
Pros
- The output token limit has increased from 64K to 1 million, allowing a single response to handle long-output tasks such as large-scale code migration.
- Among 18 officially published comparisons, it leads in 12 and is tied for first in 1, covering scenarios such as long-horizon software engineering, end-to-end business execution, legal agents, financial agents, and vulnerability remediation.
- It leads on the Vals Index (weighted by the GDP share of U.S. industries).
- In third-party evaluations, it scored 53 on the Artificial Analysis Intelligence Index, tied with GPT-6 Astra and entering the top three intelligence tier; it topped the Arena.ai Text Arena leaderboard with 1,525 points.
- Internal usage results are notable: Google engineers have already used it for debugging, migration of C/C++ to Rust codebases of up to more than 800,000 lines, and algorithm design; Argon agents freed 300 TiB of memory in data centers by discovering optimization opportunities, and one agent rewrote a video decoding component to be 2.7 times faster than the original Rust version.
- Pricing is cost-effective: introductory pricing is $2 per million input tokens and $10 per million output tokens, with a 95% discount for cached input tokens.
Pricing
Introductory pricing is $2 per million input tokens and $10 per million output tokens, with a 95% discount for cached input tokens; after the introductory period ends, it returns to $4 per million input tokens and $20 per million output tokens. For pricing information not provided, refer to the official website.
Summary
Google Gemini 4 Argon is a frontier model released by Google DeepMind on September 30, 2026, the first in the Gemini 4 series, with its output token limit increased from 64K to 1 million. It focuses on long-horizon software engineering, enterprise knowledge work such as legal and finance, cybersecurity defense, and creative writing. Among 18 officially published comparisons, it leads in 12 and is tied for first in 1, and in third-party evaluations it entered the top three intelligence tier and topped Text Arena. It is first made available through the Fairwind program to trusted cyber defenders, and subsequently opened to paid API customers and Google AI Ultra subscribers.
Version History
- Gemini 4 Argon release (2026-09-30): Google DeepMind released Argon, the first frontier model of the Gemini 4 series, raising the output token limit from 64K to 1M for long-horizon software engineering, enterprise knowledge work in law and finance, and cyber defense. Google reports leads in 12 of 18 comparisons: 77.9% on DeepSWE v1.1, 51.3% on Zapier AutomationBench, 91.7% on LVBench, and a tie for first on CWE-bench v1 with GPT-6 Astra. Introductory pricing is $2 per million input tokens and $10 per million output tokens with 95% off cached input, later rising to $4 and $20. Access begins with trusted cyber defenders via the Fairwind Program, then paid API customers and Google AI Ultra subscribers