Google announced Gemini 4 Argon on September 30, 2026. The model is designed for extended software engineering, enterprise knowledge work and defensive cybersecurity. Koray Kavukcuoglu, SVP of Google DeepMind and Chief AI Architect at Google, said the first external rollout is going to trusted cyber defenders through the Fairwind Program.
Google is expanding access gradually while participating in the U.S. government’s voluntary process for pre-release model access. The company plans to use feedback from early testers to refine its safeguards before opening Argon to developers, enterprises and consumers. The initial API price is $2 per million input tokens and $10 per million output tokens. Cached input tokens receive a 95% discount. After the introductory period, the prices will rise to $4 per 1M input tokens and $20 per 1M output tokens.
Argon is already being used by thousands of Googlers for specialized coding, research and writing tasks. Google says it helped quantum computing researchers reduce the spacetime resources of bottleneck subroutines, measured as qubits × gates. One result beat the published baseline by 40% within minutes. Agents also analyzed fleet-wide profiling data and identified memory optimizations that could free more than 300 TiB across Google’s data centers. The company estimates total savings of 500 TiB to 1 PiB after deployment.
Google is using Argon agents to migrate C/C++ codebases to Rust. The work ranges from tens of thousands of lines in core libraries such as re2 and libgav1 to more than 800K lines in the Fuchsia OS Zircon kernel. These changes are undergoing automated and manual audits, emulation tests and code review before production deployment.
For libgav1, an existing Rust port was optimized by replacing 32K lines of SIMD code through repeated profile-guided experiments. The agents studied compiler output and generated safe Rust that the compiler could vectorize automatically. Google reports that the result runs 2.7x faster than the Rust port, produces identical video output and is closer to the optimized C++ implementation.
Argon’s output limit is 1M tokens, compared with 64K tokens previously. Google presents the larger limit as a way to sustain reasoning and generation across long, multi-step workflows.
On DeepSWE v1.1, a benchmark for real-world long-horizon software engineering, Argon scores 77.9%. Google also reports leading results on the Vals Index, which measures economic impact across finance, coding, legal and tax work with sector weights based on U.S. GDP. The model leads on Vals Finance Agent v2 and Harvey’s Legal Agent Benchmark. It ranks first on Zapier’s AutomationBench with 51.3%.
The model also handles visual knowledge work. Google cites professional chart analysis, long-video detail extraction and actions based on document sequences. Argon scores 91.7% on LVBench, which evaluates long-video understanding.
Google trained Argon for defensive cybersecurity. The model can autonomously locate, confirm and fix critical software vulnerabilities. Trusted defenders and Google’s internal teams will receive Argon without cyber guardrails. Wiz is using it through the Scan for Good initiative, which seeks and remediates high-risk exposures in critical public infrastructure at no charge. In an early test, Argon found a critical vulnerability that exposed sensitive personal information in healthcare software used by hospitals worldwide. Google says previous frontier models had missed the issue.
Argon ties for first on CWE-bench v1 with a score of 68%, building on the results of 3.8 Flash Cyber on CWE-bench v0. Google’s internal vulnerability benchmark found exposures in complex codebases covering 20 programming languages. On Wiz’s black-box penetration-testing benchmark, Argon outperformed 3.8 Flash Cyber in mapping attack surfaces, finding vulnerabilities and generating proof-of-concept evidence.
Before a broad launch, Google is strengthening safeguards in four areas. The model is trained to reject harmful cyber and chemical, biological, radiological and nuclear requests while supporting legitimate dual-use research under the Frontier Safety Framework. Google is also improving monitoring of internal activations, as described in arXiv paper 2601.11516. Internal and external red teams tested the safeguards with manual and automated attacks.
Argon was also trained against indirect prompt injection. Google calls it its most resilient model so far in this area and reports a leading result on Gray Swan’s Indirect Prompt Injection benchmark. Misalignment controls monitor the model’s chain-of-thought and actions and can stop execution. Similar monitoring was used during training, with alerts sent to a dedicated incident-response team. Google kept those findings out of the training process to reduce the risk that the model would learn to evade monitoring.
Google is hardening the sandboxed environments used for high-risk training and evaluation. The systems are isolated and sealed before those activities begin, following the company’s agent control roadmap. The wider rollout will start with paid API customers and Google AI Ultra subscribers after further testing and feedback from trusted testers. Google also positions Argon as a tool for creative writing, developers, professionals and enterprises.




Comments
No comments yet — be the first.
Open the discussion
No account or password needed — just enter your e-mail and we’ll send you a one-time sign-in link. First time here? You’re set up automatically.
Your rating will be applied automatically after you sign in.
Check your inbox
We’ve sent a sign-in link to …. Open it on this device — this tab will sign you in automatically.
Nothing arrived? Check your spam folder — and mark the mail as "Not spam" so it lands in your inbox next time.