Gemini 4 Argon: Google Launches Its Frontier Model, First to Cyberdefenders
Before developers and the general public, Google is handing its new model to a handful of security specialists, with no cyber guardrails, and promising 1 million output tokens.

Who covers it
Picked up by 3 outlets · 2 top-tier · 1 tech
In brief
Google DeepMind has announced Gemini 4 Argon, a frontier model built for long-horizon coding, enterprise knowledge work (legal, finance), and cyberdefense. Rollout is staged: first a hand-picked group of defenders via the Fairwind program, then paying API customers and Google AI Ultra subscribers. The post pairs striking internal results (300 TiB of memory freed, massive migrations to Rust) with a battery of benchmarks where Argon comes out on top.
🍺 Bar-stool version
Google just released a new brain, and the first thing it's doing with it is lending it to the people whose job is stopping other brains from causing trouble. The model finds vulnerabilities, patches them, rewrites huge swaths of code in Rust, and frees up memory in data centers the way you'd find 300 bucks in an old coat pocket — except here it's 300 terabytes. For the chosen white hats, it ships unshackled, which is basically handing over the vault keys to the people you've judged trustworthy, hoping you judged right. If it works out, infosec shifts into a new gear; if it leaks, it also shifts gears — just in the wrong direction.
Key takeaways
- 1
Gemini 4 Argon is first rolled out to trusted cyberdefenders via the Fairwind program, without cyber guardrails, ahead of a wider release to paying API customers and Google AI Ultra subscribers.
- 2
Introductory pricing is set at $2 per million input tokens and $10 per million output tokens, with 95% off cached input tokens.
- 3
The output limit jumps from 64K to 1M tokens, enabling reasoning trajectories of several hundred thousand tokens in a single pass.
- 4
Internally, Argon agents have freed over 300 TiB of memory across Google's data centers, with an estimated 500 TiB to 1 PiB in total savings.
- 5
Argon is migrating C/C++ to Rust across Google, including Fuchsia's Zircon kernel (800K+ lines), and made the Rust port of the libgav1 video decoder 2.7x faster.
- 6
On benchmarks: 77.9% on DeepSWE v1.1, 51.3% and first place on Zapier's AutomationBench, 91.7% on LVBench, and a tied first place (68%) on CWE-bench v1.
- 7
Google monitors the model's chain of thought to detect misalignment and is careful not to feed those signals back into training, to avoid teaching the model to evade oversight.
A launch in reverse
Google DeepMind presents Gemini 4 Argon as its new frontier model, built to sustain deep reasoning over long, complex tasks. Three domains are targeted: real-world software engineering, enterprise knowledge work (legal, finance), and cyberdefense.
The timeline is unusual. The model isn't going to developers first — it's going to a group of trusted cyberdefenders via the Fairwind program. Google says it is participating in the US government's voluntary pre-release model access process.
The general public, businesses, and developers will come next, "as soon as possible," starting with paying API customers and Google AI Ultra subscribers. No date is given. Introductory pricing, however, is announced: $2 per million input tokens, $10 for output, and 95% off cached input.
What Argon is already doing at Google
The post emphasizes internal use, with thousands of Googlers reportedly using it for specialized code, deep research, and writing. Three quantified examples are highlighted.
In quantum computing, Argon helps optimize the space-time resources (qubits × gates) of critical subroutines; in one case, it beat the published benchmark by 40% in a matter of minutes. On infrastructure, a team of Argon agents analyzed fleet-wide profiling telemetry to apply memory optimizations on its own: over 300 TiB freed, with 500 TiB to 1 PiB in expected savings.
The third project is migrating C/C++ code to Rust, from libraries like re2 and libgav1 to Fuchsia's Zircon kernel and its 800K+ lines. On libgav1, the agents replaced 32K lines of SIMD code with safe Rust that the compiler vectorizes on its own, producing a decoder 2.7x faster than the existing Rust port, with identical video output. Google notes these rewrites still go through automated and manual audits, emulation testing, and review before reaching production.
1 million output tokens
The most concrete technical change for users is the output limit, which jumps from 64K to 1M tokens. Google calls it the highest in the industry.
The idea: let the model produce hundreds of thousands of tokens in a single trajectory, solving a hard problem "in one go" rather than through repeated back-and-forth. It's a bet on long agentic tasks, and a potentially steep cost at $10 per million generated tokens.
Enterprise and multimodal: a flood of benchmarks
On code, Argon sets a new state of the art on DeepSWE v1.1 (77.9%), which measures long software engineering tasks. Outside of code, it tops the Vals Index, which weighs finance, code, law, and tax according to their share of US GDP, and claims strong results on Vals Finance Agent v2 and Harvey's Legal Agent Benchmark.
On AutomationBench, Zapier's benchmark for end-to-end execution of business functions, Argon ranks first with 51.3%. On vision, it reaches 91.7% on LVBench, focused on long video understanding, and Google highlights chart analysis and taking action from a series of documents.
Cyberdefense, the core of the announcement
Argon was specifically trained for defense: finding, validating, and fixing critical vulnerabilities autonomously. For trusted defenders and internal teams, it will ship without cyber guardrails, to unlock its full capability.
Wiz is already using it in its Scan for Good initiative, which protects critical public infrastructure for free. The model spotted a critical vulnerability exposing sensitive personal data in healthcare software used by hospitals worldwide — a flaw that previous models had missed.
On CWE-bench v1, which evaluates vulnerability remediation, Argon ties for first place with 68%. Google also cites progress on its internal benchmark (vulnerabilities found across codebases spanning 20 languages) and on Wiz's black-box pentesting benchmark, compared to its predecessor 3.8 Flash Cyber.
Guardrails before the rollout
Google details four workstreams. Against misuse (cyber and CBRN), the model refuses dangerous requests while preserving legitimate dual-use research, and Google monitors the model's internal activations to catch misuse, with testing by internal and external red teams.
Against indirect prompt injection, Argon is said to be the most robust model on the market on Gray Swan's IPI benchmark, thanks to automated red teaming and adversarial training. For misalignment, Google monitors the model's chain of thought and actions and can halt execution.
Notably: alerts from training-time monitoring are not fed back into training, to avoid teaching the model to evade oversight. Google calls on the rest of the industry to preserve reasoning transparency. Finally, sandbox environments are isolated and sealed off before high-risk training and evaluations.
“Safely releasing frontier capabilities at this level requires a phased approach.”
“For trusted defenders and our own internal teams at Google, we'll be releasing Argon without cyber guardrails.”
“We strongly encourage the rest of the industry to preserve reasoning transparency in these pivotal moments of increased capabilities.”
Why it matters
Gemini 4 Argon confirms a shift already underway across the industry: the most powerful models no longer launch all at once — they pass through cybersecurity and the state first. Reserving an unshackled model for "trusted" defenders is defensible, but it raises a governance question the post doesn't address: who chooses these defenders, by what criteria, and what happens if access leaks? Substantively, the internal examples (memory freed, Rust migration, 2.7x gain on libgav1) are more compelling than the litany of benchmarks, often designed by partners and sometimes won only by a tie. Some perspective is needed too: the Zircon migration is still under audit, no availability date is given, and the pricing is explicitly "introductory." Perhaps the most interesting passage is the quietest one: by asking the industry to preserve readable chains of thought, and by avoiding training against its own monitoring, Google is acknowledging that the ability to read a model's reasoning has become a fragile security asset.
Free account
You just read an AI Sources article
Create a free account: a month of archives in full, your own sources summed up like this one, your notes and highlights.
For you
Put it to work on your sources.
Free: a month of articles and three sources of your own. Pro: the whole archive and your sources, from €8/month.
For your team
The same machine, on your topics.
A space in your colours, your watch angles, your curators. Pilot open to three companies.
Read next
#geminiTodayGemini 4 Argon: Google Releases Its Most Powerful Model, But Defenders Get It First
Google announces a frontier model that rewrites kernels in Rust and hunts down vulnerabilities, then locks it away while it tightens the guardrails.
Source · Le blog de Google (The Keyword) · Gemini 4 Argon: our next era of frontier intelligence
#evaluationTodaySteering a LLM has a price: Apple measures what control costs in fluency
A study from Apple and Pompeu Fabra University shows that the most effective control methods are often the ones that damage generated text the most.
Source · Apple Machine Learning Research · On the Effectiveness-Fluency Trade-Off in LLM Conditioning: A Systematic Study
#regulationTodayCalifornia: Newsom Signs 13 AI Laws, From HR to Gene-Synthesis Labs
Algorithmic layoffs, workplace surveillance, deepfakes, synthetic DNA: Sacramento keeps stacking up safeguards while Washington looks elsewhere.
Source · Governor of California · California's nation-leading AI framework just got stronger, Governor Newsom signs more first-in-the-nation worker protections and more