Primary sourceAIArticle··6 min read

Gemini 4 Argon: Google Releases Its Most Powerful Model, But Defenders Get It First

Google announces a frontier model that rewrites kernels in Rust and hunts down vulnerabilities, then locks it away while it tightens the guardrails.

Gemini 4 Argon: Google Releases Its Most Powerful Model, But Defenders Get It First
Source : Koray Kavukcuoglu · Google · 30 September 2026View original ↗

Who covers it

Picked up by 40 outlets · 6 top-tier · 6 tech · 6 forums · 30 social posts

In brief

Google has unveiled Gemini 4 Argon, its new frontier model built for long-form coding, high-value knowledge work (law, finance), and cyberdefense. For now it's only accessible to hand-picked cyber defenders through the Fairwind Program, ahead of a gradual rollout to paying API customers and Google AI Ultra subscribers. The announcement combines chart-topping benchmarks, aggressive launch pricing, and an unusually detailed safety chapter.

🍺 Bar-stool version

Google built an AI so good at hacking that it decided to lend it only to the good guys — and even stripped out the safeguards for them, because the good guys, as everyone knows, can totally be trusted. In the meantime, it's already freed up 300 TB of memory across Google's data centers and rewritten code that's generations old, which is basically like hiring an intern who cleans out the basement in one afternoon. The rest of us, meanwhile, are stuck waiting for 'soon,' a word that in tech can mean anything from next week to the next ice age. What actually matters: models are getting good enough to find flaws nobody else spotted, and the question isn't whether they'll do it anymore — it's who they'll do it for.

Key takeaways

  1. 1

    Gemini 4 Argon is first being rolled out to trusted cyber defenders through the Fairwind Program, as part of the U.S. government's voluntary pre-launch access process.

  2. 2

    Introductory pricing is set at $2 per million input tokens and $10 per million output tokens, with a 95% discount on cached tokens.

  3. 3

    The output limit jumps from 64K to 1 million tokens, enabling reasoning and generation spanning hundreds of thousands of tokens in a single pass.

  4. 4

    Google claims the top spot on DeepSWE v1.1 (77.9%), Zapier's AutomationBench (51.3%), LVBench (91.7%), and the Vals Index, plus a tied lead on CWE-bench v1 (68%).

  5. 5

    Internally, Argon agents are migrating code from C/C++ to Rust, including Fuchsia's 800K+ line Zircon kernel, and have freed up over 300 TiB of memory in data centers.

  6. 6

    Authorized defenders receive a version without cyber guardrails; working with Wiz, the model found a critical vulnerability in hospital software used worldwide.

  7. 7

    On safety, Google monitors the model's chain of reasoning and is urging the industry to preserve reasoning transparency to catch misalignment.

A two-phase launch

Gemini 4 Argon is presented as Google's new frontier model, built for deep reasoning on long, multi-step tasks. Its three main arenas: real-world software engineering, enterprise knowledge work (law, finance, tax), and cyberdefense.

But the general public doesn't have access yet. The model is first going out to a circle of cyber defenders through the Fairwind Program, while Google takes part in the voluntary pre-launch access process set up by the U.S. government.

Access will then expand in stages, starting with paying API customers and Google AI Ultra subscribers. No date has been given, just a 'rolling out soon.' Launch pricing, however, is already set: $2 per million input tokens, $10 per million output tokens, and a 95% discount on cached tokens.

What Argon is already doing at Google

According to Koray Kavukcuoglu, thousands of employees already use Argon daily, from debugging to algorithm design. Several concrete examples have been cited.

In quantum computing, the model optimized the spacetime resources (qubits × gates) of a subroutine, beating the published benchmark by 40% in just a few minutes. On infrastructure, a team of Argon agents analyzed profiling telemetry across the entire fleet and autonomously applied memory optimizations: over 300 TiB freed up, with total gains estimated between 500 TiB and 1 PiB.

The most striking project is the migration from C/C++ to Rust, covering libraries like re2 and libgav1 up to Fuchsia's Zircon kernel (800K+ lines). On libgav1, Google's open-source video decoder, agents replaced 32K lines of SIMD code with safe Rust that the compiler auto-vectorizes, producing a decoder 2.7 times faster than the existing Rust port, with identical video output. Google notes that these rewrites go through automated and manual audits before any production deployment.

A million output tokens and chart-topping benchmarks

The most notable technical change is the output limit, jumping from 64K to 1 million tokens. The idea: let the model think and produce hundreds of thousands of tokens in a single trajectory to solve a hard problem in one go.

On coding, Argon scores 77.9% on DeepSWE v1.1, which measures long software engineering tasks. Outside of coding, it claims the top spot on the Vals Index (weighted by each sector's contribution to U.S. GDP), Vals Finance Agent v2, Harvey's Legal Agent Benchmark, and Zapier's AutomationBench (51.3%).

Google also highlights visual understanding applied to office work: chart analysis, details spotted in long videos, actions derived from series of documents. On LVBench, dedicated to long videos, the model reaches 91.7%.

Cyberdefense as the showcase

Argon was specifically trained to find, validate, and fix vulnerabilities autonomously. For trusted defenders and internal teams, Google ships it without cyber guardrails, so they can tap into its full capabilities.

Wiz is already using it in its Scan for Good initiative, which protects critical public infrastructure for free. The model reportedly discovered a critical vulnerability exposing sensitive personal data in healthcare software used by hospitals worldwide — a flaw previous frontier models had missed.

On CWE-bench v1, which evaluates vulnerability remediation, Argon ties for first place with 68%. Google also claims clear progress over 3.8 Flash Cyber, notably on an internal benchmark covering 20 languages and on Wiz's black-box pentesting benchmark, which requires analyzing live production web systems without source code.

Four safety workstreams before the wider rollout

Against malicious use (cyber and CBRN), the model refuses dangerous requests while preserving legitimate dual-use research, in line with the Frontier Safety Framework. Google is notably strengthening monitoring of the model's internal activations to catch misuse, with testing from internal and external red teams.

Against indirect prompt injections, Argon is presented as Google's most resistant model yet, topping Gray Swan's IPI benchmark thanks to automated red teaming and adversarial training.

Against misalignment, Google monitors the model's chain of reasoning and actions, and can interrupt execution. The same system was used during training, carefully avoiding feeding findings back into training so as not to teach the model to evade monitoring. Finally, sandbox environments are isolated and sealed before high-risk training runs and evaluations.

“Safely releasing frontier capabilities at this level requires a phased approach.”
“For trusted defenders and our own internal teams at Google, we'll be releasing Argon without cyber guardrails so they can leverage its full frontier-level cybersecurity defense capabilities.”
“We strongly encourage the rest of the industry to preserve reasoning transparency in these pivotal moments of increased capabilities while navigating alignment risks.”

Why it matters

Beyond the benchmark race, this announcement says a lot about where the industry stands. First, cyberdefense has become the launch pitch for a frontier model: companies aren't just selling an assistant anymore, they're selling an offensive-defensive capability deemed too powerful to distribute freely — powerful enough to hand over without guardrails to partners chosen by Google itself. That's consistent, but it concentrates real power in the hands of a few players who get to decide who counts as 'trusted.' Second, the internal examples (Rust, memory, quantum) are more convincing than the benchmarks, many of which are recent, niche, or in-house (Vals, Harvey, Zapier, Google's and Wiz's internal benchmarks), and some scores aren't even quantified in the post. The very low launch pricing and the 1-million-token output clearly target long agentic workloads, where competition is now playing out. Finally, the explicit call to preserve reasoning transparency is a strong signal: Google is admitting that monitoring the chain-of-thought is one of the few reliable tools against misalignment — and that it could disappear if the industry optimizes poorly. One simple question remains: as long as Argon stays 'coming soon,' all of this is just a promise.

Free account

You just read an AI Sources article

Create a free account: a month of archives in full, your own sources summed up like this one, your notes and highlights.

#ai#llm#google#gemini#cybersecurity#agents
Original source
Gemini 4 Argon: our next era of frontier intelligence
Koray Kavukcuoglu
Open the article ↗

For you

Put it to work on your sources.

Free: a month of articles and three sources of your own. Pro: the whole archive and your sources, from €8/month.

For your team

The same machine, on your topics.

A space in your colours, your watch angles, your curators. Pilot open to three companies.

Read next