Skip to content

AI·News & analysis

Google's newest AI model is so good at finding security holes, you can't use it yet

Google announced Gemini 4 Argon, its new frontier AI model, but is initially releasing it only to trusted cyber defenders through a vetting program rather than the general public.

By Dan Kost aka Poseidan8 min read
Google CEO Sundar Pichai standing in front of a world map graphic
Photo: Maurizio Pesce from Milan, Italy / Wikimedia Commons, CC BY 2.0

Tide

Wave

Sci-fi

6/10

Reality

Demo

An AI so good at hacking defense, Google won't let just anyone use it.How we rate

The Squeeze

Google announced Gemini 4 Argon, a new frontier AI model that ties for first place on a major cybersecurity vulnerability-finding benchmark and leads several others in coding and enterprise work.

Instead of launching broadly, Google is releasing it first, without its usual cyber guardrails, to vetted cyber defenders through a program called Fairwind, citing the model's real dual-use risk. It's a notable moment: an AI company restricting its own most capable model specifically because it's too good at a sensitive task for unrestricted public release.

What to know

  1. Google announced Gemini 4 Argon, its new frontier AI model built for complex, long-horizon tasks across software engineering, enterprise knowledge work, and cybersecurity defense.
  2. Argon is rolling out first to vetted cyber defenders through Google's Fairwind Program, without the usual cyber guardrails, rather than launching broadly to consumers or developers.
  3. The model ties for first place on CWE-bench v1, a benchmark measuring the ability to find and fix security vulnerabilities, scoring 68%, and Google's security partner Wiz used it to uncover critical healthcare software vulnerabilities other frontier models missed.
  4. Argon also leads several other benchmarks, including state-of-the-art coding performance on DeepSWE v1.1 and the top spot on the Vals Index, which measures finance, legal, and tax work.
  5. Google is strengthening safeguards against misuse, prompt-injection attacks, and model misalignment before making Argon broadly available to developers, enterprises, and consumers.

Google just announced its most capable AI model yet, and its very first move afterward was to make sure almost nobody can actually use it right away.

What is Gemini 4 Argon, and what makes it different?

Gemini 4 Argon is Google's new frontier AI model, built specifically to "sustain deep reasoning across complex, long-horizon workflows." Google says it delivers top-tier performance across three areas: real-world software engineering, enterprise knowledge work like legal and finance, and cybersecurity defense.

One major technical upgrade: Argon supports an industry-leading 1 million token output limit, up from 64,000 tokens in prior models. That lets it work through far deeper, longer problem-solving sequences in a single continuous session, instead of needing a task broken into smaller pieces.

Why can't most people actually use it yet?

The catch: Google is releasing Argon first to vetted cyber defenders through a program called Fairwind, deliberately without the cybersecurity guardrails built into its public-facing models. That's specifically so trusted security professionals can use Argon's full vulnerability-finding capability, rather than a version deliberately held back for safety.

In real life it's like a locksmith company building a master key that can open almost any lock, and initially handing it out only to licensed locksmiths instead of selling it in stores, precisely because of how well it works.

Why it matters: that's a genuinely unusual move in AI right now. Most companies race to put their newest, most capable model in front of as many users as possible as quickly as they can manage. Google deliberately restricting Argon's early access is a direct acknowledgment that this particular model's capabilities carry real dual-use risk worth taking seriously.

How good is Argon actually at finding security vulnerabilities?

Argon ties for first place on CWE-bench v1, a benchmark specifically measuring a model's ability to find and patch software security vulnerabilities, with a top score of 68%. Google's security partner Wiz reportedly used Argon to uncover critical vulnerabilities in healthcare software that other frontier AI models had missed entirely.

  • CWE-bench v1: 68%, tied for first place.
  • Real-world validation: found healthcare software vulnerabilities other models missed.
  • Access model: unrestricted cyber capabilities, vetted defenders only, via Fairwind.

Is Argon good at anything besides cybersecurity?

Yes, extensively. Argon achieves state-of-the-art performance on DeepSWE v1.1, a real-world software engineering benchmark, hitting 77.9%. It also leads the Vals Index, which measures finance, legal, and tax-related work, scores 91.7% on long-video understanding (LVBench), and ranks first on AutomationBench with 51.3%, a benchmark for completing multi-step automated tasks.

Why it matters: those results span well beyond cybersecurity specifically, suggesting Argon's underlying reasoning improvements apply broadly across demanding professional tasks, not just the one specific security-focused capability Google is being most cautious about releasing to the public right now.

What safety work is happening before wider release?

Google said it's strengthening safeguards in four specific areas before Argon becomes broadly available:

  • Misuse prevention: refusing genuinely harmful requests while still protecting legitimate dual-use security research.
  • Prompt-injection defense: leading performance on an industry benchmark for indirect prompt-injection attacks.
  • Misalignment monitoring: tracking the model's own reasoning process and halting execution when something looks wrong.
  • Sandbox hardening: isolating high-risk training and evaluation environments more thoroughly.

Why it matters: that list reads less like standard release boilerplate and more like a real acknowledgment that a model this capable at finding and exploiting vulnerabilities needs meaningfully more safety infrastructure than a typical chatbot release, before it reaches people whose intentions Google can't vet.

How much will it cost, and when can regular users try it?

Introductory pricing is currently set at $2 per million input tokens and $10 per million output tokens, with a 95% discount on cached inputs. Standard pricing rises to $4 and $20 per million tokens respectively once that introductory period ends.

Google said Argon will become available to developers, enterprises, and everyday consumers "as soon as possible," starting with paid API customers and Google AI Ultra subscribers, though no specific public release date has been confirmed yet.

Is this Google retaking an AI lead from OpenAI and Anthropic?

Background: some coverage framed Argon's benchmark results as Google reclaiming a performance edge over rival frontier models from OpenAI and Anthropic, both of which have released their own advanced models in recent months. Google's own announcement focused mainly on Argon's specific coding, enterprise, and cybersecurity strengths rather than drawing direct head-to-head comparisons itself.

Why it matters: regardless of how the competitive framing shakes out, Argon's unusual restricted launch is arguably the more notable story here. It's a rare example of a major AI lab treating "this model is extremely capable" as a reason to slow down public access, rather than purely a marketing point to lead with.

How did Google's Gemini models get here?

Background: Google first announced Gemini 1.0 in December 2023, its first natively multimodal model family, released in three sizes: Ultra, Pro, and Nano. Google rebranded its Bard chatbot as Gemini in February 2024, the same month it introduced Gemini 1.5 Pro with a then-groundbreaking 1 million token context window, a figure that's now become the output limit for Argon rather than just its input capacity.

The most recent prior generation, Gemini 3, arrived in November 2025, bringing extended reasoning and agentic capabilities, with a faster Flash variant that reportedly outperformed the larger Pro model on coding tasks while running three times faster. Argon represents the next full step beyond that generation, arriving less than a year later.

Why it matters: that pace, major model generations arriving roughly every few months to about a year apart, reflects just how quickly frontier AI capabilities have been advancing recently. It also means Argon's unusually cautious, restricted rollout stands out specifically because it breaks from Google's own recent pattern of releasing new Gemini generations broadly and quickly, rather than gating access behind a vetting program from day one.

What it means for you

  • You won't be able to use Argon directly yet, unless you're a vetted cybersecurity professional accepted into the Fairwind Program or an early paid API/Google AI Ultra customer once that access opens.
  • If you work in security, Argon's vulnerability-finding capability is worth watching closely once broader access opens, given its benchmark results and the real-world healthcare vulnerability discovery already reported.
  • The eventual public release is likely to arrive with more guardrails than the defender-only version, so expect the consumer-facing Argon to behave somewhat differently than what cyber defenders get access to now.
  • This restrained rollout approach may become a template for future powerful AI models, particularly ones with strong dual-use capabilities in sensitive domains like security.

The bottom line

Google built an AI model good enough at finding security vulnerabilities that its first move was limiting who could actually use it, a genuinely different posture than the usual "ship it to everyone immediately" approach that's defined most major AI model launches this year.

Whether that caution reflects real, carefully considered risk management or just careful optics around a genuinely dual-use capability, the practical result is the same for now: one of the most capable AI models Google has ever built is sitting mostly out of reach for anyone who isn't already a trusted security professional, at least until Google decides the remaining safeguards are ready.

Key facts

Model
Gemini 4 Argon
Cybersecurity benchmark
CWE-bench v1, 68% (tied #1)
Initial access
Vetted cyber defenders (Fairwind)
Output limit
1 million tokens (up from 64K)
Intro pricing
$2/$10 per million tokens (in/out)

Got questions?

Quick answers, plain words

What is Gemini 4 Argon?

Google's new frontier AI model, designed to sustain deep reasoning across complex, long-horizon workflows. Google says it delivers top-tier performance in real-world software engineering, enterprise knowledge work like legal and finance, and cybersecurity defense specifically.

Why can't the public use it yet?

Google is initially releasing Argon only to vetted cyber defenders through its Fairwind Program, without the cybersecurity guardrails built into its public models, so that defenders can use its full vulnerability-finding capability. Broader release is planned once Google strengthens additional safeguards against misuse.

What is the Fairwind Program?

Google's vetting process for granting trusted cybersecurity professionals and organizations, along with voluntary U.S. government pre-release participants, early access to Argon's unrestricted cyber-defense capabilities before the model becomes broadly available.

How good is Argon at finding security vulnerabilities, specifically?

It ties for first place on CWE-bench v1, a benchmark measuring a model's ability to find and patch software vulnerabilities, with a top score of 68%. Google's security partner Wiz reportedly used Argon to uncover critical vulnerabilities in healthcare software that other frontier AI models had missed.

What else is Argon good at besides cybersecurity?

It achieves state-of-the-art performance on DeepSWE v1.1, a real-world software engineering benchmark, leads the Vals Index covering finance, legal, and tax work, scores 91.7% on long-video understanding, and ranks first on AutomationBench, a benchmark for completing multi-step automated tasks.

What's technically different about Argon compared to earlier Gemini models?

Argon supports an industry-leading 1 million token output limit, up from 64,000 tokens in prior models, letting it work through much deeper, longer problem-solving sequences in a single continuous session rather than needing to be split into smaller chunks.

How much does Argon cost to use?

Introductory pricing is $2 per million input tokens and $10 per million output tokens, with a 95% discount on cached inputs. Standard pricing after the introductory period rises to $4 and $20 per million tokens respectively.

What safety measures is Google adding before wider release?

Google says it's strengthening four areas: preventing misuse while still allowing legitimate dual-use security research, defending against indirect prompt-injection attacks, monitoring the model's reasoning process to catch and stop misalignment, and hardening isolated sandbox environments used for high-risk training and evaluation.

When will regular users be able to use Gemini 4 Argon?

Google said it plans to make Argon available to developers, enterprises, and consumers 'as soon as possible,' starting with paid API customers and Google AI Ultra subscribers, though no specific public release date has been announced.

Is this Google reclaiming a competitive lead over OpenAI and Anthropic?

Some coverage framed Argon's benchmark results as Google retaking a performance lead over rival frontier models from OpenAI and Anthropic, though Google's own announcement focused primarily on Argon's specific coding, enterprise, and cybersecurity capabilities rather than direct competitive comparisons.

SourcesGoogle
Topics and tagsGoogle, Cybersecurity, google, gemini

The daily newsletter

Tech news you'll actually get.

One short email a day. Five minutes. Plain words. The daily email is launching soon. Join the early list.

Free. Early list: we'll email you when the first issue goes out.

More in brief