News

Gemini 4 Argon is Google's most capable model yet, and almost nobody can use it

Google announced Gemini 4 Argon on 30 September 2026, its first model under the Gemini 4 name and the one it's built specifically to find, validate, and patch security vulnerabilities on its own. It isn't available to the public. It isn't available to most Google AI subscribers either. And Google's own benchmark table tells a noticeably rosier story than the independent testing that followed it.

อัปเดตเมื่อ 4 Oct 20268 min read
Quick answer

What shipped: Gemini 4 Argon, a new frontier model from Google DeepMind, announced 30 September 2026. Google says it delivers frontier-level performance on real-world software engineering, enterprise knowledge work like legal and finance, and especially cybersecurity defense, including autonomously finding, validating, and patching vulnerabilities.

Who can actually use it: Almost no one, yet. Argon is rolling out first to a vetted set of cyber defenders through Google's Fairwind Program. Google says it will reach developers, enterprises, and consumers "as soon as possible," starting with paid API customers and Google AI Ultra subscribers, with no committed date. Gemini Notebook and the free Gemini app aren't mentioned anywhere in the announcement.

Pricing, where it applies: $2 per million input tokens and $10 per million output tokens at launch, rising later to $4/$20, with cached input discounted 95%. The output token limit jumps to 1,000,000, up from 64,000 on Google's prior Gemini 3-generation models.

The gap most launch-day coverage skipped: Google's own comparison table shows Argon ahead of Claude Opus 5.5 and GPT-6 Astra on most of the benchmarks it chose to publish. Artificial Analysis's independent testing, using its own standardized methodology, puts Argon in a tie with GPT-6 Astra and a step behind Opus 5.5 overall, though the per-benchmark gaps are narrow, not dramatic. Argon is still the cheapest of the three to run per completed task.

Google's official blog post 'Gemini 4 Argon: our next era of frontier intelligence,' dated Sep 30, 2026, bylined Koray Kavukcuoglu, SVP Google DeepMind and Chief AI Architect, Google.
blog.google, September 2026.

What Argon actually is

Argon is the first model Google has released under the Gemini 4 name. No Gemini 4 Pro or Gemini 4 Flash exists yet, so for now Argon is Gemini 4, a single release rather than a full generation rollout the way Gemini 3 launched with multiple tiers at once. Google's own announcement leans hard on three use cases: real-world software engineering, enterprise knowledge work such as legal and financial analysis, and cybersecurity defense, where the company says Argon can autonomously locate a vulnerability, confirm it's real, and write a working patch without a human directing each step.

The technical change most likely to matter for agentic or long-running tasks is the output token limit: Google raised it to 1,000,000 tokens, up from 64,000 on its prior Gemini 3-generation models. That's a 15x jump in how much a single response can contain, which matters specifically for tasks like patch generation or long document drafting that need to reason continuously without restarting a new call partway through.

The Fairwind rollout, and why consumers aren't in it yet

Argon's cyber-focused capabilities ship first through Google's Fairwind Program, which gives a vetted group, government cyber-defense authorities and other trusted organizations, access ahead of everyone else, as part of the voluntary pre-release access process the US government has set up for frontier AI. Fairwind participants also get a guardrail-free version of Argon, paired with what Google describes as misalignment mitigations that monitor the model's chain-of-thought for signs it's reasoning its way around its own restrictions. Access is restricted to internal cybersecurity, incident-response, and penetration-testing staff, and participating organizations are required to enforce multi-factor authentication. Google frames the whole arrangement under its Frontier Safety Framework: a dual-use model capable of finding exploitable bugs is also a model capable of writing them, so broader access comes with broader risk.

For everyone else, Google's own language is the only public timeline that exists: Argon will reach "developers, enterprises, and consumers, starting with paid API customers and Google AI Ultra subscribers," with no date attached. There's no mention of the free Gemini app tier, no mention of Google AI Plus or Pro subscribers, and no mention of Gemini Notebook anywhere in the announcement, the pricing page, or the coverage that followed it. If you're not a paid API customer or an AI Ultra subscriber, there's currently no stated path to trying Argon at all.

Pricing: introductory now, double later

Where Argon is available through the API, it's priced at $2 per million input tokens and $10 per million output tokens, introductory pricing Google says will rise to $4/$20 later. Cached input is discounted 95%, which matters a lot for workflows that repeatedly reprocess the same large codebase or document set, the kind of task Argon is specifically positioned for. Google hasn't published a date for when the introductory pricing ends.

Google says it wins. Independent testing says it's closer to a toss-up

Google's launch post backs Argon with a wide benchmark table: 77.9% on DeepSWE v1.1 (long-horizon software engineering), 91.7% on LVBench (long video understanding), 68% on CWE-Bench v1 (vulnerability detection), and a first-place 51.3% on AutomationBench, Zapier's benchmark for end-to-end business task execution. By Google's own count, Argon leads on roughly 12 to 13 of the 18 to 19 benchmarks in its comparison table against Claude Opus 5.5 and GPT-6 Astra.

Artificial Analysis, which runs its own standardized benchmark suite rather than relying on a company's self-selected table, found a less flattering picture. Its Intelligence Index, an aggregate score meant for apples-to-apples comparison across labs, put Argon at 53, GPT-6 Astra at 53 (tied with Argon), and Claude Opus 5.5 at 58, ahead of both. On Terminal-Bench 4.0 specifically, a benchmark for operating a real computer through a terminal, Artificial Analysis measured Argon at 57%, Astra at 59%, and Opus 5.5 at 60%, a narrow spread rather than a blowout. The pattern holds across both measures: a model Google's own table shows winning most head-to-head comparisons lands in a tie with one rival and a step behind the other once a neutral party runs the tests, even if the gap on any single benchmark isn't dramatic.

Metric                         Argon      GPT-6 Astra   Opus 5.5
----------------------------------------------------------------------
Google-published wins          12-13 of 18-19 benchmarks (self-reported)
AA Intelligence Index            53            53            58
Terminal-Bench 4.0               57%           59%           60%
AA cost per completed task      $1.99         $3.26         $5.98

Sources: Google's launch benchmark table (self-reported, tally
not independently reproduced); Artificial Analysis independent
testing (neutral methodology). Argon wins on Google's own table
and on cost. On AA's neutral Intelligence Index and Terminal-Bench
4.0, it's tied with Astra and a step behind Opus 5.5, though the
per-benchmark gaps are narrow, not dramatic.

Argon's clearest real advantage in the independent data is cost, not raw intelligence: Artificial Analysis measured $1.99 per completed task on its evaluation suite, against $3.26 for GPT-6 Astra and $5.98 for Claude Opus 5.5. That's a genuine win, but it comes with an asterisk. Artificial Analysis also found Argon notably verbose, averaging about 62,000 output tokens per task against roughly 27,000 for GPT-6 Astra, more than double. Argon is cheap mainly because its per-token price is low, not because it's an efficient model that reasons in fewer tokens than its rivals.

There's also an unresolved dispute worth flagging rather than ignoring: The Hacker News reported on the Fairwind access restrictions, and separately, Bloomberg reported, citing people with direct access to the effort, that Google's own engineers found Argon performs worse on real coding tasks than the benchmarks suggest, specifically citing weak front-end design work (how apps and websites look and feel). Google disputed the characterization, saying it would be "inaccurate to say Gemini 4 is underperforming in areas such as coding." Neither side's account has been independently settled, and it's the kind of gap between benchmark performance and daily-use experience worth watching for in user reports over the coming weeks rather than taking either side's word for.

Where this leaves Gemini Notebook users

If you use Gemini Notebook today, nothing about your notebook changes because of this release. Argon isn't mentioned in Google's announcement as reaching the Gemini app, Gemini Notebook, Sheets, or any consumer surface, the same gap we found when Gemini 3.8 Flash shipped three weeks earlier. Argon's rollout path (Fairwind first, then paid API and AI Ultra, with everything else undated) suggests Gemini Notebook, if it gets Gemini 4-generation output at all, is a later step rather than a near-term one.

If your actual goal is the best model output available right now rather than waiting on Argon, the independent numbers above point toward Claude Opus 5.5 for raw task quality, already broadly available in ways Argon currently isn't.

People also ask

What is Gemini 4 Argon?

Google DeepMind's first model under the Gemini 4 name, announced 30 September 2026. It's built for sustained-reasoning tasks: real-world software engineering, enterprise knowledge work like legal and financial analysis, and cybersecurity defense, including autonomously finding and patching vulnerabilities.

Can I use Gemini 4 Argon right now?

Only if you're part of Google's Fairwind Program (vetted cyber defenders), a paid API customer, or a Google AI Ultra subscriber, and even then Google hasn't given a firm availability date. There's no free-tier or Gemini Notebook access, and no stated timeline for when there will be.

How much does Gemini 4 Argon cost?

$2 per million input tokens and $10 per million output tokens at introductory pricing, rising later to $4/$20. Cached input is discounted 95%. Google hasn't said when the introductory rate ends.

Is Gemini 4 Argon better than Claude Opus 5.5 or GPT-6 Astra?

Depends whose numbers you trust. Google's own benchmark table shows Argon leading most comparisons. Artificial Analysis's independent testing puts Argon roughly tied with GPT-6 Astra and behind Claude Opus 5.5 on its aggregate Intelligence Index, though Argon is the cheapest of the three per completed task.

Does Gemini 4 Argon power Gemini Notebook?

No, and there's no indication it will soon. Google's announcement doesn't mention Gemini Notebook or the free Gemini app at all; the stated rollout path goes Fairwind Program first, then paid API customers and Google AI Ultra subscribers, with everything else undated.

ส่งออก NotebookLM ของคุณได้ในคลิกเดียว

ส่วนขยาย Chrome ฟรี รองรับ PDF, Word และ Markdown ประมวลผลบนเครื่องของคุณเอง — ไม่มีการอัปโหลด

อ่านต่อ