Wain AI/Tech Blog

AI news and trends worldwide, updated nearly every day

Google Announces Gemini 4 Argon - Prices Are Published, but Fairwind Defenders Get It First

Image: Google official blog

Google Announces Gemini 4 Argon - Prices Are Published, but Fairwind Defenders Get It First

Google announced a new model, Gemini 4 Argon, on September 30, 2026. Introductory pricing is $2 input and $10 output per million tokens, and the output token ceiling moves from 64K to 1M. Access starts with trusted cyber defenders in the Fairwind Program, with no date given for developers, enterprises or consumers.

On September 30, 2026, Google announced Gemini 4 Argon, which it presents as its new frontier model. The post is bylined by Koray Kavukcuoglu, SVP at Google DeepMind and Chief AI Architect at Google, and Google’s claim for the model is frontier-level results on demanding workloads in three areas: software engineering done on real codebases, enterprise knowledge work such as legal and finance, and defensive cybersecurity1.

What you cannot do is use it. At announcement, the model has gone only to part of the group of vetted organisations inside the Fairwind Program, Google’s security initiative — the ones it calls trusted cyber defenders. Developers, enterprises and consumers are to be brought in by stages while Google takes part in what it calls “the U.S. government’s voluntary process for pre-release model access”, and the company says that putting frontier capability of this level into the world safely calls for exactly that kind of staging1.

The price list arrived before the model did

Numbers are in the post. During the introductory period a million input tokens cost $2 and a million output tokens cost $10, and cached input is discounted 95 percent against the input rate. A footnote records what happens afterwards: $4 and $20 respectively1.

Those first two figures match what OpenAI charges for GPT-6.1 Sol in standard processing on short context — also $2 and $102. The match stops at the table, though. Google’s post says nothing about whether Argon’s rate varies by processing tier or by context length, so there is no way to tell whether the same job would produce the same bill.

The other thing missing is a date. When Google priced Gemini 3.8 Flash, the footnote named the day the introductory rate ended and the rate that applied from the next day onward. Here there is a post-introductory rate with no calendar attached to it. For anyone building an annual API budget, neither the date you can start using this nor the date it gets more expensive is currently knowable.

Output ceiling: 64K to 1M

On the specification side the striking change is on the output end. Google says it has raised the output token ceiling from the previous 64K to 1M and calls that figure industry-leading. Its stated reasoning is that room to produce several hundred thousand tokens inside a single trajectory adds depth to the reasoning and lets a hard problem be settled in one pass1.

For reference, GPT-6.1 Sol — the model at the matching price — tops out at 128,000 output tokens3. If what you have in mind is a long uninterrupted generation, the kind of large migration or automated refactor where stopping halfway costs you the context, the size of that budget is a design input. Again: once it is something you can call.

Every benchmark figure here is Google’s own

Google reports the following, all of it as published by Google. On DeepSWE v1.1, which looks at long-horizon software engineering, 77.9 percent. On AutomationBench, a Zapier benchmark for running core business functions end to end, 51.3 percent and first place. On LVBench, for understanding long video, 91.7 percent. Google also says Argon leads the Vals Index, a composite that rates work in finance, coding, law and tax by its economic impact, with each sector weighted by how much it contributes to U.S. GDP. It claims comparable leadership on Vals Finance Agent v2 and on Harvey’s Legal Agent Benchmark, but publishes no scores for those two1.

On the security side, Google reports 68 percent on CWE-bench v1, tied for the top spot1. Note that Google itself distinguishes the versions: the figures it cites for Gemini 3.8 Flash Cyber are on CWE-bench v0. The gap is not readable as generational improvement.

Some recipients get it with the guardrails off

How the model is handed out has moved on from September’s Flash Cyber. Google states that trusted defenders and its own internal teams will receive Argon with its cyber guardrails removed, so that the full defensive capability is available to them1. For Flash Cyber, the wording was that the model shipped with looser mitigations on the cybersecurity side.

The worked example is Wiz, which Google says is running Argon inside Scan for Good, an initiative that protects critical public infrastructure at no charge. By Google’s account the model turned up a serious vulnerability leaking personal information through healthcare software in hospital use worldwide, one that earlier frontier models had not caught. Google positions Argon as able to locate, confirm and patch serious software vulnerabilities on its own1.

Not distributing a highly capable model to everyone is a pattern that also showed up in early September, when OpenAI determined that Astra had crossed the “Critical” threshold for cybersecurity. The two companies reached that point by different reasoning and different procedures.

What the monitoring finds does not go back into training

Of the four areas of safeguards Google lists, two are worth reading if you operate agents yourself.

The first is monitoring for misalignment. Google says it has mitigations in place that watch the chain of thought and the actions, halting execution where that is warranted, and that a comparable setup watched its training runs and paged a dedicated incident response team. It then says it was careful not to feed what the monitoring found back into training, the reason given being the risk of shaping Argon’s reasoning toward evading that monitoring1. Recycling monitoring logs into the next round of training data is an obvious move; the argument here is that doing it teaches the evasion. Google also presses other labs to keep reasoning transparent.

The second is resistance to indirect prompt injection. Google describes Argon as the most resistant model it has built to date and says it leads on Gray Swan’s Indirect Prompt Injection (IPI) benchmark. On the misuse side, it says it has improved techniques that watch the model’s internal activations for signs of abuse, and that red teams inside and outside the company tested how well the safeguards hold using a mix of manual and automated attacks1.

Not a candidate yet

Google says the broad rollout begins with customers who pay for the API, and with subscribers to Google AI Ultra. On timing, the post offers nothing more precise than as soon as possible1.

In practical model-selection terms, Argon does not belong on the shortlist for now. The pricing and the output budget tell you what Google was aiming at; they are not numbers you can quote in an estimate. What actually matters is that a model you reach only through an application and a review has now appeared on the flagship frontier line rather than in a cybersecurity-specific spin-off. The more your security work already runs on models, the sooner you need an answer to whether you can apply at all, and to which internal teams access would be confined.

Sources

  1. Gemini 4 Argon: our next era of frontier intelligence - Google official blog (September 30, 2026)
  2. Pricing - OpenAI official documentation (used for the price comparison, accessed October 2, 2026)
  3. GPT-6.1 Sol - OpenAI official documentation (used for the max output token comparison, accessed October 2, 2026)

We publish the latest AI news nearly every day.

Subscribe via RSS Get new posts the moment they go live.

Search other keywords →