Artificial Analysis needed 110M output tokens to run its evaluation of Gemini 4 Argon. The median for comparable models is 82M. Since output is the expensive side of the bill, that gap matters more than the headline rate of $2 per 1M input tokens and $10 per 1M output tokens that Google announced on September 30, 2026.
I'm writing a day after the announcement, working from public data and without access to the model. Right now, only the cyber defenders in Google's Fairwind Program can use it. Disclaimer: my analysis is based on publicly available data, so treat every number below as a starting point for your own testing.
The Sticker Price Is Not the Task Price
The introductory rate is a launch promotion, and the list rate is double.
| Token type | Introductory (per 1M) | After the intro period (per 1M) |
|---|---|---|
| Input | $2 | $4 |
| Output | $10 | $20 |
The rate is only half of the equation. The other half is how many tokens the model burns to solve a problem. At 110M versus an 82M median, my rough calculation puts Argon about a third more verbose than its peers, and that surplus is billed at the highest rate on the card.
The numbers speak for themselves: Argon ranks #8 of 223 models on intelligence and #77 of 223 on cost-effectiveness. Artificial Analysis lists an average cost per task of $1.99 on its Intelligence Index. The source doesn't say whether that figure uses the introductory or the list price, so I can't confirm which scenario it reflects. If it's the introductory one, doubling the rates pushes it up.
There is one clear exception. Google offers a 95% discount on cached input tokens. For RAG systems and agents that resend long, repetitive prompts, input dominates the bill and that discount changes the math. A pipeline that mostly reads fares far better than an agent that writes long code. What decides it is the ratio of cached input to output in your own workload.
Benchmarks: Where Argon Wins, Where It Doesn't
| Benchmark | Argon | Comparison | Source |
|---|---|---|---|
| DeepSWE v1.1 | 77.9% | Claude Opus 5.5: 74.2%; GPT-6 Astra: 74.1% | Google (via Yahoo Finance) |
| AutomationBench | 51.3% | #1 | |
| LVBench | 91.7% | No rival published in the sources reviewed | |
| CWE-bench v1 | 68% | Tied for first | |
| FrontierSWE v2 | 55.0% | GPT-6 Astra: 65.5% | TBreak |
| Terminal-Bench 4.0 | 57.4% | Claude Opus 5.5: 66.4% | TBreak |
Argon leads DeepSWE v1.1 by almost 4 points over Opus 5.5. It trails by 10.5 points on FrontierSWE v2 and by 9 on Terminal-Bench 4.0. A model that wins one software engineering test and drops two others in the same field isn't "the best at coding." It's the best at one particular definition of coding. TBreak and ExplainX report the losses that Google's own post doesn't highlight.
Independent measurement is kinder to the overall picture. Artificial Analysis scores Argon (High) at 53 on its Intelligence Index, #8 of 223 models, against a peer median of 26. One caveat remains: the figures labeled Google come from the vendor, and Yahoo Finance stresses that none of them has been independently verified. For a primer on reading vendor benchmark claims with some distance, see our breakdown of GPT-5.2 and the ARC-AGI benchmark.
Fairwind, Missing Dates and the 1M Output Cap
Koray Kavukcuoglu, SVP at Google DeepMind and Chief AI Architect, signed the official announcement. Argon rolls out first to a group of "trusted cyber defenders" through the Fairwind Program, and they get it without cybersecurity guardrails. Wiz already uses it in its Scan for Good initiative to protect critical public infrastructure for free. Google says access will widen to developers, enterprises and consumers, starting with paid API customers and Google AI Ultra subscribers.
What the post never states is a date. It doesn't say when the introductory pricing ends, and it doesn't say what criteria open the Fairwind door. None of the secondary coverage I checked this morning fills that gap either. It's frustrating when a vendor publishes a rate that doubles with no end date attached: you can't close a quarterly budget on that, and an end date is the least a launch page should carry.
Can you use Argon today? Unless you're an accepted Fairwind defender, no. ExplainX advises keeping production on current models such as Gemini 3.8 Flash until a versioned endpoint exists.
The "1M" in the announcement is also easy to misread. It is the cap on OUTPUT tokens, up from 64K. That's a maximum capacity, not the typical length of a response, and it isn't a context window: Artificial Analysis lists 1M tokens of context separately. Two different parameters share one number. The cap enables code migrations or long reasoning chains in a single inference, and it also enables very long invoices on a model that is already verbose. For a view of where Google stands against its rivals, see our Gemini 3 vs ChatGPT comparison.
The Bottom Line: A Budget Test You Can Run
In my decade of analyzing productivity tools, one rule has held: never buy on the promotional price. Model your budget at $4/$20, not $2/$10. If the use case doesn't pencil out at list price, it doesn't pencil out.
Here's what this actually means for your team once access opens. Argon pays off only if its accuracy gain on your own evaluations covers both the doubled rate and the extra third of output. Measure output tokens per task with your own prompts and compare them with your current model instead of trusting the Artificial Analysis median. Check your cached-input share, because the 95% discount only matters when that share is high. If you do software engineering work, put Terminal-Bench and FrontierSWE in your test suite, since that's where Google doesn't lead. And hold off on moving production until a versioned endpoint ships.
If I had to bet, Argon is a serious candidate on intelligence and a doubtful one on efficiency. Until Fairwind stops being the only door, the productive move is to build your evaluations now and budget at list price. A model that can't show its cost on a fixed date shouldn't expect to be trusted with a fixed budget.




