Gemini 4 Argon is Google's new frontier model, announced on September 30, 2026. It raises Gemini's output limit from 64K to 1 million tokens, scores 77.9% on the DeepSWE v1.1 long-horizon coding benchmark, and will cost $2 per million input tokens and $10 per million output tokens at introductory rates. Today only trusted cyber defenders in Google's Fairwind Program can use it. Paid API customers and Google AI Ultra subscribers are next, with no date given.
Our read: Argon is the strongest published option for long-running knowledge work and very long outputs. On terminal and computer-use agent work it still trails GPT-6 Astra and Claude Opus 5.5. Pick it per workload, not as a new default.
1M
Output tokens, up from 64K
77.9%
DeepSWE v1.1, vendor-reported
$2 / $10
Intro price per 1M in / out
$4 / $20
Standard price after intro
What is Gemini 4 Argon?
Argon is the first model in Google's Gemini 4 series. Google built it for deep, multi-step reasoning across long tasks, and highlights three areas: real-world software engineering, enterprise knowledge work in law, finance, and tax, and defensive cybersecurity.
It arrived one day after OpenAI announced its always-on Dots agents. The two launches point the same way. Labs now compete on how long a model can work unattended, not only on how well it answers one question.
Google has not disclosed Argon's input context window, a public model ID, or the length of the introductory pricing period. Plan for all three to change before general availability.
Gemini 4 Argon benchmarks: where it leads and where it trails
Every number below comes from Google's own comparison table. None has been independently reproduced yet, so treat them as vendor-reported.
| Benchmark | Argon | GPT-6 Astra | Opus 5.5 |
|---|---|---|---|
| DeepSWE v1.1 (long-horizon coding) | 77.9% | 74.1% | 74.2% |
| Vals Index (knowledge work) | 68.9% | 63.1% | 67.0% |
| Vals Finance Agent v2 | 65.4% | 53.5% | 58.6% |
| AutomationBench | 51.3% | 41.4% | 42.5% |
| GraphWalks 256K to 1M (long context) | 84.2% | 71.8% | 66.8% |
| FrontierSWE v2 (agentic coding) | 55.0% | 65.5% | 62.3% |
| Terminal-Bench 4.0 | 57.4% | 58.2% | 66.4% |
| OSWorld-2.0 (computer use, offline) | 69.2% | 72.6% | Not reported |
The pattern is clear. Argon leads on knowledge work, long context, and DeepSWE. It trails by about 10 points on FrontierSWE v2 and by 9 points on Terminal-Bench 4.0. A coding agent that lives in a shell is the weakest fit in this table.
Two benchmarks that measure similar things also disagree. Argon wins DeepSWE by 3.7 points and loses FrontierSWE by 10.5. That gap is a reason to test on your own repository instead of choosing from a leaderboard. Our guide to reading coding-agent benchmark results explains why harness and task budget move these scores.
The only named outside user so far is Wiz. Google says Wiz used Argon in its Scan for Good work and found a critical flaw in hospital software that earlier models missed. That is still Google's account of a partner result, not an independent evaluation.
What 1 million output tokens changes
Earlier Gemini models stopped at 64K output tokens. A cap like that forces long jobs, such as migrating a module or drafting a full contract set, into many calls that each lose some context. Argon can carry one reasoning trajectory and one artifact through a single response.
Google says it uses this internally for C and C++ to Rust migrations, including more than 800,000 lines in the Fuchsia Zircon kernel. It also reports a Rust port of the libgav1 decoder that runs 2.7 times faster. These are Google's own results on Google's own code.
The catch is that one call can now be very expensive and very slow. A single response that uses the full output budget costs $10 at the introductory rate and $20 at the standard rate. If an agent loops, that cost repeats. Set an explicit maximum output per call, and checkpoint long jobs so one failure does not throw away hours of work.
What a long agent run costs on Argon
Per-token prices only matter once you know the shape of a run. We modeled a 60-turn agent session over a 150,000-token working context with a 90% cache hit rate and 3,000 output tokens per turn:
# USD per 1M tokens, from published list prices on 2026-10-01.
PRICES = {
"argon-intro": {"input": 2.00, "cached": 0.10, "output": 10.00},
"argon-standard": {"input": 4.00, "cached": 0.20, "output": 20.00},
"gpt-6-astra": {"input": 10.00, "cached": 1.00, "output": 50.00},
}
def run_cost(price, uncached_in, cached_in, out):
return (uncached_in * price["input"]
+ cached_in * price["cached"]
+ out * price["output"]) / 1_000_000
# One long-horizon agent run: 60 turns over a 150K-token working context.
turns, context, cache_hit, out_per_turn = 60, 150_000, 0.90, 3_000
cached_in = int(turns * context * cache_hit)
uncached_in = turns * context - cached_in
out = turns * out_per_turn
for name, price in PRICES.items():
print(f"{name:15} ${run_cost(price, uncached_in, cached_in, out):6.2f}")
Output:
argon-intro $ 4.41
argon-standard $ 8.82
gpt-6-astra $ 26.10
On list prices, Argon at standard rates costs about a third of GPT-6 Astra for the same token counts. The script leaves out cache writes and storage. Google has not priced those for Argon, and OpenAI charges $12.50 per million cache-write tokens on Astra, so Astra's real figure is higher. It also holds token counts constant, which real runs never do. Models tokenize differently, and a model that writes longer answers or needs more turns can erase a per-token discount. Compare cost per finished task, not cost per token. Our GPT-6 Astra 272K token limit guide shows how a single threshold can double a bill.
Who can use Gemini 4 Argon today
Google is releasing Argon in phases:
- Now: trusted cyber defenders in the Fairwind Program, plus Google's internal teams. Fairwind participants get the model without its cyber guardrails.
- Next: paid Gemini API customers and Google AI Ultra subscribers, "as soon as possible," with no date.
- Later: broader developer, enterprise, and consumer access, after Google iterates on guardrails with early testers.
Google also says Argon went through the U.S. government's voluntary pre-release model access process. Every other customer will get a version with cyber guardrails on. If your product depends on security-adjacent behavior, such as code scanning, expect refusals the Fairwind version does not have.
How to prepare before the API opens
Do the work now that does not depend on access:
- Build a held-out task set. Pick 20 to 50 real tasks from your backlog with automatic pass or fail checks. Use tasks the model cannot have seen.
- Record a baseline. Run your current model on that set and capture pass rate, cost per task, and wall-clock time.
- Abstract the model call. Keep model IDs and price tables in config, since Argon's ID and standard pricing date are not public yet.
- Add output caps and checkpoints. A 1M-token ceiling needs a budget per call and a resume point per job.
When the API opens, you can run the same set on day one and decide from your own numbers.
When Gemini 4 Argon is not worth using
Skip Argon for terminal-heavy coding agents and computer-use agents, where Google's own table shows GPT-6 Astra or Opus 5.5 ahead. Skip it for short, high-volume calls such as classification or extraction, where a Flash-class model is far cheaper. And skip it for anything on a fixed launch date, because access has no date and the introductory price will double.
Use it when the job is long, document-heavy, or needs one very long output: contract and filing analysis, financial research, large code migrations, or long video understanding.
FAQ
When will Gemini 4 Argon be available in the API?
Google has not given a date. As of October 1, 2026, Argon is available only to trusted cyber defenders in the Fairwind Program and to Google's internal teams. Google says paid API customers and Google AI Ultra subscribers come next, "as soon as possible," followed by broader developer, enterprise, and consumer access.
How much does Gemini 4 Argon cost?
Introductory API pricing is $2 per million input tokens and $10 per million output tokens, with cached input 95% cheaper. After the introductory period, pricing rises to $4 input and $20 output per million tokens. Google has not said how long the introductory period lasts, so budget at the standard rate.
Is Gemini 4 Argon better than GPT-6 Astra?
It depends on the task. Google's vendor-reported table shows Argon ahead on DeepSWE v1.1, the Vals Index, finance and legal agent benchmarks, and long-context tests. GPT-6 Astra leads on FrontierSWE v2 and OSWorld-2.0. Neither result has been independently reproduced, so test both models on your own tasks.
