News Google
Google announces Gemini 4 Argon: benchmarks, pricing and access

Google announces Gemini 4 Argon. Compare benchmarks against GPT-6 Astra and Claude Opus 5.5, introductory API pricing and who gets access first.
Google announced Gemini 4 Argon on September 30. Its new model targets long professional workflows, coding and defensive cybersecurity. Access starts with participants in Fairwind, Google’s program for cyber defenders. Public availability will expand in stages.
The announcement comes a day after OpenAI DevDay 2026 and two days after Claude Sonnet 5.5. To compare these releases, look at the task, the cost of completing it and who can use the models today.
Gemini 4 Argon vs. GPT-6 Astra and Claude Opus 5.5
These charts reproduce six evaluations from the launch table. Switch between tests to compare results on a consistent scale from 0 to 100. The “Leading” label marks the highest score among the four models shown.
Choose the task. Compare the models.
AutomationBench
Higher is betterCompleting business workflows from start to finish. Score, %.
- Gemini 4 ArgonLeading
- 51.3%
- GPT-6 Astra
- 41.4%
- Claude Fable 5.1
- 31.4%
- Claude Opus 5.5
- 42.5%
Argon scores 8.8 percentage points above Opus 5.5 on this evaluation.
DeepSWE v1.1
Higher is betterLong-horizon software engineering tasks. Score, %.
- Gemini 4 ArgonLeading
- 77.9%
- GPT-6 Astra
- 74.1%
- Claude Fable 5.1
- 67.4%
- Claude Opus 5.5
- 74.2%
Argon leads the four models in Google's table. The gap to Opus 5.5 is 3.7 points.
Terminal-Bench 4.0
Higher is betterAgent tasks in a terminal. Score, %.
- Gemini 4 Argon
- 57.4%
- GPT-6 Astra
- 58.2%
- Claude Fable 5.1
- 57.9%
- Claude Opus 5.5Leading
- 66.4%
Opus 5.5 leads here. Argon trails by 9 points. A model can lead one benchmark and trail another.
FrontierSWE v2
Higher is betterA separate software engineering evaluation. Score, %.
- Gemini 4 Argon
- 55.0%
- GPT-6 AstraLeading
- 65.5%
- Claude Fable 5.1
- 56.3%
- Claude Opus 5.5
- 62.3%
Astra has the highest score in this group. This qualifies Argon's lead on DeepSWE.
LVBench
Higher is betterLong-video understanding. Score, %.
- Gemini 4 ArgonLeading
- 91.7%
- GPT-6 Astra
- 87.5%
- Claude Fable 5.1
- 79.7%
- Claude Opus 5.5
- 83.7%
Google uses 1 frame per second for Gemini, 800 frames for Astra, 300 for Fable and 600 for Opus due to API limits. Measures understanding, not video generation.
CWE-bench v1
Higher is betterRemediating software vulnerabilities. Score, %.
- Gemini 4 ArgonLeading
- 68.0%
- GPT-6 AstraLeading
- 68.0%
- Claude Fable 5.1
- 58.0%
- Claude Opus 5.5
- 67.0%
Argon and Astra tie at 68%. This test measures vulnerability remediation, rather than all cybersecurity tasks.
The selection includes wins, ties and tests where another model comes first. Those differences help you decide what to evaluate with your own materials. A strong coding score tells you little about whether a model will preserve your brand’s voice in a carousel draft.
Google provides the evaluation configurations and methodology. Changing an agent’s tools, reasoning budget or task set can change its score. We retain the values in Google’s table, without mixing them with other providers’ runs.
What the independent Vals AI evaluation shows
The Vals Index table places Argon first at 68.90%, followed by Sonnet 5.5 at 67.04% and Opus 5.5 at 66.97%. GPT-6 Astra reaches 63.13% and GPT-6.1 Sol reaches 61.15% in the same table, checked on September 30.
Vals combines finance, coding, legal and tax tasks using economic weights. Its index helps compare professional work; it does not evaluate marketing campaigns. The page’s written conclusions still describe Sonnet and Opus as the leaders. We use the current table for these figures, which already includes Argon.
One million output tokens: what changes
Google raises the output ceiling from 64K to one million tokens. That gives a long execution more room. Keep this separate from the context window, which describes what a model can receive.
- Previous limit
- 64K
- Gemini 4 Argon
- 1M
This is an output ceiling. Actual usage and cost depend on how many tokens the task consumes.
For an agent researching a problem, writing code and checking changes, more room may let it continue without hitting that output limit. A long task can also consume many tokens. Once access makes testing possible, we suggest measuring tokens and calls per deliverable alongside human review time.
Gemini 4 Argon pricing: the introductory rate changes later
Google announced these API rates, in US dollars per million tokens:
| Tokens | Introductory | Afterwards |
|---|---|---|
| Input | US$2 | US$4 |
| Output | US$10 | US$20 |
Google has not announced when the introductory period ends.
The announcement also offers cached input tokens at a 95% discount to the input rate. It does not specify when the introductory period ends.
A calculated example: 10,000 input tokens and 2,000 output tokens cost US$0.04 at the initial rate or US$0.08 at the later rate. This excludes tools and caching; research involving multiple calls uses more. A one-million-token output ceiling does not mean every response will use it all.
When will Gemini 4 Argon be available?
Google is starting with cyber defenders and trusted testers. Its announcement says expansion will begin with paid API customers and Google AI Ultra subscribers. It gives no date for general availability.
Access affects the practical choice: check your account and current documentation before preparing an integration or changing an automation’s model.
What a marketing team could test
Once access arrives, we suggest a short test with a deliverable you can check:
- A report: supply post metrics from the same account and request conclusions with calculations and sources.
- Research: request a trend summary with links, dates and a clear distinction between available releases and future promises.
- A draft: provide a brief and brand examples; check which facts it preserves and which sections you need to rewrite.
HeyMark lets you work with your accounts’ posts and metrics, and bring those data to Claude or ChatGPT. These materials let you assess a deliverable against your brand’s facts. We have not tested Argon with HeyMark or announced a Gemini integration.
The launch results give us reasons to try it. A decision to use it at work will depend on deliverables you can review, their cost and actual access.
Source
Plan your next post in HeyMark.
Keep the idea, review the draft with your team, and see how it performed in the accounts you connected.