HeyMark
Content

News Google

Google announces Gemini 4 Argon: benchmarks, pricing and access

5 min read
Google's official cover: Gemini 4 Argon and a luminous number four on a blue background.

Google announces Gemini 4 Argon. Compare benchmarks against GPT-6 Astra and Claude Opus 5.5, introductory API pricing and who gets access first.

Google announced Gemini 4 Argon on September 30. Its new model targets long professional workflows, coding and defensive cybersecurity. Access starts with participants in Fairwind, Google’s program for cyber defenders. Public availability will expand in stages.

The announcement comes a day after OpenAI DevDay 2026 and two days after Claude Sonnet 5.5. To compare these releases, look at the task, the cost of completing it and who can use the models today.

Gemini 4 Argon vs. GPT-6 Astra and Claude Opus 5.5

These charts reproduce six evaluations from the launch table. Switch between tests to compare results on a consistent scale from 0 to 100. The “Leading” label marks the highest score among the four models shown.

30.09.2026 · Google DeepMind

Choose the task. Compare the models.

Argon

AutomationBench

Higher is better

Completing business workflows from start to finish. Score, %.

Gemini 4 ArgonLeading
51.3%
GPT-6 Astra
41.4%
Claude Fable 5.1
31.4%
Claude Opus 5.5
42.5%

Argon scores 8.8 percentage points above Opus 5.5 on this evaluation.

DeepSWE v1.1

Higher is better

Long-horizon software engineering tasks. Score, %.

Gemini 4 ArgonLeading
77.9%
GPT-6 Astra
74.1%
Claude Fable 5.1
67.4%
Claude Opus 5.5
74.2%

Argon leads the four models in Google's table. The gap to Opus 5.5 is 3.7 points.

Terminal-Bench 4.0

Higher is better

Agent tasks in a terminal. Score, %.

Gemini 4 Argon
57.4%
GPT-6 Astra
58.2%
Claude Fable 5.1
57.9%
Claude Opus 5.5Leading
66.4%

Opus 5.5 leads here. Argon trails by 9 points. A model can lead one benchmark and trail another.

FrontierSWE v2

Higher is better

A separate software engineering evaluation. Score, %.

Gemini 4 Argon
55.0%
GPT-6 AstraLeading
65.5%
Claude Fable 5.1
56.3%
Claude Opus 5.5
62.3%

Astra has the highest score in this group. This qualifies Argon's lead on DeepSWE.

LVBench

Higher is better

Long-video understanding. Score, %.

Gemini 4 ArgonLeading
91.7%
GPT-6 Astra
87.5%
Claude Fable 5.1
79.7%
Claude Opus 5.5
83.7%

Google uses 1 frame per second for Gemini, 800 frames for Astra, 300 for Fable and 600 for Opus due to API limits. Measures understanding, not video generation.

CWE-bench v1

Higher is better

Remediating software vulnerabilities. Score, %.

Gemini 4 ArgonLeading
68.0%
GPT-6 AstraLeading
68.0%
Claude Fable 5.1
58.0%
Claude Opus 5.5
67.0%

Argon and Astra tie at 68%. This test measures vulnerability remediation, rather than all cybersecurity tasks.

The selection includes wins, ties and tests where another model comes first. Those differences help you decide what to evaluate with your own materials. A strong coding score tells you little about whether a model will preserve your brand’s voice in a carousel draft.

Google provides the evaluation configurations and methodology. Changing an agent’s tools, reasoning budget or task set can change its score. We retain the values in Google’s table, without mixing them with other providers’ runs.

What the independent Vals AI evaluation shows

The Vals Index table places Argon first at 68.90%, followed by Sonnet 5.5 at 67.04% and Opus 5.5 at 66.97%. GPT-6 Astra reaches 63.13% and GPT-6.1 Sol reaches 61.15% in the same table, checked on September 30.

Vals combines finance, coding, legal and tax tasks using economic weights. Its index helps compare professional work; it does not evaluate marketing campaigns. The page’s written conclusions still describe Sonnet and Opus as the leaders. We use the current table for these figures, which already includes Argon.

One million output tokens: what changes

Google raises the output ceiling from 64K to one million tokens. That gives a long execution more room. Keep this separate from the context window, which describes what a model can receive.

More room for a long responseOutput token limit announced by Google
Previous limit
64K
Gemini 4 Argon
1M

This is an output ceiling. Actual usage and cost depend on how many tokens the task consumes.

For an agent researching a problem, writing code and checking changes, more room may let it continue without hitting that output limit. A long task can also consume many tokens. Once access makes testing possible, we suggest measuring tokens and calls per deliverable alongside human review time.

Gemini 4 Argon pricing: the introductory rate changes later

Google announced these API rates, in US dollars per million tokens:

API · USD per million tokens
TokensIntroductoryAfterwards
InputUS$2US$4
OutputUS$10US$20

Google has not announced when the introductory period ends.

The announcement also offers cached input tokens at a 95% discount to the input rate. It does not specify when the introductory period ends.

A calculated example: 10,000 input tokens and 2,000 output tokens cost US$0.04 at the initial rate or US$0.08 at the later rate. This excludes tools and caching; research involving multiple calls uses more. A one-million-token output ceiling does not mean every response will use it all.

When will Gemini 4 Argon be available?

Google is starting with cyber defenders and trusted testers. Its announcement says expansion will begin with paid API customers and Google AI Ultra subscribers. It gives no date for general availability.

Access affects the practical choice: check your account and current documentation before preparing an integration or changing an automation’s model.

What a marketing team could test

Once access arrives, we suggest a short test with a deliverable you can check:

  • A report: supply post metrics from the same account and request conclusions with calculations and sources.
  • Research: request a trend summary with links, dates and a clear distinction between available releases and future promises.
  • A draft: provide a brief and brand examples; check which facts it preserves and which sections you need to rewrite.

HeyMark lets you work with your accounts’ posts and metrics, and bring those data to Claude or ChatGPT. These materials let you assess a deliverable against your brand’s facts. We have not tested Argon with HeyMark or announced a Gemini integration.

The launch results give us reasons to try it. A decision to use it at work will depend on deliverables you can review, their cost and actual access.

Source

Google DeepMind · Gemini 4 Argon

Plan your next post in HeyMark.

Keep the idea, review the draft with your team, and see how it performed in the accounts you connected.