GPT Image 2.5 is live — OpenAI's newest image model, targeted edits that leave the rest of the frame alone
Gemini 4 Argon Pricing: Intro Rate, Standard Rate, Real Cost
2026/10/01

Gemini 4 Argon Pricing: Intro Rate, Standard Rate, Real Cost

Gemini 4 Argon pricing explained: the $2/$10 intro rate, the $4/$20 standard rate, cached input, and what real tasks cost next to GPT-6 Astra and Opus 5.5.

Gemini 4 Argon pricing has two numbers, and most coverage only quotes the first. Google launched the model at an introductory price of $2 per million input tokens and $10 per million output tokens, with cached input 95% cheaper than regular input[1]. A footnote in the same announcement says that "after the introductory period expires, the price of $4 per 1M input tokens and $20 per 1M output tokens will apply"[1]. Google did not say how long the introductory period lasts.

So any budget built on Gemini 4 Argon should carry both rates. This article lays out the full price card, works through four real task shapes, and compares the bill with two of the three models in Google's own benchmark table: GPT-6 Astra and Claude Opus 5.5.

TL;DR

  • Introductory rate: $2 per 1M input tokens, $10 per 1M output tokens, cached input at 95% off the input price, so $0.10 per 1M[1].
  • Standard rate after the intro period: $4 input and $20 output per 1M tokens. Google has not published the cached-input rate for this phase[1].
  • Versus GPT-6 Astra: Astra lists at $10 input and $50 output per 1M tokens, with higher rates above 272K input tokens[2]. Argon's intro rate is one fifth of that.
  • Versus Claude Opus 5.5: Opus 5.5 lists at $4 input and $20 output[3], exactly Argon's standard rate.
  • No free tier and no developer access yet: the first developer wave is "paid API customers", and Google has not given a date[1].

The full Gemini 4 Argon price card

Everything Google has published about Gemini 4 Argon pricing fits in one table. Cells marked "Not published" are things Google has not said, not zeros.

ItemIntroductory rateStandard rate (after intro)
Input, per 1M tokens$2.00$4.00
Output, per 1M tokens$10.00$20.00
Cached input, per 1M tokens$0.10 (95% off input)Not published
Long-context surchargeNot publishedNot published
Batch discountNot publishedNot published
Length of intro periodNot published

Sources: Google's announcement and its footnote[1].

Two things stand out. First, output is five times the price of input, the same ratio GPT-6 Astra and Claude Opus 5.5 use[2][3]. Second, the cache discount is steep. At $0.10 per million, re-reading a large cached codebase on every agent turn costs a twentieth of reading it fresh.

How much is 1 million tokens on Gemini 4 Argon?

It depends which direction the tokens go. One million input tokens costs $2 at the introductory rate and $4 after it. One million output tokens costs $10 at the introductory rate and $20 after it[1].

The output case matters more for Argon than for other models, because Argon can actually produce a million tokens in one response. Google raised the output limit to 1M tokens, up from 64K[1]. GPT-6 Astra stops at 128,000 output tokens[2], and Claude Opus 5.5 at 128K on its standard synchronous API, or 300K through its Batch API with a beta header[4]. A full-length Argon response therefore costs $10 in output alone at the intro rate, and $20 at the standard rate.

What real tasks cost: Argon vs GPT-6 Astra vs Claude Opus 5.5

List prices hide the shape of the work, so here are four task shapes with the arithmetic done. All figures use each provider's published standard API list price for a single request, no batch discounts. Astra's rates double for input and rise 1.5x for output once a prompt passes 272K input tokens[2], and that rule is applied where it triggers.

TaskArgon introArgon standardClaude Opus 5.5GPT-6 Astra
Code review: 100K in, 20K out$0.40$0.80$0.80$2.00
Agent turn: 100K new + 400K cached in, 50K out$0.74Not published$1.48$6.55
Long report: 200K in, 600K out$6.40$12.80Not possible in one responseNot possible in one response
Maximum output: 100K in, 1M out$10.20$20.40Not possible in one responseNot possible in one response

How the agent row works: Argon intro is 100K × $2 + 400K × $0.10 + 50K × $10 per million. Opus 5.5 uses its $0.20 cache-hit rate[3]. Astra's 500K-token prompt crosses 272K, so input becomes $20, cached input $2 and output $75 per million[2]. The table ignores the one-time cost of writing the cache, which every provider charges separately.

The pattern is simple. At the introductory rate, Gemini 4 Argon is half the price of Claude Opus 5.5 and a fifth of GPT-6 Astra on ordinary requests. After the intro period, Argon and Opus 5.5 cost the same per token, and Astra stays two and a half times more. The bottom two rows are not a price comparison at all: only Argon can return that much text in one call.

When the intro rate ends

Google says only that the $4 and $20 rates apply "after the introductory period expires"[1]. There is no date, and the developer API is not open yet, so it is possible that part of the introductory period passes before most developers can make a request.

Plan with the standard rate. If a workload only makes sense at $2 and $10, it is a workload that breaks on a date nobody has announced. If it works at $4 and $20, the introductory period is a discount on top.

What a 1M-token output budget means in practice

Long outputs are where Gemini 4 Argon's pricing gets interesting, and also where it can surprise you. Google says Argon can "generate hundreds of thousands of tokens in a single trajectory"[1]. That is the feature that let Argon agents rewrite 32K lines of SIMD code in libgav1 and work on an 800K+ line kernel migration[1].

It also means one unattended request can spend $20 at the standard rate. Three habits keep that under control:

  1. Set a max output budget on every request. Treat the 1M limit as headroom, not a default.
  2. Cache the stable context. A repository or document set that every turn rereads should sit in the cache at $0.10 per million, not in fresh input at $2.
  3. Measure output per task, not per token. A model that finishes in one long response can cost less than one that needs five shorter retries, even at a higher token price.

Where reAPI fits

Gemini 4 Argon is not available on reAPI yet, because Google has not opened developer access[1]. We will publish reAPI pricing on the Gemini 4 Argon model page when it launches here.

If you want to start building now, Claude Opus 5.5 is live on reAPI and lists at the same per-token price as Argon's standard rate on Anthropic's own platform[3]. Current reAPI rates are on each model page. For background on why output caps matter for cost, see our guide to LLM max output tokens.

FAQ

How much does Gemini 4 Argon cost?

At launch, $2 per 1M input tokens and $10 per 1M output tokens, with cached input at $0.10 per 1M. After the introductory period, $4 input and $20 output per 1M tokens[1].

Is Gemini 4 Argon free?

No. Google has announced only paid pricing, and the first developer access goes to paid API customers[1].

How long does the Gemini 4 Argon introductory price last?

Google has not said. The announcement states only that the standard rate applies after the introductory period expires[1].

Is Gemini 4 Argon cheaper than GPT-6 Astra?

Yes, on list price. Astra costs $10 input and $50 output per 1M tokens[2], five times Argon's intro rate and two and a half times its standard rate.

Is Gemini 4 Argon cheaper than Claude Opus 5.5?

During the introductory period, yes, by half. After it, both list at $4 input and $20 output per 1M tokens[1][3].

Does Gemini 4 Argon need a Google AI Ultra subscription?

Not for the API, which is priced per token. Google AI Ultra is the consumer plan Google named for the first consumer access[1].

How do I estimate a Gemini 4 Argon bill?

Multiply input tokens by the input rate, cached input tokens by the cached rate and output tokens by the output rate, each per million. Then run the same numbers at the standard rate, since the intro period has no published end.

Budgeting for Gemini 4 Argon

The headline is real: at $2 and $10, Gemini 4 Argon undercuts every model in Google's own comparison table. But it is a promotional price on a model most developers cannot call yet, with an end date nobody has published. Build the business case on $4 and $20, which puts Argon level with Claude Opus 5.5 and well under GPT-6 Astra, and count the introductory period as upside.

The bigger pricing story is output length. No other model in this comparison can bill you for a million output tokens in one response, so Gemini 4 Argon pricing is cheapest exactly where it is also most capable, and most expensive when a long run goes unchecked.

References

  1. Google. Gemini 4 Argon: our next era of frontier intelligence. Retrieved October 2026 from blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon
  2. OpenAI. GPT-6 Astra model page. Retrieved October 2026 from developers.openai.com/api/docs/models/gpt-6-astra
  3. Anthropic. Pricing. Retrieved October 2026 from platform.claude.com/docs/en/about-claude/pricing
  4. Anthropic. Models overview. Retrieved October 2026 from platform.claude.com/docs/en/about-claude/models/overview