GPT Image 2.5 is live — OpenAI's newest image model, targeted edits that leave the rest of the frame alone
Gemini 4 Argon vs GPT-6 Astra vs Claude Opus 5.5: Benchmarks
2026/10/01

Gemini 4 Argon vs GPT-6 Astra vs Claude Opus 5.5: Benchmarks

Gemini 4 Argon vs GPT-6 Astra vs Claude Opus 5.5 on all 18 benchmarks in Google's table, plus output limits, pricing and which model to pick for each job.

Gemini 4 Argon leads Google's own benchmark table, but not everywhere. In the 19-row table Google published with the launch on September 30, 2026, Argon has the top score among itself, GPT-6 Astra and Claude Opus 5.5 on 13 rows, ties on one, and loses five: three to Astra and two to Opus 5.5[1]. The losses cluster around terminal work, harder software benchmarks and computer use, which is exactly where many developers will judge it.

This comparison puts every number from Google's table side by side, adds the specifications and prices each provider publishes, and explains the caveats in Google's methodology. Search interest after launch asked the same things: "gemini 4 argon benchmarks", "gemini 4 argon performance relative to other models", "gemini 4 vs gpt 6 astra", "gemini 4 vs claude". One practical difference sits above all the scores: Astra and Opus 5.5 are generally available, and Argon is not[1].

TL;DR

  • Overall: among the three models, Gemini 4 Argon has the best score on 13 of 19 rows in Google's table, GPT-6 Astra on 3, Claude Opus 5.5 on 2, with one tie[1].
  • Argon's clearest leads: Harvey's Legal Agent Benchmark (19.6% vs 5.4% and 3.8%), long-context GraphWalks 256K–1M (84.2% vs 71.8% and 66.8%) and LVBench long video (91.7%)[1].
  • Where Argon trails: Opus 5.5 leads Terminal-bench 4.0 (66.4% vs 57.4%) and PostTrainBench; Astra leads FrontierSWE v2, Terminal-Bench Science and OSWorld-2.0[1].
  • Output: Argon can return 1M tokens in one response. Astra and Opus 5.5 stop at 128K on their standard APIs[1][2][4].
  • Price per 1M tokens: Argon $2/$10 introductory, then $4/$20; Opus 5.5 $4/$20; Astra $10/$50[1][2][3].

Gemini 4 Argon benchmarks: the full table

These are the numbers from the benchmark table in Google's announcement. Google's table also includes Claude Fable 5.1, shown here in its own column so the "vs Fable" question has an answer too. The best score among Argon, Astra and Opus 5.5 is in bold; a dash means Google reported no result[1].

AreaBenchmarkGemini 4 ArgonGPT-6 AstraClaude Opus 5.5Claude Fable 5.1
Knowledge workVals Index68.9%63.1%67.0%65.8%
Knowledge workAutomationBench51.3%41.4%42.5%31.4%
Knowledge workVals Finance Agent v265.4%53.5%58.6%58.9%
Knowledge workHarvey's Legal Agent Benchmark19.6%5.4%3.8%6.7%
Agentic codingDeepSWE v1.177.9%74.1%74.2%67.4%
Agentic codingFrontierSWE v255.0%65.5%62.3%56.3%
Agentic codingVibe Code Bench91.9%89.6%90.3%90.3%
Agentic codingTerminal-bench 4.057.4%58.2%66.4%57.9%
ML engineeringPostTrainBench45.3%44.3%49.3%40.2%
Science and mathTerminal-Bench Science 0.157.6%68.1%63.3%52.6%
Science and mathLABBench 288.8%85.4%73.1%68.6%
Science and mathRiemannBench76.0%72.0%69.6%65.6%
Long contextGraphWalks, up to 128K99.7%98.7%90.6%91.4%
Long contextGraphWalks, 256K to 1M84.2%71.8%66.8%65.0%
Computer useAgent's Last Exam39.5%34.2%38.2%—
Computer useOSWorld-2.0 (offline subset)69.2%72.6%——
MultimodalChartography71.6%71.0%66.3%46.2%
MultimodalLVBench91.7%87.5%83.7%79.7%
CybersecurityCWE-bench v168.0%68.0%67.0%58.0%

Head to head with Claude Opus 5.5 alone, Gemini 4 Argon is ahead on 14 of the 18 rows where both have a score and behind on four: FrontierSWE v2, Terminal-bench 4.0, PostTrainBench and Terminal-Bench Science[1].

Read the table with Google's methodology in hand

Every score in that table comes from Google's announcement, and Google's methodology page explains where each number came from. Three details change how much weight the gaps deserve[5]:

  1. Most rival scores are self-reported. Google says results for non-Gemini models are "sourced from providers' self reported numbers unless otherwise mentioned", using each rival's maximum available reasoning setting where reported.
  2. Some Argon scores are self-computed. DeepSWE v1.1 and Terminal-bench 4.0 results for Argon were run by Google, while Astra, Fable and Opus figures come from public leaderboards or system cards. That is normal practice, but it means the harness is not always identical.
  3. Inputs were not always equal. On LVBench, Google used 1 frame per second for Gemini but 800 frames for Astra, 600 for Opus 5.5 and 300 for Fable 5.1 "due to API limitations". The models did not see the same frame budget, so treat the long-video gap with some care.

None of this makes the table wrong. It does mean a 1-point lead, such as Vals Index or CWE-bench, is a tie for practical purposes, while gaps of 12 to 17 points, like Harvey's Legal Agent Benchmark or GraphWalks at 256K to 1M, are real signals.

Specs and price side by side

Benchmarks are only half the decision. Here is what each provider publishes about limits and price.

Gemini 4 ArgonGPT-6 AstraClaude Opus 5.5
AvailabilityFairwind Program partners only; developers "as soon as possible"[1]Generally available[2]Generally available[4]
Context windowNot published1,050,000 tokens[2]1M tokens[4]
Max output per response1M tokens[1]128,000 tokens[2]128K (300K via Batch API beta)[4]
Input typesText plus charts, documents and long video, per Google[1]Text, image[2]Text, image[4]
Input / output per 1M tokens$2 / $10 intro, then $4 / $20[1]$10 / $50, more above 272K input[2]$4 / $20[3]
Cached input per 1M tokens$0.10 during intro[1]$1[2]$0.20 cache hit[3]

Price per token and price per task are different things; the Gemini 4 Argon pricing breakdown works through four task shapes on all three models.

Gemini 4 Argon vs GPT-6 Astra

Argon wins most of the knowledge-work and long-context rows against Astra, and by wide margins on the legal and finance agent tests: 19.6% vs 5.4% on Harvey's benchmark and 65.4% vs 53.5% on Vals Finance Agent v2[1]. It also costs a fifth of Astra's list price during the introductory period, and two fifths after it[1][2].

Astra is the stronger pick for computer use and harder software benchmarks. It leads OSWorld-2.0 (72.6% vs 69.2%), FrontierSWE v2 (65.5% vs 55.0%) and Terminal-Bench Science (68.1% vs 57.6%)[1]. The 10.5-point gaps on FrontierSWE v2 and Terminal-Bench Science are the largest losses Argon takes in its own launch table, so teams whose work looks like those benchmarks should test before switching. Astra also edges Argon on Terminal-bench 4.0, 58.2% to 57.4%.

Gemini 4 Argon vs Claude Opus 5.5

This is the closest pairing on coding. Argon edges Opus 5.5 on DeepSWE v1.1 (77.9% vs 74.2%) and Vibe Code Bench (91.9% vs 90.3%), while Opus 5.5 wins Terminal-bench 4.0 by 9 points (66.4% vs 57.4%) and PostTrainBench by 4 (49.3% vs 45.3%)[1]. If your agents live in a shell, Opus 5.5 holds up well against the newer model.

Outside coding, Argon pulls away. The legal benchmark gap is 15.8 points, long-context GraphWalks at 256K to 1M is 17.4 points, and LABBench 2 is 15.7 points[1]. After Argon's introductory period ends, the two models list at the same $4 and $20 per million tokens[1][3], so the choice comes down to the work, not the price.

Which model to use for which job

  • Legal and finance research, long document sets, long video: Gemini 4 Argon, once you can get it. These are its widest leads.
  • Very long single outputs, such as full migrations or reports past 128K tokens: Gemini 4 Argon is the only one of the three that can do it in one response.
  • Terminal-heavy coding agents: Claude Opus 5.5, which leads Terminal-bench 4.0 by a clear margin.
  • Computer use and the hardest SWE tasks: GPT-6 Astra, which leads OSWorld-2.0 and FrontierSWE v2.
  • Anything you need to ship this month: Astra or Opus 5.5, because Gemini 4 Argon has no public API yet[1].

On reAPI, Claude Opus 5.5 and GPT-6 Astra are both live behind one API key and one chat completions endpoint, so you can run the same evaluation on both today and add Gemini 4 Argon when it arrives. The Gemini 4 Argon model page tracks its availability.

FAQ

Is Gemini 4 Argon better than GPT-6 Astra?

On most of Google's table, yes: Argon has the higher score on 14 of 19 rows against Astra, with one tie. Astra leads on FrontierSWE v2, Terminal-Bench Science, OSWorld-2.0 and, narrowly, Terminal-bench 4.0[1].

Is Gemini 4 Argon better than Claude Opus 5.5?

On 14 of the 18 comparable rows in Google's table. Opus 5.5 leads on Terminal-bench 4.0, PostTrainBench, FrontierSWE v2 and Terminal-Bench Science[1].

How does Gemini 4 Argon compare with Claude Fable 5.1?

Argon scores higher than Fable 5.1 on 15 of the 17 rows where Google reports both. Fable 5.1 is slightly ahead on FrontierSWE v2 (56.3% vs 55.0%) and Terminal-bench 4.0 (57.9% vs 57.4%)[1].

What is Gemini 4 Argon's best benchmark result?

GraphWalks up to 128K at 99.7%. The most meaningful wins are the large gaps: Harvey's Legal Agent Benchmark at 19.6% against 5.4% for Astra and 3.8% for Opus 5.5, and GraphWalks 256K to 1M at 84.2%[1].

Are these benchmark scores independent?

Not fully. They come from Google's announcement; Google says most rival scores are self-reported by each provider, and it ran some Argon tests itself[5].

Which is cheapest: Gemini 4 Argon, GPT-6 Astra or Claude Opus 5.5?

Gemini 4 Argon at its introductory $2/$10 rate. After that, Argon and Opus 5.5 both list at $4/$20, and Astra at $10/$50 per 1M tokens[1][2][3].

Choosing between Argon, Astra and Opus 5.5

Google's table makes a strong case for Gemini 4 Argon as the best general model of the three, with the biggest leads in legal, finance, long context and long video, and a 1M-token output nobody else matches. It also shows where the case is weakest: terminal and computer-use agents, where Opus 5.5 and Astra still win.

Two caveats keep this from being a final verdict. The scores come from the vendor of the model that wins most of them, and Argon is the one model in the comparison you cannot call yet. Build your evaluation on Astra and Opus 5.5 now, and rerun it when Gemini 4 Argon opens to developers.

References

  1. Google. Gemini 4 Argon: our next era of frontier intelligence (including benchmark table). Retrieved October 2026 from blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon
  2. OpenAI. GPT-6 Astra model page. Retrieved October 2026 from developers.openai.com/api/docs/models/gpt-6-astra
  3. Anthropic. Pricing. Retrieved October 2026 from platform.claude.com/docs/en/about-claude/pricing
  4. Anthropic. Models overview. Retrieved October 2026 from platform.claude.com/docs/en/about-claude/models/overview
  5. Google DeepMind. Gemini 4 Argon model evaluation: approach, methodology and results. Retrieved October 2026 from deepmind.google/models/evals-methodology/gemini-4-argon