model fatıgue
First look

Does Gemini 4 Argon really beat Opus 5.5 at half the price?

Published 2 Oct 2026

Pressing play loads the video from YouTube (Google), in privacy-enhanced mode. Privacy

This is the video written out, with every figure in full, each linked to its source in the table below.

Beats Opus at half the price?▶ 0:00

On Tuesday, 29 September, OpenAI put out GPT-6.1 Sol. On Wednesday, 30 September, Google announced Gemini 4 Argon, and nobody outside a small group can use it yet. It is Google's first big model after a summer of smaller Flash releases (Gemini 3.6 Flash on 21 July, Gemini 3.7 Flash on 13 August and Gemini 3.8 Flash on 2 September, by Artificial Analysis's dates), and Google's own table puts it ahead of Opus 5.5 at half Opus's price per token.

That table compares Argon with Anthropic's Claude Opus 5.5 on 17 benchmarks that have an Opus score, counting GraphWalks's two context lengths once, and Argon has the higher score on 13 of them. For Opus, Google reports the highest setting it could find a result for. Some launch videos made that their headline, like "BREAKING: Gemini 4 Argon Beats GPT-6 & Claude on 12 Tests (Half Opus Price)" and "Google Is Back: Gemini 4 Argon Beats Opus 5.5 at Half the Price".

We can't run Argon ourselves. A probe of Google's Vertex AI for the two Argon model ids it tried got a 404 for both ("not found or your project does not have access to it"), and Google's API model list and pricing page don't list it. So nothing on this page is our measurement. Three independent testers have published broad results of their own: Artificial Analysis, Vals AI, and the Arena, where people vote on answers. This page asks what the video asks: does Argon beat Opus, at what, and is it really half the price?

The basics▶ 0:59

Google DeepMind announced Gemini 4 Argon on 30 September. It is going first to trusted cyber defenders in Google's Fairwind program and to early testers. Google says paid API customers and Google AI Ultra subscribers come next, and gives no date.

The price is $2 per million input tokens and $10 per million output tokens, with cached input at 95% off the input price, so $0.10. That's an introductory price, and Google doesn't say how long it lasts. After it, Argon costs $4 in and $20 out, the same as Opus 5.5.

Price per million tokens

API list prices

Gemini 4 Argon, introductoryGemini 4 Argon, afterOpus 5.5GPT-6.1 Sol
input$2$4$4$2
cached input$0.10not stated$0.20$0.10
output$10$20$20$10
Google's announcement and its footnote, Anthropic's Opus 5.5 page, OpenAI's GPT-6.1 Sol post. Argon's cached price is our arithmetic from Google's 95% off the input price, which Google states only with the introductory price.

Google says Argon can write up to 1M tokens in one answer, up from 64K in earlier Gemini models. On Artificial Analysis's Intelligence Index it scores 52.6, 11.7 points more than Gemini 3.8 Flash's 40.9, which was Google's best model on that index until now.

The overall score▶ 1:53

Artificial Analysis's Intelligence Index is an independent score built from ten tests. Artificial Analysis ran Argon at its high setting, the one setting Vals and the Arena used as well, and has no other Argon row yet.

Argon scores 52.6, at $1.99 per index task. Opus 5.5 at its own high setting scores 53.6, for $1.82: a point higher, for $0.17 less. At max, Opus reaches 57.6, 5.0 points above Argon. Argon sits level with OpenAI's GPT-6 Astra at max, 52.7.

Argon at high against Opus 5.5 at every setting

Intelligence Index score against cost per task. Argon has been run at high only. The yellow lines join the two pairs the text compares.

Opus 5.5Gemini 4 Argon
405060$0.10$0.30$1$3$10Opus 5.5 at low: score 42.3, $0.55 per task · Artificial Analysis, read 1 Oct 2026, 07:05 CESTlowOpus 5.5 at medium: score 51.2, $1.34 per task · Artificial Analysis, read 1 Oct 2026, 07:05 CESTmediumOpus 5.5 at high: score 53.6, $1.82 per task · Artificial Analysis, read 1 Oct 2026, 07:05 CESThighOpus 5.5 at xhigh: score 56.0, $3.46 per task · Artificial Analysis, read 1 Oct 2026, 07:05 CESTxhighOpus 5.5 at max: score 57.6, $5.98 per task · Artificial Analysis, read 1 Oct 2026, 07:05 CESTmaxOpus 5.5Gemini 4 Argon at high: score 52.6, $1.99 per task · Artificial Analysis, read 1 Oct 2026, 07:05 CESTGemini 4 Argonscorecost per task, log scale
405060$0.10$0.30$1$3$10Opus 5.5 at low: score 42.3, $0.55 per task · Artificial Analysis, read 1 Oct 2026, 07:05 CESTOpus 5.5 at medium: score 51.2, $1.34 per task · Artificial Analysis, read 1 Oct 2026, 07:05 CESTOpus 5.5 at high: score 53.6, $1.82 per task · Artificial Analysis, read 1 Oct 2026, 07:05 CESTOpus 5.5 at xhigh: score 56.0, $3.46 per task · Artificial Analysis, read 1 Oct 2026, 07:05 CESTOpus 5.5 at max: score 57.6, $5.98 per task · Artificial Analysis, read 1 Oct 2026, 07:05 CESTOpus 5.5Gemini 4 Argon at high: score 52.6, $1.99 per task · Artificial Analysis, read 1 Oct 2026, 07:05 CESTGemini 4 Argonscorecost per task, log scale
Artificial Analysis Intelligence Index v4.3.2, average per task, API list prices including caching, Argon at its introductory price; Opus 5.5 with fallback. Read 1 October 2026.
The numbers in this chart
settingOpus 5.5 scoreOpus 5.5 cost per index taskGemini 4 Argon scoreGemini 4 Argon cost per index task
low42.3$0.55––
medium51.2$1.34––
high53.6$1.8252.6$1.99
xhigh56.0$3.46––
max57.6$5.98––

Artificial Analysis and Vals both run Opus 5.5 with a fallback: when it refuses a task, Anthropic's servers usually pass the task to an older Claude model, and the result counts as Opus 5.5's. That applies to every Opus 5.5 score from Artificial Analysis and Vals on this page. Artificial Analysis labels every Opus 5.5 row "Default Fallback" without publishing how many tasks that was on the index. The table in the coding section below gives the published share for each test on this page, and goes into what it changes.

Where Argon is ahead▶ 2:39

The index averages ten tests, and that hides where each model wins. These are the results where Argon is ahead that Google didn't produce itself.

On Harvey's Legal Agent Benchmark, which Vals runs, Argon scores 19.6% against 3.8% for Opus at max. On Vals' Finance Agent test, multi-step financial research, it scores 65.4% against 58.6%. On Artificial Analysis's own run of AutomationBench, business workflows across apps scored with partial credit, it scores 77.5% against 69.5% for Opus at max. That is the highest AutomationBench score of the 175 model-and-setting rows on Artificial Analysis's list. Google's table quotes Zapier's AutomationBench leaderboard instead, where Argon has 51.3% from Zapier's own run, not Artificial Analysis's.

On long software-engineering tasks, Google's headline claim holds up. Google reports 77.9% on DeepSWE from its own run. Artificial Analysis's coding-agent run, each model working inside its maker's own tool, gives Argon 78.8% against 68.4% for Claude Code with Opus 5.5 at max.

Where Argon is ahead

Results Google didn't produce itself: Argon at high, Opus 5.5 at max

Opus 5.5 at maxGemini 4 Argon at high
score, %Argon minus Opus, pointsHarvey's legal agents (Vals)Harvey's legal agents (Vals), Opus 5.5 at max: 3.8% · Vals AI, read 1 Oct 2026, 04:40 CEST3.8%Harvey's legal agents (Vals), Gemini 4 Argon at high: 19.6% · Vals AI, read 1 Oct 2026, 07:05 CEST19.6%+15.8Finance Agent (Vals)Finance Agent (Vals), Opus 5.5 at max: 58.6% · Vals AI, read 1 Oct 2026, 04:40 CEST58.6%Finance Agent (Vals), Gemini 4 Argon at high: 65.4% · Vals AI, read 1 Oct 2026, 07:05 CEST65.4%+6.8AutomationBench (AA)AutomationBench (AA), Opus 5.5 at max: 69.5% · Model Fatigue, arithmetic on Artificial Analysis's figures, read 1 Oct 2026, 07:05 CEST69.5%AutomationBench (AA), Gemini 4 Argon at high: 77.5% · Model Fatigue, arithmetic on Artificial Analysis's figures, read 1 Oct 2026, 07:05 CEST77.5%+8.0DeepSWE (AA)DeepSWE (AA), Opus 5.5 at max: 68.4% · Model Fatigue, arithmetic on Artificial Analysis's figures, read 1 Oct 2026, 09:09 CEST68.4%DeepSWE (AA), Gemini 4 Argon at high: 78.8% · Model Fatigue, arithmetic on Artificial Analysis's figures, read 1 Oct 2026, 09:09 CEST78.8%+10.4
score, %Argon minus Opus, pointsHarvey's legal agents (Vals)Harvey's legal agents (Vals), Opus 5.5 at max: 3.8% · Vals AI, read 1 Oct 2026, 04:40 CEST3.8%Harvey's legal agents (Vals), Gemini 4 Argon at high: 19.6% · Vals AI, read 1 Oct 2026, 07:05 CEST19.6%+15.8Finance Agent (Vals)Finance Agent (Vals), Opus 5.5 at max: 58.6% · Vals AI, read 1 Oct 2026, 04:40 CEST58.6%Finance Agent (Vals), Gemini 4 Argon at high: 65.4% · Vals AI, read 1 Oct 2026, 07:05 CEST65.4%+6.8AutomationBench (AA)AutomationBench (AA), Opus 5.5 at max: 69.5% · Model Fatigue, arithmetic on Artificial Analysis's figures, read 1 Oct 2026, 07:05 CEST69.5%AutomationBench (AA), Gemini 4 Argon at high: 77.5% · Model Fatigue, arithmetic on Artificial Analysis's figures, read 1 Oct 2026, 07:05 CEST77.5%+8.0DeepSWE (AA)DeepSWE (AA), Opus 5.5 at max: 68.4% · Model Fatigue, arithmetic on Artificial Analysis's figures, read 1 Oct 2026, 09:09 CEST68.4%DeepSWE (AA), Gemini 4 Argon at high: 78.8% · Model Fatigue, arithmetic on Artificial Analysis's figures, read 1 Oct 2026, 09:09 CEST78.8%+10.4
Vals AI and Artificial Analysis, read 1 October 2026. DeepSWE is from Artificial Analysis's Coding Agent Index: Claude Code with Opus 5.5 against Antigravity CLI with Argon. Differences are our arithmetic. Of Opus's tasks, an older model answered 1.3% on Finance Agent and 0.6% of attempts on DeepSWE; Artificial Analysis publishes no share for AutomationBench, and Vals none for Harvey's benchmark.
The numbers in this chart
Opus 5.5 at maxGemini 4 Argon at highArgon minus Opus, points
Harvey's legal agents (Vals)3.8%19.6%+15.8
Finance Agent (Vals)58.6%65.4%+6.8
AutomationBench (AA)69.5%77.5%+8.0
DeepSWE (AA)68.4%78.8%+10.4

People also prefer Argon's answers on the Arena's text leaderboard, where it is ranked #1 (1525, plus or minus 9), against 1504 for Opus 5.5 at high, ranked #4. The Arena still marks Argon's rating as preliminary.

Where Opus is ahead▶ 3:48

In a terminal, Opus leads. On Vals' Terminal-Bench 4.0, Opus at max scores 65.2% to Argon's 57.6%, and Google's own table shows a similar gap, 66.4% against 57.4%. On ProgramBench, which asks a model to rebuild a program from scratch, Opus fully resolves 18.5% of the tasks and Argon 2.5%.

On the Arena's web-development board, Opus 5.5 at max is ranked #1 (1818) and Argon #8 (1679, preliminary).

On Artificial Analysis's two tests of finished work, Opus at high is clearly ahead of Argon at high. AA-Briefcase has a model produce deliverables like spreadsheets, presentations and memos: Opus 1704, Argon 1494. On GDPval-AA, where an AI judge compares the work, Opus scores 1692 and Argon 1611.

Where Opus is ahead

Argon at high, Opus 5.5 at max

Opus 5.5 at maxGemini 4 Argon at high
score, %Argon minus Opus, pointsTerminal-Bench 4.0 (Vals)Terminal-Bench 4.0 (Vals), Opus 5.5 at max: 65.2% · Vals AI, read 1 Oct 2026, 04:40 CEST65.2%Terminal-Bench 4.0 (Vals), Gemini 4 Argon at high: 57.6% · Vals AI, read 1 Oct 2026, 07:05 CEST57.6%−7.6ProgramBench (Vals)ProgramBench (Vals), Opus 5.5 at max: 18.5% · Vals AI, read 1 Oct 2026, 04:40 CEST18.5%ProgramBench (Vals), Gemini 4 Argon at high: 2.5% · Vals AI, read 1 Oct 2026, 07:05 CEST2.5%−16.0Terminal-Bench v4 (AA)Terminal-Bench v4 (AA), Opus 5.5 at max: 63.1% · Model Fatigue, arithmetic on Artificial Analysis's figures, read 1 Oct 2026, 09:09 CEST63.1%Terminal-Bench v4 (AA), Gemini 4 Argon at high: 56.1% · Model Fatigue, arithmetic on Artificial Analysis's figures, read 1 Oct 2026, 09:09 CEST56.1%−7.0SWE-Atlas-QnA (AA)SWE-Atlas-QnA (AA), Opus 5.5 at max: 66.4% · Model Fatigue, arithmetic on Artificial Analysis's figures, read 1 Oct 2026, 09:09 CEST66.4%SWE-Atlas-QnA (AA), Gemini 4 Argon at high: 56.5% · Model Fatigue, arithmetic on Artificial Analysis's figures, read 1 Oct 2026, 09:09 CEST56.5%−9.9
score, %Argon minus Opus, pointsTerminal-Bench 4.0 (Vals)Terminal-Bench 4.0 (Vals), Opus 5.5 at max: 65.2% · Vals AI, read 1 Oct 2026, 04:40 CEST65.2%Terminal-Bench 4.0 (Vals), Gemini 4 Argon at high: 57.6% · Vals AI, read 1 Oct 2026, 07:05 CEST57.6%−7.6ProgramBench (Vals)ProgramBench (Vals), Opus 5.5 at max: 18.5% · Vals AI, read 1 Oct 2026, 04:40 CEST18.5%ProgramBench (Vals), Gemini 4 Argon at high: 2.5% · Vals AI, read 1 Oct 2026, 07:05 CEST2.5%−16.0Terminal-Bench v4 (AA)Terminal-Bench v4 (AA), Opus 5.5 at max: 63.1% · Model Fatigue, arithmetic on Artificial Analysis's figures, read 1 Oct 2026, 09:09 CEST63.1%Terminal-Bench v4 (AA), Gemini 4 Argon at high: 56.1% · Model Fatigue, arithmetic on Artificial Analysis's figures, read 1 Oct 2026, 09:09 CEST56.1%−7.0SWE-Atlas-QnA (AA)SWE-Atlas-QnA (AA), Opus 5.5 at max: 66.4% · Model Fatigue, arithmetic on Artificial Analysis's figures, read 1 Oct 2026, 09:09 CEST66.4%SWE-Atlas-QnA (AA), Gemini 4 Argon at high: 56.5% · Model Fatigue, arithmetic on Artificial Analysis's figures, read 1 Oct 2026, 09:09 CEST56.5%−9.9
Vals AI and Artificial Analysis, read 1 October 2026. Differences are our arithmetic. Every Opus score here includes tasks an older Claude model answered after a refusal: 11.1%, 12.0%, 12.1% and 14.0%.
The numbers in this chart
Opus 5.5 at maxGemini 4 Argon at highArgon minus Opus, points
Terminal-Bench 4.0 (Vals)65.2%57.6%−7.6
ProgramBench (Vals)18.5%2.5%−16.0
Terminal-Bench v4 (AA)63.1%56.1%−7.0
SWE-Atlas-QnA (AA)66.4%56.5%−9.9

Bloomberg reported on 1 October, citing people with direct access, that while Gemini 4 "has performed well on benchmarks widely used to gauge model efficacy, it does less well when employees actually put it to work", and that the model "struggles to handle certain coding tasks". Its sources asked not to be named. Google pushed back, as Finimize reported it, "saying it's inaccurate to characterize the model as underperforming in coding".

Coding, all together▶ 4:42

So who is better at coding? It's mixed.

On four other coding tests Vals runs, Argon is ahead of Opus at max: narrowly on Vibe Code Bench v1.1 (+1.6 points), Code Migration (+1.5) and IOI (+4.9), and clearly on SRE Bench (+10.7).

Artificial Analysis's Coding Agent Index is three tests. Argon wins DeepSWE, and Opus wins the other two, SWE-Atlas-QnA and Artificial Analysis's run of Terminal-Bench v4. Overall, Claude Code with Opus 5.5 at max scores 66.0 and Argon in Google's own Antigravity CLI scores 63.8, 2.2 points lower.

Artificial Analysis's Coding Agent Index

Each model in its maker's own agent

Claude Code, Opus 5.5 at maxAntigravity CLI, Gemini 4 Argon at high
index66.063.8
DeepSWE v1.168.4%78.8%
SWE-Atlas-QnA66.4%56.5%
Terminal-Bench v463.1%56.1%
cost per task$13.04$5.84
time per task64 min35 min
attempts that fell back to an older model78 of 9090 of 905
Artificial Analysis, read 1 October 2026 (the Argon row appeared that morning). Cost at API list prices, Argon at its introductory price.

Then there is the fallback. Counted on its own, a refused task would score nothing, so Opus's own scores on these tests would be lower than what's published. How much lower depends on the share of tasks an older model answered:

How much of Opus 5.5's score an older Claude model answered

Opus 5.5 at max, Argon at high. The share is of Opus's tasks (Vals) or attempts (Artificial Analysis) passed to Claude Opus 5 or Opus 4.8 after a refusal

testOpus 5.5Gemini 4 Argonanswered by an older model
Vals Index67.0%68.9%4.0%
Vals, Finance Agent58.6%65.4%1.3%
Vals, Terminal-Bench 4.065.2%57.6%11.1%
Vals, ProgramBench18.5%2.5%12.0%
Vals, Vibe Code Bench v1.190.3%91.9%8.0%
Vals, Code Migration66.7%68.2%81.5%
Vals, SRE Bench33.6%44.3%82.8%
Vals, IOI95.1%100.0%5.6%
Vals, CyberBench v1.155.4%77.9%52.6%
Vals, Terminal-Bench Science47.1%44.3%2.9%
AA agent index, DeepSWE v1.168.4%78.8%0.6%
AA agent index, SWE-Atlas-QnA66.4%56.5%14.0%
AA agent index, Terminal-Bench v463.1%56.1%12.1%
Vals AI's Opus 5.5 page (its per-test fallback labels) and Argon page; Artificial Analysis's Coding Agent Index (attempt counts per test, the shares our arithmetic). Vals reports no fallback on Harvey's legal benchmark; Argon's runs show none anywhere. Vals' page has a toggle for scores with fallbacks counted as failures; our saved copy shows the default view.

On the terminal tests that share is 11.1% of Vals' tasks and 12.1% of Artificial Analysis's attempts. On Code Migration and SRE Bench it is 81.5% and 82.8%, so those two say more about the older models than about Opus 5.5. Vals' own note shows the size of it on SRE Bench: counted as failures, Opus's score falls from 33.6% to 5.3%. On DeepSWE in the agent index it is 0.6%, and Argon's runs show none.

Opus's lead on Vals' Terminal-Bench 4.0 is about 15 tasks of 198, against 22 that an older model answered. When Vals first published Opus 5.5 on 22 September, its note said that counting the 30 fallback-assisted tasks it had then (of 198) as failures moved Opus's Terminal-Bench 4.0 score from 61.6% to 53.5%, below Argon's 57.6% today. Vals' page now shows Opus at 65.2%, and the matching fallback-free value sits behind a toggle on that page which our saved copy doesn't carry. In the agent index, Opus's 2.2 points lead compares with 8.6% of its attempts falling back. So Opus's lead in a terminal, and its lead on the coding-agent index, aren't settled.

How often it guesses▶ 5:57

Google's launch post doesn't mention it, but the independent numbers show how rarely Argon guesses when it doesn't know. AA-Omniscience is a knowledge test where a model can answer, or say it doesn't know; a right answer adds to the score, a wrong one takes away from it, and not answering does neither.

Opus 5.5 at high gets 64.6% of the questions right and 24.0% wrong. Argon gets 49.9% right and only 7.6% wrong; the remaining 42.5% it doesn't answer, or answers only in part. It knows less, and it bluffs much less, so the test's score comes out about the same: 42.4 for Argon, 40.6 for Opus. We split the shares out of Artificial Analysis's accuracy and score; the split reproduces the hallucination rate Artificial Analysis publishes.

Right, wrong, and not answered

AA-Omniscience, both models at high

Opus 5.5 at highGemini 4 Argon at high
share of questions, %Argon minus Opusanswered rightanswered right, Opus 5.5 at high: 64.6% · Model Fatigue, arithmetic on Artificial Analysis's figures, read 1 Oct 2026, 07:05 CEST64.6%answered right, Gemini 4 Argon at high: 49.9% · Model Fatigue, arithmetic on Artificial Analysis's figures, read 1 Oct 2026, 07:05 CEST49.9%answered wronganswered wrong, Opus 5.5 at high: 24.0% · Model Fatigue, arithmetic on Artificial Analysis's figures, read 1 Oct 2026, 07:05 CEST24.0%answered wrong, Gemini 4 Argon at high: 7.6% · Model Fatigue, arithmetic on Artificial Analysis's figures, read 1 Oct 2026, 07:05 CEST7.6%not answered, or in partnot answered, or in part, Opus 5.5 at high: 11.5% · Model Fatigue, arithmetic on Artificial Analysis's figures, read 1 Oct 2026, 07:05 CEST11.5%not answered, or in part, Gemini 4 Argon at high: 42.5% · Model Fatigue, arithmetic on Artificial Analysis's figures, read 1 Oct 2026, 07:05 CEST42.5%
share of questions, %Argon minus Opusanswered rightanswered right, Opus 5.5 at high: 64.6% · Model Fatigue, arithmetic on Artificial Analysis's figures, read 1 Oct 2026, 07:05 CEST64.6%answered right, Gemini 4 Argon at high: 49.9% · Model Fatigue, arithmetic on Artificial Analysis's figures, read 1 Oct 2026, 07:05 CEST49.9%answered wronganswered wrong, Opus 5.5 at high: 24.0% · Model Fatigue, arithmetic on Artificial Analysis's figures, read 1 Oct 2026, 07:05 CEST24.0%answered wrong, Gemini 4 Argon at high: 7.6% · Model Fatigue, arithmetic on Artificial Analysis's figures, read 1 Oct 2026, 07:05 CEST7.6%not answered, or in partnot answered, or in part, Opus 5.5 at high: 11.5% · Model Fatigue, arithmetic on Artificial Analysis's figures, read 1 Oct 2026, 07:05 CEST11.5%not answered, or in part, Gemini 4 Argon at high: 42.5% · Model Fatigue, arithmetic on Artificial Analysis's figures, read 1 Oct 2026, 07:05 CEST42.5%
Artificial Analysis, read 1 October 2026. The split into right, wrong and the rest is our arithmetic from Artificial Analysis's accuracy and score; it reproduces Artificial Analysis's published hallucination rate.
The numbers in this chart
Opus 5.5 at highGemini 4 Argon at highArgon minus Opus
answered right64.6%49.9%
answered wrong24.0%7.6%
not answered, or in part11.5%42.5%

Artificial Analysis's hallucination rate is the wrong answers as a share of everything a model didn't get right. Argon's is 15.1%. Among the 34 model-and-setting rows on Artificial Analysis's list that score above forty-five on the index, none is lower; the next is Qwen3.8 Max (0902) at 28.8%. In legal and finance work, where a confident wrong answer does the most damage, that could matter more than the test's score.

How often each model guesses when it doesn't know

Hallucination rate: wrong answers as a share of everything not answered right

0%10%20%30%40%50%60%70%80%Gemini 4 Argon, highGemini 4 Argon, high: 15.1% · Artificial Analysis, read 1 Oct 2026, 07:05 CEST15.1%Qwen3.8 Max (0902)Qwen3.8 Max (0902): 28.8% · Artificial Analysis, read 1 Oct 2026, 07:05 CEST28.8%Grok 4.7, xhighGrok 4.7, xhigh: 29.3% · Artificial Analysis, read 1 Oct 2026, 07:05 CEST29.3%GPT-6 Astra, highGPT-6 Astra, high: 44.8% · Artificial Analysis, read 1 Oct 2026, 07:05 CEST44.8%GPT-6.1 Sol, maxGPT-6.1 Sol, max: 54.3% · Artificial Analysis, read 1 Oct 2026, 07:05 CEST54.3%Opus 5.5, maxOpus 5.5, max: 58.6% · Artificial Analysis, read 1 Oct 2026, 07:05 CEST58.6%Opus 5.5, highOpus 5.5, high: 67.6% · Artificial Analysis, read 1 Oct 2026, 07:05 CEST67.6%Fable 5.1, maxFable 5.1, max: 72.6% · Artificial Analysis, read 1 Oct 2026, 07:05 CEST72.6%
0%10%20%30%40%50%60%70%80%Gemini 4 Argon, highGemini 4 Argon, high: 15.1% · Artificial Analysis, read 1 Oct 2026, 07:05 CEST15.1%Qwen3.8 Max (0902)Qwen3.8 Max (0902): 28.8% · Artificial Analysis, read 1 Oct 2026, 07:05 CEST28.8%Grok 4.7, xhighGrok 4.7, xhigh: 29.3% · Artificial Analysis, read 1 Oct 2026, 07:05 CEST29.3%GPT-6 Astra, highGPT-6 Astra, high: 44.8% · Artificial Analysis, read 1 Oct 2026, 07:05 CEST44.8%GPT-6.1 Sol, maxGPT-6.1 Sol, max: 54.3% · Artificial Analysis, read 1 Oct 2026, 07:05 CEST54.3%Opus 5.5, maxOpus 5.5, max: 58.6% · Artificial Analysis, read 1 Oct 2026, 07:05 CEST58.6%Opus 5.5, highOpus 5.5, high: 67.6% · Artificial Analysis, read 1 Oct 2026, 07:05 CEST67.6%Fable 5.1, maxFable 5.1, max: 72.6% · Artificial Analysis, read 1 Oct 2026, 07:05 CEST72.6%
Artificial Analysis's AA-Omniscience, read 1 October 2026. Argon has the lowest rate of the 34 model-and-setting rows scoring above forty-five on the index; the two after it are the next lowest of those. The Opus 5.5 and Fable 5.1 rows are run with fallback. Of every row Artificial Analysis lists for Opus 5.5, Fable 5.1, Sonnet 5.5, GPT-6 Astra and GPT-6.1 Sol, GPT-6 Astra at high has the lowest rate and Fable 5.1 at max the highest.
The numbers in this chart
hallucination rate
Gemini 4 Argon, high15.1%
Qwen3.8 Max (0902)28.8%
Grok 4.7, xhigh29.3%
GPT-6 Astra, high44.8%
GPT-6.1 Sol, max54.3%
Opus 5.5, max58.6%
Opus 5.5, high67.6%
Fable 5.1, max72.6%

Half the price?▶ 6:49

Per token, Argon is half of Opus, for now. Per task, it depends on which Opus you compare it with.

Against Opus at max, Argon at high costs a third as much on Artificial Analysis's index ($1.99 against $5.98) and scores 5.0 points less. In the coding-agent index it's $5.84 a task against $13.04, under half (0.45 of the cost).

Against Opus at the same quality, the gap goes away: Opus at high scores a point more than Argon, for $0.17 less per task. Argon writes more to get there, 61.6K output tokens per index task against 35.6K for Opus at high, and spends more on the input side too, $1.37 against $1.11.

All of those Argon prices are introductory. When that period ends, every token price doubles. Google states the cached-input discount only for the introductory price, so we give a range: with the cached rate doubled too, or left where it is. We price cache writes at the input rate in both, as Artificial Analysis does at the introductory price. By our arithmetic, Argon then costs $3.66 to $3.98 per index task, about twice Opus at high (2.01 to 2.18 times its cost), and the coding agent's $5.84 becomes $10.51 to $11.70, close to Opus at max's $13.04. We checked the method on the agent run first: Argon's mean tokens at the introductory price come to $5.85, which matches Artificial Analysis's figure.

Per task, at the introductory price and after it

Argon at high against the Opus setting each comparison uses

comparisonGemini 4 Argon nowGemini 4 Argon after the introductory priceOpus 5.5scores
AA index, Opus at max$1.99$3.66 to $3.98$5.9852.6 against 57.6
AA index, Opus at high$1.99$3.66 to $3.98$1.8252.6 against 53.6
AA coding agent, Opus at max$5.84$10.51 to $11.70$13.0463.8 against 66.0
Vals Index, Opus at maxalready at the full price$15.68$32.1468.9% against 67.0%
Artificial Analysis and Vals AI, read 1 October 2026. The after-price columns are our arithmetic: every token price doubled, with cached input either doubled or left at the introductory rate, and cache writes at the input price. Vals prices Argon at the full price already.

Vals priced Argon at the full price, $4 in and $20 out, from the start. Its Vals Index costs $15.68 per test for Argon against $32.14 for Opus at max, about half, on a different mix of tasks from Artificial Analysis's. Google describes the Vals Index as spanning finance, coding, legal and tax work, each weighted by its share of US GDP, and Argon scores higher on it (68.9% against 67.0%).

Google's table, checked▶ 7:49

Google's methodology page says where each row of its table comes from. The Vals Index, Finance Agent and Harvey's legal benchmark rows are sourced from Vals AI, Vibe Code Bench from Vals' public leaderboard, AutomationBench from Zapier's public leaderboard, and several more from public leaderboards. Argon's DeepSWE and Terminal-Bench 4.0 scores are Google's own runs.

Google's benchmark table, with where each row comes from

Argon against GPT-6 Astra, Fable 5.1 and Opus 5.5, as Google published it. Marked rows: Opus scores higher

benchmarkGemini 4 ArgonGPT-6 AstraFable 5.1Opus 5.5where Google's numbers come from
Vals Index68.9%63.1%65.8%67.0%Vals AI
AutomationBench51.3%41.4%31.4%42.5%Zapier's public leaderboard
Vals Finance Agent65.4%53.5%58.9%58.6%Vals AI
Harvey's Legal Agent Benchmark19.6%5.4%6.7%3.8%Vals AI
DeepSWE v1.177.9%74.1%67.4%74.2%Argon: Google's own run; Astra: public leaderboard; Fable and Opus: their system cards
FrontierSWE55.0%65.5%56.3%62.3%Proximal's public leaderboard
Vibe Code Bench91.9%89.6%90.3%90.3%Vals AI's public leaderboard
Terminal-Bench 4.057.4%58.2%57.9%66.4%Argon: Google's own run; others: public leaderboard
PostTrainBench45.3%44.3%40.2%49.3%Google's runs for every model
Terminal-Bench Science57.6%68.1%52.6%63.3%Argon: Google's own run, with a longer verifier timeout; others: public leaderboard
LABBench 288.8%85.4%68.6%73.1%Google's runs for every model
RiemannBench76.0%72.0%65.6%69.6%Surge's public leaderboard
GraphWalks, shorter contexts99.7%98.7%91.4%90.6%Google's runs for every model
GraphWalks, longest contexts84.2%71.8%65.0%66.8%Google's runs for every model
Agent's Last Exam39.5%34.2%no score38.2%Argon: Google's own run; others: public leaderboard
OSWorld-2.069.2%72.6%no scoreno scoreArgon: Google's own run; Astra: OpenAI's post
Chartography71.6%71.0%46.2%66.3%Surge's public leaderboard
LVBench91.7%87.5%79.7%83.7%Google's runs for every model
CWE-bench68.0%68.0%58.0%67.0%public leaderboard
Google DeepMind's model page and its evaluation methodology page, read 1 October 2026. For Opus 5.5 Google reports the highest setting it found a result for. Counting GraphWalks once, 17 benchmarks have an Opus score, and Argon's is higher on 13.

Argon's Vals numbers in the table match Vals' own pages: the Vals Index 68.9%, Finance Agent 65.4%, Harvey's benchmark 19.6% and Vibe Code Bench v1.1 91.9%, rounded. DeepSWE, the headline number, is Google's own run, and Artificial Analysis's agent run gives a similar score, 78.8% against Google's 77.9%.

On Terminal-Bench Science, Google's figure for Argon, 57.6%, is "self computed, with 6x verifier timeout to address timeout issues with verification", according to the methodology page. Vals runs the same test and has Argon at 44.3% and Opus 5.5 at max at 47.1%, with its own caution that its numbers are "not directly comparable to the official" leaderboard, which pairs each model with its own agent. So the gap between Google's figure and Vals' says little about Argon on its own.

GPT-6.1 Sol isn't in the table, but it came out the day before Argon.

Against the rest▶ 8:12

GPT-6 Astra at max scores 52.7, level with Argon, at $3.26 a task, so for now Argon costs about sixty percent of what Astra does (0.61 times). GPT-6.1 Sol launched at the same price per token as Argon's introductory price ($2 in, $0.10 cached, $10 out). At max it scores 51.8, within a point of Argon, for $0.72 a task, a little over a third of Argon's cost (0.36 times). In Codex at xhigh it scores 62.9 on the coding-agent index for $1.04 a task, against Argon's 63.8 for $5.84. Anthropic's Fable 5.1 at max, which Artificial Analysis also runs with fallback, scores 53.4, for $7.63.

Against the rest of the top tier

Intelligence Index score against cost per task, four models at the settings Artificial Analysis lists; Fable 5.1 is in the text above. The yellow lines join the two pairs the text compares.

GPT-6.1 SolGPT-6 AstraOpus 5.5Gemini 4 Argon
405060$0.10$0.30$1$3$10GPT-6.1 Sol at low: score 42.1, $0.13 per task · Artificial Analysis, read 1 Oct 2026, 07:05 CESTlowGPT-6.1 Sol at medium: score 47.8, $0.21 per task · Artificial Analysis, read 1 Oct 2026, 07:05 CESTmediumGPT-6.1 Sol at high: score 50.2, $0.32 per task · Artificial Analysis, read 1 Oct 2026, 07:05 CESThighGPT-6.1 Sol at xhigh: score 51.0, $0.39 per task · Artificial Analysis, read 1 Oct 2026, 07:05 CESTxhighGPT-6.1 Sol at max: score 51.8, $0.72 per task · Artificial Analysis, read 1 Oct 2026, 07:05 CESTmaxGPT-6.1 SolGPT-6 Astra at low: score 45.8, $0.82 per task · Artificial Analysis, read 1 Oct 2026, 07:05 CESTGPT-6 Astra at medium: score 49.6, $1.54 per task · Artificial Analysis, read 1 Oct 2026, 07:05 CESTGPT-6 Astra at high: score 50.9, $1.73 per task · Artificial Analysis, read 1 Oct 2026, 07:05 CESTGPT-6 Astra at xhigh: score 52.4, $2.31 per task · Artificial Analysis, read 1 Oct 2026, 07:05 CESTGPT-6 Astra at max: score 52.7, $3.26 per task · Artificial Analysis, read 1 Oct 2026, 07:05 CESTGPT-6 AstraOpus 5.5 at low: score 42.3, $0.55 per task · Artificial Analysis, read 1 Oct 2026, 07:05 CESTOpus 5.5 at medium: score 51.2, $1.34 per task · Artificial Analysis, read 1 Oct 2026, 07:05 CESTOpus 5.5 at high: score 53.6, $1.82 per task · Artificial Analysis, read 1 Oct 2026, 07:05 CESTOpus 5.5 at xhigh: score 56.0, $3.46 per task · Artificial Analysis, read 1 Oct 2026, 07:05 CESTOpus 5.5 at max: score 57.6, $5.98 per task · Artificial Analysis, read 1 Oct 2026, 07:05 CESTOpus 5.5Gemini 4 Argon at high: score 52.6, $1.99 per task · Artificial Analysis, read 1 Oct 2026, 07:05 CESTGemini 4 Argonscorecost per task, log scale
405060$0.10$0.30$1$3$10GPT-6.1 Sol at low: score 42.1, $0.13 per task · Artificial Analysis, read 1 Oct 2026, 07:05 CESTGPT-6.1 Sol at medium: score 47.8, $0.21 per task · Artificial Analysis, read 1 Oct 2026, 07:05 CESTGPT-6.1 Sol at high: score 50.2, $0.32 per task · Artificial Analysis, read 1 Oct 2026, 07:05 CESTGPT-6.1 Sol at xhigh: score 51.0, $0.39 per task · Artificial Analysis, read 1 Oct 2026, 07:05 CESTGPT-6.1 Sol at max: score 51.8, $0.72 per task · Artificial Analysis, read 1 Oct 2026, 07:05 CESTGPT-6.1 SolGPT-6 Astra at low: score 45.8, $0.82 per task · Artificial Analysis, read 1 Oct 2026, 07:05 CESTGPT-6 Astra at medium: score 49.6, $1.54 per task · Artificial Analysis, read 1 Oct 2026, 07:05 CESTGPT-6 Astra at high: score 50.9, $1.73 per task · Artificial Analysis, read 1 Oct 2026, 07:05 CESTGPT-6 Astra at xhigh: score 52.4, $2.31 per task · Artificial Analysis, read 1 Oct 2026, 07:05 CESTGPT-6 Astra at max: score 52.7, $3.26 per task · Artificial Analysis, read 1 Oct 2026, 07:05 CESTGPT-6 AstraOpus 5.5 at low: score 42.3, $0.55 per task · Artificial Analysis, read 1 Oct 2026, 07:05 CESTOpus 5.5 at medium: score 51.2, $1.34 per task · Artificial Analysis, read 1 Oct 2026, 07:05 CESTOpus 5.5 at high: score 53.6, $1.82 per task · Artificial Analysis, read 1 Oct 2026, 07:05 CESTOpus 5.5 at xhigh: score 56.0, $3.46 per task · Artificial Analysis, read 1 Oct 2026, 07:05 CESTOpus 5.5 at max: score 57.6, $5.98 per task · Artificial Analysis, read 1 Oct 2026, 07:05 CESTOpus 5.5Gemini 4 Argon at high: score 52.6, $1.99 per task · Artificial Analysis, read 1 Oct 2026, 07:05 CESTGemini 4 Argonscorecost per task, log scale
Artificial Analysis Intelligence Index v4.3.2, API list prices including caching, Argon at its introductory price; Opus 5.5 with fallback. Read 1 October 2026.
The numbers in this chart
settingGPT-6.1 Sol scoreGPT-6.1 Sol cost per index taskGPT-6 Astra scoreGPT-6 Astra cost per index taskOpus 5.5 scoreOpus 5.5 cost per index taskGemini 4 Argon scoreGemini 4 Argon cost per index task
low42.1$0.1345.8$0.8242.3$0.55––
medium47.8$0.2149.6$1.5451.2$1.34––
high50.2$0.3250.9$1.7353.6$1.8252.6$1.99
xhigh51.0$0.3952.4$2.3156.0$3.46––
max51.8$0.7252.7$3.2657.6$5.98––

Why it's gated▶ 8:59

The launch pages we read don't link a system card for Argon. Google's methodology page links Anthropic's Opus 5.5 system card, not one for Argon. So this part is Google's own account.

Google says that "safely releasing frontier capabilities at this level requires a phased approach", starting with cyber defenders, and that it is "actively engaged in the U.S. government's voluntary process for pre-release model access". For trusted defenders and Google's own teams, it is releasing Argon "without cyber guardrails". Before a wider release, Google describes four kinds of safeguard: defending against misuse, defending against prompt injection attacks, monitoring for misalignment (watching Argon's chain of thought and actions and stopping execution when necessary), and hardening the environments it is tested in.

Two numbers here don't come from Google's own runs. On Vals' CyberBench v1.1, Argon scores 77.9% against 55.4% for Opus at max, and Opus refused 52.6% of those tasks (61), which older Claude models then answered. On CWE-bench, which tests fixing security vulnerabilities and whose public leaderboard Google quotes, it is 68.0% against 67.0%.

Should you switch?▶ 9:56

You can't yet. For when Argon opens to paid API customers and Google AI Ultra subscribers, this is what the evidence says so far.

For legal, finance and business-automation agents, it's worth trying first: independent tests put it well ahead of Opus there, and it guesses wrong far less often. For coding it's mixed, so test both on your own work before you move. Budget for the full price, not the introductory one. If you want something close today, GPT-6.1 Sol scores within a point of Argon on Artificial Analysis's index for a little over a third of the cost per task.

Where the evidence points, so far

For when Argon opens to paid API customers and Google AI Ultra subscribers

workwhat the independent tests show
legal and finance agentsArgon well ahead on Vals' tests (19.6% against 3.8%; 65.4% against 58.6%)
business workflows across appsArgon ahead on Artificial Analysis's AutomationBench (77.5% against 69.5%)
codingmixed: Argon ahead on DeepSWE, Opus in a terminal and on ProgramBench, and part of Opus's scores are older models' answers
finished work: spreadsheets, presentations, memosOpus at high ahead on both of Artificial Analysis's tests
web front endsOpus ranked #1 on the Arena's board, Argon #8
not measured yetreal use outside Google, Argon at settings other than high on Artificial Analysis, Vals and the Arena, how long the introductory price lasts
Every figure from the sections above, read 1 October 2026.

Hold all of this loosely. It's under a day of evidence, and Artificial Analysis, Vals and the Arena each ran Argon at high only. When model-economics' sweep stopped looking, early on 1 October, it hadn't found a verified write-up from anyone outside Google using Argon for real work. Real use, the other settings, and how long the introductory price lasts would settle it.

Just before this video, we looked at GPT-6.1 Sol, OpenAI's release from the day before. That's our previous video.

Every number

These are all 407 figures behind the video and this page, grouped by whose they are, with the page each came from and when we read it. Figures marked ⟳ can move. When a re-read finds a change, the new value shows next to the one from the video.

Artificial Analysis Intelligence Index, by model and effort setting

Artificial Analysis · Artificial Analysis, Gemini 4 Argon model page (Intelligence Index v4.3.2, with every model it lists) · read 1 Oct 2026, 07:05 CEST · ⟳ re-read 1 Oct 2026, 10:09 CEST, unchanged · Intelligence Index v4.3.2, average per task; cost at API list prices including caching, Argon at its introductory price. Every Opus 5.5 and Fable 5.1 row is labelled 'Default Fallback' by Artificial Analysis. · Argon has one row only, at high: Artificial Analysis, Vals and the Arena all ran it at high.

Intelligence index score

lowmediumhighxhighmax
Gemini 4 Argon––52.6––
Opus 5.542.351.253.656.057.6
GPT-6.1 Sol42.147.850.251.051.8
GPT-6 Astra45.849.650.952.452.7
Fable 5.146.848.951.253.253.4

Cost per index task, list prices incl. caching

lowmediumhighxhighmax
Gemini 4 Argon––$1.99––
Opus 5.5$0.55$1.34$1.82$3.46$5.98
GPT-6.1 Sol$0.13$0.21$0.32$0.39$0.72
GPT-6 Astra$0.82$1.54$1.73$2.31$3.26
Fable 5.1$2.37$2.98$3.91$5.98$7.63

Output tokens per index task

lowmediumhighxhighmax
Gemini 4 Argon––61.6K––
Opus 5.510.2K25.7K35.6K65.7K119.2K
GPT-6.1 Sol4.0K8.1K13.2K17.6K38.1K
GPT-6 Astra4.4K9.6K11.8K16.9K27.2K
Fable 5.121.6K27.9K38.1K60.5K78.1K

Cost per index task on the input side (uncached, cache reads and writes)

lowmediumhighxhighmax
Gemini 4 Argon––$1.37––
Opus 5.5$0.35$0.82$1.11$2.15$3.60

Cost per index task on cache reads

lowmediumhighxhighmax
Gemini 4 Argon––$0.32––

Automationbench-aa score (business workflows across apps, partial credit)

lowmediumhighxhighmax
Gemini 4 Argon––77.5%––
Opus 5.552.9%61.2%63.2%65.0%69.5%

Gdp.pdf score (questions over long documents)

lowmediumhighxhighmax
Gemini 4 Argon––21.8%––
Opus 5.525.6%25.6%28.8%26.6%26.2%

Terminal-bench 4.0 score, aa's run

lowmediumhighxhighmax
Gemini 4 Argon––57.1%––
Opus 5.531.3%52.5%56.6%59.6%59.6%

Gdpval-aa elo (finished work, judged by an ai judge)

lowmediumhighxhighmax
Gemini 4 Argon––1611––
Opus 5.512241576169218201846

Aa-briefcase elo (deliverables like spreadsheets, presentations and memos)

lowmediumhighxhighmax
Gemini 4 Argon––1494––
Opus 5.512851642170417801822

Humanity's last exam score

lowmediumhighxhighmax
Gemini 4 Argon––57.1%––
Opus 5.548.3%54.7%55.6%57.5%61.4%

Aa-omniscience, share of questions answered right

lowmediumhighxhighmax
Gemini 4 Argon––49.9%––
Opus 5.563.5%64.5%64.6%65.4%66.2%

Aa-omniscience, share answered wrong (accuracy minus score; reproduces aa's hallucination rate)

lowmediumhighxhighmax
Gemini 4 Argon––7.6%––
Opus 5.524.7%24.2%24.0%22.7%19.8%

Aa-omniscience, share not answered or answered in part (the rest)

lowmediumhighxhighmax
Gemini 4 Argon––42.5%––
Opus 5.511.8%11.2%11.5%11.9%14.0%

Aa-omniscience score (right answers minus wrong ones)

lowmediumhighxhighmax
Gemini 4 Argon––42.4––
Opus 5.538.940.340.642.646.4

Aa-omniscience hallucination rate (wrong answers as a share of everything not answered right)

lowmediumhighxhighmax
Gemini 4 Argon––15.1%––
Opus 5.567.6%68.4%67.6%65.7%58.6%
GPT-6.1 Sol51.6%51.6%49.4%50.9%54.3%
GPT-6 Astra46.9%46.5%44.8%48.3%51.3%
Fable 5.165.6%69.1%68.8%70.5%72.6%

Artificial Analysis

Intelligence Index score, Gemini 3.8 Flash at high40.9 ⟳
release date of Gemini 3.6 Flash, as Artificial Analysis records it2026-07-21
release date of Gemini 3.7 Flash, as Artificial Analysis records it2026-08-13
release date of Gemini 3.8 Flash, as Artificial Analysis records it2026-09-02
model-and-setting rows on Artificial Analysis's list scoring above 45 on the index34
model rows on Artificial Analysis's list in the saved page175
of the rows above 45 on the index, the one with the lowest hallucination rateGemini 4 Argon (High)
its hallucination rate15.1%
of the rows above 45 on the index, the one with the second-lowest hallucination rateQwen3.8 Max (0902)
its hallucination rate28.8%
of the rows above 45 on the index, the one with the third-lowest hallucination rateGrok 4.7 (Xhigh)
its hallucination rate29.3%
the top AutomationBench-AA score of every row on the list: whoseGemini 4 Argon (High)
model-and-setting rows with an AutomationBench-AA score175
Intelligence Index, Argon minus Gemini 3.8 Flash (our arithmetic)+11.7 points
Intelligence Index, Argon at high minus Opus 5.5 at high (our arithmetic)−1.0 points
Intelligence Index, Opus 5.5 at max above Argon at high (our arithmetic)5.0 points
cost per task, how much less Opus 5.5 at high costs than Argon at high (our arithmetic)$0.17
cost per task, Argon at high against Opus 5.5 at max (our arithmetic)0.33
cost per task, Argon at high against GPT-6 Astra at max (our arithmetic)0.61
Intelligence Index, Argon at high minus GPT-6 Astra at max (our arithmetic)−0.1 points
cost per task, GPT-6.1 Sol at max against Argon at high (our arithmetic)0.36
Intelligence Index, Argon at high minus GPT-6.1 Sol at max (our arithmetic)+0.8 points
output tokens per index task, Argon at high against Opus 5.5 at high (our arithmetic)1.7×
AutomationBench-AA, Argon at high, as a percentage (our arithmetic)77.5%
AutomationBench-AA, Opus 5.5 at max, as a percentage (our arithmetic)69.5%
AutomationBench-AA, Opus 5.5 at high, as a percentage (our arithmetic)63.2%
GDP.pdf, Argon at high, as a percentage (our arithmetic)21.8%
GDP.pdf, Opus 5.5 at max, as a percentage (our arithmetic)26.2%
GDP.pdf, Opus 5.5 at high, as a percentage (our arithmetic)28.8%
Terminal-Bench 4.0, Argon at high, as a percentage (our arithmetic)57.1%
Terminal-Bench 4.0, Opus 5.5 at max, as a percentage (our arithmetic)59.6%
Terminal-Bench 4.0, Opus 5.5 at high, as a percentage (our arithmetic)56.6%
AutomationBench-AA, Argon at high minus Opus 5.5 at max (our arithmetic)+8.0
cost per index task for Argon at the price after the introductory period, every part of the price doubled, cached input included (our arithmetic)$3.98
the same with cached input left at the introductory rate (Google doesn't state the later cached rate) (our arithmetic)$3.66
Argon at the full price (upper end) against Opus 5.5 at high, cost per index task (our arithmetic)2.18
Argon at the full price (lower end) against Opus 5.5 at high, cost per index task (our arithmetic)2.01
AA-Omniscience, Opus 5.5 at high's right-answer share above Argon's (our arithmetic)14.7 points
AA-Omniscience, Argon at high, acc, as a percentage (our arithmetic)49.9%
AA-Omniscience, Opus 5.5 at high, acc, as a percentage (our arithmetic)64.6%
AA-Omniscience, Argon at high, wrong, as a percentage (our arithmetic)7.6%
AA-Omniscience, Opus 5.5 at high, wrong, as a percentage (our arithmetic)24.0%
AA-Omniscience, Argon at high, rest, as a percentage (our arithmetic)42.5%
AA-Omniscience, Opus 5.5 at high, rest, as a percentage (our arithmetic)11.5%

Artificial Analysis

Coding Agent Index, Antigravity CLI with Gemini 4 Argon at high63.8 ⟳
Coding Agent Index, Antigravity CLI with Gemini 4 Argon at high, cost per task$5.84 ⟳
Coding Agent Index, Antigravity CLI with Gemini 4 Argon at high, agent time per task35 min ⟳
Coding Agent Index, Antigravity CLI with Gemini 4 Argon at high: attempts that fell back to an older model0
Coding Agent Index, Antigravity CLI with Gemini 4 Argon at high: attempts905
DeepSWE v1.1 in the Coding Agent Index, Antigravity CLI with Gemini 4 Argon at high78.8% ⟳
DeepSWE v1.1, Antigravity CLI with Gemini 4 Argon at high: attempts that fell back0
DeepSWE v1.1, Antigravity CLI with Gemini 4 Argon at high: attempts336
SWE-Atlas-QnA in the Coding Agent Index, Antigravity CLI with Gemini 4 Argon at high56.5% ⟳
SWE-Atlas-QnA, Antigravity CLI with Gemini 4 Argon at high: attempts that fell back0
SWE-Atlas-QnA, Antigravity CLI with Gemini 4 Argon at high: attempts371
Terminal-Bench v4 in the Coding Agent Index, Antigravity CLI with Gemini 4 Argon at high56.1% ⟳
Terminal-Bench v4, Antigravity CLI with Gemini 4 Argon at high: attempts that fell back0
Terminal-Bench v4, Antigravity CLI with Gemini 4 Argon at high: attempts198
Coding Agent Index, Claude Code with Opus 5.5 at max66.0 ⟳
Coding Agent Index, Claude Code with Opus 5.5 at max, cost per task$13.04 ⟳
Coding Agent Index, Claude Code with Opus 5.5 at max, agent time per task64 min ⟳
Coding Agent Index, Claude Code with Opus 5.5 at max: attempts that fell back to an older model78
Coding Agent Index, Claude Code with Opus 5.5 at max: attempts909
DeepSWE v1.1 in the Coding Agent Index, Claude Code with Opus 5.5 at max68.4% ⟳
DeepSWE v1.1, Claude Code with Opus 5.5 at max: attempts that fell back2
DeepSWE v1.1, Claude Code with Opus 5.5 at max: attempts339
SWE-Atlas-QnA in the Coding Agent Index, Claude Code with Opus 5.5 at max66.4% ⟳
SWE-Atlas-QnA, Claude Code with Opus 5.5 at max: attempts that fell back52
SWE-Atlas-QnA, Claude Code with Opus 5.5 at max: attempts372
Terminal-Bench v4 in the Coding Agent Index, Claude Code with Opus 5.5 at max63.1% ⟳
Terminal-Bench v4, Claude Code with Opus 5.5 at max: attempts that fell back24
Terminal-Bench v4, Claude Code with Opus 5.5 at max: attempts198
Coding Agent Index, Codex with GPT-6.1 Sol at xhigh62.9 ⟳
Coding Agent Index, Codex with GPT-6.1 Sol at xhigh, cost per task$1.04 ⟳
Coding Agent Index, Codex with GPT-6.1 Sol at xhigh, agent time per task16 min ⟳
Coding Agent Index, Argon: mean input tokens per task13.62M ⟳
Coding Agent Index, Argon: mean cached input tokens per task11.85M ⟳
Coding Agent Index, Argon: mean output tokens per task112K ⟳
DeepSWE v1.1, Antigravity CLI with Gemini 4 Argon at high, as a percentage (our arithmetic)78.8%
SWE-Atlas-QnA, Antigravity CLI with Gemini 4 Argon at high, as a percentage (our arithmetic)56.5%
Terminal-Bench v4, Antigravity CLI with Gemini 4 Argon at high, as a percentage (our arithmetic)56.1%
DeepSWE v1.1, Claude Code with Opus 5.5 at max, as a percentage (our arithmetic)68.4%
SWE-Atlas-QnA, Claude Code with Opus 5.5 at max, as a percentage (our arithmetic)66.4%
Terminal-Bench v4, Claude Code with Opus 5.5 at max, as a percentage (our arithmetic)63.1%
DeepSWE v1.1, Claude Code with Opus 5.5 at max: share of attempts that fell back (our arithmetic)0.6%
SWE-Atlas-QnA, Claude Code with Opus 5.5 at max: share of attempts that fell back (our arithmetic)14.0%
Terminal-Bench v4, Claude Code with Opus 5.5 at max: share of attempts that fell back (our arithmetic)12.1%
Coding Agent Index, Claude Code with Opus 5.5 at max: share of all attempts that fell back (our arithmetic)8.6%
Coding Agent Index, Claude Code with Opus 5.5 at max above Argon in Antigravity CLI (our arithmetic)2.2 points
Coding Agent Index cost per task, Argon against Claude Code with Opus 5.5 at max (our arithmetic)0.45
DeepSWE v1.1 in the agent index, Argon minus Opus 5.5 (our arithmetic)+10.4
SWE-Atlas-QnA in the agent index, Argon minus Opus 5.5 (our arithmetic)−9.9
Terminal-Bench v4 in the agent index, Argon minus Opus 5.5 (our arithmetic)−7.0
Argon's agent-index cost per task at the price after the introductory period, from its mean tokens, cached input at 95% off the new input price (our arithmetic)$11.70
the same with cached input left at the introductory rate (our arithmetic)$10.51
Argon's agent-index cost per task recomputed from its mean tokens at the introductory price (checks that Artificial Analysis priced it at the introductory rate) (our arithmetic)$5.85
Coding Agent Index cost per task, Codex with GPT-6.1 Sol at xhigh against Argon (our arithmetic)0.18

Google

Argon, introductory API price per million input tokens$2
Argon, introductory API price per million output tokens$10
Argon, cached input discount off the input price (stated with the introductory price)95%
Argon, API price per million input tokens after the introductory period$4
Argon, API price per million output tokens after the introductory period$20
Argon's output token limit, Google says (tokens)1M
the previous Gemini output limit, Google says (tokens)64K
DeepSWE v1.1, Argon, Google's own run (its headline)77.9%
Argon, introductory price per million cached input tokens (our arithmetic)$0.10
Argon's introductory price per token against Opus 5.5's (input; output is the same ratio) (our arithmetic)0.50

Anthropic

Anthropic, Claude Opus 5.5 (launch page) · read 30 Sep 2026, 02:28 CEST
Opus 5.5, API price per million input tokens$4
Opus 5.5, API price per million output tokens$20
Opus 5.5, API price per million cache-read tokens$0.20

OpenAI

OpenAI, Introducing GPT-6.1 Sol · read 29 Sep 2026, 22:32 CEST
GPT-6.1 Sol, API price per million input tokens$2
GPT-6.1 Sol, API price per million cached input tokens$0.10
GPT-6.1 Sol, API price per million output tokens$10

Google

Google's table, Vals Index, Gemini 4 Argon68.9%
Google's table, Vals Index, GPT-6 Astra63.1%
Google's table, Vals Index, Fable 5.165.8%
Google's table, Vals Index, Opus 5.567.0%
Google's table, AutomationBench, Gemini 4 Argon51.3%
Google's table, AutomationBench, GPT-6 Astra41.4%
Google's table, AutomationBench, Fable 5.131.4%
Google's table, AutomationBench, Opus 5.542.5%
Google's table, Vals Finance Agent v2, Gemini 4 Argon65.4%
Google's table, Vals Finance Agent v2, GPT-6 Astra53.5%
Google's table, Vals Finance Agent v2, Fable 5.158.9%
Google's table, Vals Finance Agent v2, Opus 5.558.6%
Google's table, Harvey's Legal Agent Benchmark, Gemini 4 Argon19.6%
Google's table, Harvey's Legal Agent Benchmark, GPT-6 Astra5.4%
Google's table, Harvey's Legal Agent Benchmark, Fable 5.16.7%
Google's table, Harvey's Legal Agent Benchmark, Opus 5.53.8%
Google's table, DeepSWE v1.1, Gemini 4 Argon77.9%
Google's table, DeepSWE v1.1, GPT-6 Astra74.1%
Google's table, DeepSWE v1.1, Fable 5.167.4%
Google's table, DeepSWE v1.1, Opus 5.574.2%
Google's table, FrontierSWE v2, Gemini 4 Argon55.0%
Google's table, FrontierSWE v2, GPT-6 Astra65.5%
Google's table, FrontierSWE v2, Fable 5.156.3%
Google's table, FrontierSWE v2, Opus 5.562.3%
Google's table, Vibe Code Bench, Gemini 4 Argon91.9%
Google's table, Vibe Code Bench, GPT-6 Astra89.6%
Google's table, Vibe Code Bench, Fable 5.190.3%
Google's table, Vibe Code Bench, Opus 5.590.3%
Google's table, Terminal-bench 4.0, Gemini 4 Argon57.4%
Google's table, Terminal-bench 4.0, GPT-6 Astra58.2%
Google's table, Terminal-bench 4.0, Fable 5.157.9%
Google's table, Terminal-bench 4.0, Opus 5.566.4%
Google's table, PostTrainBench, Gemini 4 Argon45.3%
Google's table, PostTrainBench, GPT-6 Astra44.3%
Google's table, PostTrainBench, Fable 5.140.2%
Google's table, PostTrainBench, Opus 5.549.3%
Google's table, Terminal-Bench Science 0.1 Science, Gemini 4 Argon57.6%
Google's table, Terminal-Bench Science 0.1 Science, GPT-6 Astra68.1%
Google's table, Terminal-Bench Science 0.1 Science, Fable 5.152.6%
Google's table, Terminal-Bench Science 0.1 Science, Opus 5.563.3%
Google's table, LABBench 2 Science, Gemini 4 Argon88.8%
Google's table, LABBench 2 Science, GPT-6 Astra85.4%
Google's table, LABBench 2 Science, Fable 5.168.6%
Google's table, LABBench 2 Science, Opus 5.573.1%
Google's table, RiemannBench Science, Gemini 4 Argon76.0%
Google's table, RiemannBench Science, GPT-6 Astra72.0%
Google's table, RiemannBench Science, Fable 5.165.6%
Google's table, RiemannBench Science, Opus 5.569.6%
Google's table, GraphWalks, Up to 128k, Gemini 4 Argon99.7%
Google's table, GraphWalks, Up to 128k, GPT-6 Astra98.7%
Google's table, GraphWalks, Up to 128k, Fable 5.191.4%
Google's table, GraphWalks, Up to 128k, Opus 5.590.6%
Google's table, GraphWalks, 256k to 1M, Gemini 4 Argon84.2%
Google's table, GraphWalks, 256k to 1M, GPT-6 Astra71.8%
Google's table, GraphWalks, 256k to 1M, Fable 5.165.0%
Google's table, GraphWalks, 256k to 1M, Opus 5.566.8%
Google's table, Agent's Last Exam, Gemini 4 Argon39.5%
Google's table, Agent's Last Exam, GPT-6 Astra34.2%
Google's table, Agent's Last Exam, Opus 5.538.2%
Google's table, OSWorld-2.0, Gemini 4 Argon69.2%
Google's table, OSWorld-2.0, GPT-6 Astra72.6%
Google's table, Chartography, Gemini 4 Argon71.6%
Google's table, Chartography, GPT-6 Astra71.0%
Google's table, Chartography, Fable 5.146.2%
Google's table, Chartography, Opus 5.566.3%
Google's table, LVBench, Gemini 4 Argon91.7%
Google's table, LVBench, GPT-6 Astra87.5%
Google's table, LVBench, Fable 5.179.7%
Google's table, LVBench, Opus 5.583.7%
Google's table, CWE-bench, Gemini 4 Argon68.0%
Google's table, CWE-bench, GPT-6 Astra68.0%
Google's table, CWE-bench, Fable 5.158.0%
Google's table, CWE-bench, Opus 5.567.0%
benchmarks in Google's table with an Opus 5.5 score, GraphWalks's two lengths counted once (our arithmetic)17
of those, the ones where Argon's score is higher than Opus 5.5's (our arithmetic)13

Google

the verifier timeout Google gave its own Terminal-Bench Science run of Argon, against the usual one6x

Vals AI

Vals AI, Vals Index, Gemini 4 Argon at high68.9%
Vals AI, Finance Agent v2, Gemini 4 Argon at high65.4%
Vals AI, Harvey's Legal Agent Benchmark, Gemini 4 Argon at high19.6%
Vals AI, CyberBench v1.1, Gemini 4 Argon at high77.9%
Vals AI, Terminal-Bench 4.0, Gemini 4 Argon at high57.6%
Vals AI, ProgramBench (fully resolved), Gemini 4 Argon at high2.5%
Vals AI, Vibe Code Bench v1.1, Gemini 4 Argon at high91.9%
Vals AI, Code Migration, Gemini 4 Argon at high68.2%
Vals AI, SRE Bench, Gemini 4 Argon at high44.3%
Vals AI, IOI, Gemini 4 Argon at high100.0%
Vals AI, Terminal-Bench Science, Gemini 4 Argon at high44.3%
Vals Index, cost per test, Argon at high (priced at $4 / $20)$15.68
the input price Vals AI priced Argon at$4
the output price Vals AI priced Argon at$20
Vals AI, Finance Agent v2, Argon at high minus Opus 5.5 at max (percentage points) (our arithmetic)+6.8
Vals AI, Harvey's Legal Agent Benchmark, Argon at high minus Opus 5.5 at max (percentage points) (our arithmetic)+15.8
Vals AI, CyberBench v1.1, Argon at high minus Opus 5.5 at max (percentage points) (our arithmetic)+22.5
Vals AI, Terminal-Bench 4.0, Argon at high minus Opus 5.5 at max (percentage points) (our arithmetic)−7.6
Vals AI, ProgramBench (fully resolved), Argon at high minus Opus 5.5 at max (percentage points) (our arithmetic)−16.0
Vals AI, Vibe Code Bench v1.1, Argon at high minus Opus 5.5 at max (percentage points) (our arithmetic)+1.6
Vals AI, Code Migration, Argon at high minus Opus 5.5 at max (percentage points) (our arithmetic)+1.5
Vals AI, SRE Bench, Argon at high minus Opus 5.5 at max (percentage points) (our arithmetic)+10.7
Vals AI, IOI, Argon at high minus Opus 5.5 at max (percentage points) (our arithmetic)+4.9
Vals Index cost per test, Argon at the full price against Opus 5.5 at max (our arithmetic)0.49
Vals AI, Terminal-Bench 4.0: Opus 5.5 at max's lead over Argon in tasks, of the 198 (our arithmetic)15

Vals AI

Vals AI, Vals Index, Opus 5.5 at max (with fallback)67.0%
Vals AI, Finance Agent v2, Opus 5.5 at max (with fallback)58.6%
Vals AI, Harvey's Legal Agent Benchmark, Opus 5.5 at max (with fallback)3.8%
Vals AI, CyberBench v1.1, Opus 5.5 at max (with fallback)55.4%
Vals AI, Terminal-Bench 4.0, Opus 5.5 at max (with fallback)65.2%
Vals AI, ProgramBench (fully resolved), Opus 5.5 at max (with fallback)18.5%
Vals AI, Vibe Code Bench v1.1, Opus 5.5 at max (with fallback)90.3%
Vals AI, Code Migration, Opus 5.5 at max (with fallback)66.7%
Vals AI, SRE Bench, Opus 5.5 at max (with fallback)33.6%
Vals AI, IOI, Opus 5.5 at max (with fallback)95.1%
Vals AI, Terminal-Bench Science, Opus 5.5 at max (with fallback)47.1%
Vals Index, cost per test, Opus 5.5 at max$32.14
Vals AI's 22 September note: Opus 5.5's Terminal-Bench 4.0 score then61.6%
the same, with fallback-assisted tasks counted as failures53.5%
fallback-assisted Terminal-Bench 4.0 tasks in that note30
Terminal-Bench 4.0 tasks198
Vals AI's note: Opus 5.5's SRE Bench score with fallback-assisted tasks counted as failures5.3%

Vals AI

Vals AI, Vals Index: share of Opus 5.5's tasks answered by an older Claude model after a refusal4.0%
Vals AI, Code Migration: share of Opus 5.5's tasks answered by an older Claude model after a refusal81.5%
Vals AI, CyberBench v1.1: share of Opus 5.5's tasks answered by an older Claude model after a refusal52.6%
Vals AI, Finance Agent v2: share of Opus 5.5's tasks answered by an older Claude model after a refusal1.3%
Vals AI, SRE Bench: share of Opus 5.5's tasks answered by an older Claude model after a refusal82.8%
Vals AI, Vibe Code Bench v1.1: share of Opus 5.5's tasks answered by an older Claude model after a refusal8.0%
Vals AI, IOI: share of Opus 5.5's tasks answered by an older Claude model after a refusal5.6%
Vals AI, ProgramBench (fully resolved): share of Opus 5.5's tasks answered by an older Claude model after a refusal12.0%
Vals AI, Terminal-Bench 4.0: share of Opus 5.5's tasks answered by an older Claude model after a refusal11.1%
Vals AI, Terminal-Bench Science: share of Opus 5.5's tasks answered by an older Claude model after a refusal2.9%
Vals AI, Terminal-Bench 4.0: Opus 5.5 tasks answered by an older Claude model after a refusal22
Vals AI, CyberBench v1.1: Opus 5.5 tasks answered by an older Claude model after a refusal61

Arena

Arena, text leaderboard (as of 30 Sep) · read 1 Oct 2026, 07:13 CEST
Arena text leaderboard, Gemini 4 Argon at high: rating (preliminary)1525
its interval, plus or minus9
Arena text leaderboard, Gemini 4 Argon's place#1
Arena text leaderboard, Opus 5.5 at high: rating1504
Arena text leaderboard, Opus 5.5 at high's place#4

Arena

Arena, WebDev leaderboard (as of 30 Sep) · read 1 Oct 2026, 07:13 CEST
Arena WebDev leaderboard, Opus 5.5 at max: rating1818
Arena WebDev leaderboard, Opus 5.5 at max's place#1
Arena WebDev leaderboard, Gemini 4 Argon at high: rating (preliminary)1679
Arena WebDev leaderboard, Gemini 4 Argon's place#8

Model Fatigue (model-economics' probe)

model-economics' probe of Google's Vertex AI for two Argon model ids · probed 1 Oct 2026, 07:07 CEST
the response code Vertex AI gave for both Argon model ids the probe tried, in both regions404

YouTube channels, as titled

Sources

These are the pages the video and this page draw on. We keep a copy of each page as we read it, so a figure can be checked against what the page said at the time.

Credits

The narration in the video is an AI voice, made with ElevenLabs.

Google's benchmark table appears in the video redrawn from the one on its Gemini model page. This page gives the same values as a table.

Music in the video: "Airport Lounge" by Kevin MacLeod (incompetech.com), licensed under Creative Commons: By Attribution 4.0.