This is the video written out, with every figure in full, each linked to its source in the table below.
Beats Opus at half the price?▶ 0:00
On Tuesday, 29 September, OpenAI put out GPT-6.1 Sol. On Wednesday, 30 September, Google announced Gemini 4 Argon, and nobody outside a small group can use it yet. It is Google's first big model after a summer of smaller Flash releases (Gemini 3.6 Flash on 21 July, Gemini 3.7 Flash on 13 August and Gemini 3.8 Flash on 2 September, by Artificial Analysis's dates), and Google's own table puts it ahead of Opus 5.5 at half Opus's price per token.
That table compares Argon with Anthropic's Claude Opus 5.5 on 17 benchmarks that have an Opus score, counting GraphWalks's two context lengths once, and Argon has the higher score on 13 of them. For Opus, Google reports the highest setting it could find a result for. Some launch videos made that their headline, like "BREAKING: Gemini 4 Argon Beats GPT-6 & Claude on 12 Tests (Half Opus Price)" and "Google Is Back: Gemini 4 Argon Beats Opus 5.5 at Half the Price".
We can't run Argon ourselves. A probe of Google's Vertex AI for the two Argon model ids it tried got a 404 for both ("not found or your project does not have access to it"), and Google's API model list and pricing page don't list it. So nothing on this page is our measurement. Three independent testers have published broad results of their own: Artificial Analysis, Vals AI, and the Arena, where people vote on answers. This page asks what the video asks: does Argon beat Opus, at what, and is it really half the price?
The basics▶ 0:59
Google DeepMind announced Gemini 4 Argon on 30 September. It is going first to trusted cyber defenders in Google's Fairwind program and to early testers. Google says paid API customers and Google AI Ultra subscribers come next, and gives no date.
The price is $2 per million input tokens and $10 per million output tokens, with cached input at 95% off the input price, so $0.10. That's an introductory price, and Google doesn't say how long it lasts. After it, Argon costs $4 in and $20 out, the same as Opus 5.5.
Price per million tokens
API list prices
| Gemini 4 Argon, introductory | Gemini 4 Argon, after | Opus 5.5 | GPT-6.1 Sol | |
|---|---|---|---|---|
| input | $2 | $4 | $4 | $2 |
| cached input | $0.10 | not stated | $0.20 | $0.10 |
| output | $10 | $20 | $20 | $10 |
Google says Argon can write up to 1M tokens in one answer, up from 64K in earlier Gemini models. On Artificial Analysis's Intelligence Index it scores 52.6, 11.7 points more than Gemini 3.8 Flash's 40.9, which was Google's best model on that index until now.
The overall score▶ 1:53
Artificial Analysis's Intelligence Index is an independent score built from ten tests. Artificial Analysis ran Argon at its high setting, the one setting Vals and the Arena used as well, and has no other Argon row yet.
Argon scores 52.6, at $1.99 per index task. Opus 5.5 at its own high setting scores 53.6, for $1.82: a point higher, for $0.17 less. At max, Opus reaches 57.6, 5.0 points above Argon. Argon sits level with OpenAI's GPT-6 Astra at max, 52.7.
Argon at high against Opus 5.5 at every setting
Intelligence Index score against cost per task. Argon has been run at high only. The yellow lines join the two pairs the text compares.
Artificial Analysis and Vals both run Opus 5.5 with a fallback: when it refuses a task, Anthropic's servers usually pass the task to an older Claude model, and the result counts as Opus 5.5's. That applies to every Opus 5.5 score from Artificial Analysis and Vals on this page. Artificial Analysis labels every Opus 5.5 row "Default Fallback" without publishing how many tasks that was on the index. The table in the coding section below gives the published share for each test on this page, and goes into what it changes.
Where Argon is ahead▶ 2:39
The index averages ten tests, and that hides where each model wins. These are the results where Argon is ahead that Google didn't produce itself.
On Harvey's Legal Agent Benchmark, which Vals runs, Argon scores 19.6% against 3.8% for Opus at max. On Vals' Finance Agent test, multi-step financial research, it scores 65.4% against 58.6%. On Artificial Analysis's own run of AutomationBench, business workflows across apps scored with partial credit, it scores 77.5% against 69.5% for Opus at max. That is the highest AutomationBench score of the 175 model-and-setting rows on Artificial Analysis's list. Google's table quotes Zapier's AutomationBench leaderboard instead, where Argon has 51.3% from Zapier's own run, not Artificial Analysis's.
On long software-engineering tasks, Google's headline claim holds up. Google reports 77.9% on DeepSWE from its own run. Artificial Analysis's coding-agent run, each model working inside its maker's own tool, gives Argon 78.8% against 68.4% for Claude Code with Opus 5.5 at max.
Where Argon is ahead
Results Google didn't produce itself: Argon at high, Opus 5.5 at max
People also prefer Argon's answers on the Arena's text leaderboard, where it is ranked #1 (1525, plus or minus 9), against 1504 for Opus 5.5 at high, ranked #4. The Arena still marks Argon's rating as preliminary.
Where Opus is ahead▶ 3:48
In a terminal, Opus leads. On Vals' Terminal-Bench 4.0, Opus at max scores 65.2% to Argon's 57.6%, and Google's own table shows a similar gap, 66.4% against 57.4%. On ProgramBench, which asks a model to rebuild a program from scratch, Opus fully resolves 18.5% of the tasks and Argon 2.5%.
On the Arena's web-development board, Opus 5.5 at max is ranked #1 (1818) and Argon #8 (1679, preliminary).
On Artificial Analysis's two tests of finished work, Opus at high is clearly ahead of Argon at high. AA-Briefcase has a model produce deliverables like spreadsheets, presentations and memos: Opus 1704, Argon 1494. On GDPval-AA, where an AI judge compares the work, Opus scores 1692 and Argon 1611.
Where Opus is ahead
Argon at high, Opus 5.5 at max
Bloomberg reported on 1 October, citing people with direct access, that while Gemini 4 "has performed well on benchmarks widely used to gauge model efficacy, it does less well when employees actually put it to work", and that the model "struggles to handle certain coding tasks". Its sources asked not to be named. Google pushed back, as Finimize reported it, "saying it's inaccurate to characterize the model as underperforming in coding".
Coding, all together▶ 4:42
So who is better at coding? It's mixed.
On four other coding tests Vals runs, Argon is ahead of Opus at max: narrowly on Vibe Code Bench v1.1 (+1.6 points), Code Migration (+1.5) and IOI (+4.9), and clearly on SRE Bench (+10.7).
Artificial Analysis's Coding Agent Index is three tests. Argon wins DeepSWE, and Opus wins the other two, SWE-Atlas-QnA and Artificial Analysis's run of Terminal-Bench v4. Overall, Claude Code with Opus 5.5 at max scores 66.0 and Argon in Google's own Antigravity CLI scores 63.8, 2.2 points lower.
Artificial Analysis's Coding Agent Index
Each model in its maker's own agent
| Claude Code, Opus 5.5 at max | Antigravity CLI, Gemini 4 Argon at high | |
|---|---|---|
| index | 66.0 | 63.8 |
| DeepSWE v1.1 | 68.4% | 78.8% |
| SWE-Atlas-QnA | 66.4% | 56.5% |
| Terminal-Bench v4 | 63.1% | 56.1% |
| cost per task | $13.04 | $5.84 |
| time per task | 64 min | 35 min |
| attempts that fell back to an older model | 78 of 909 | 0 of 905 |
Then there is the fallback. Counted on its own, a refused task would score nothing, so Opus's own scores on these tests would be lower than what's published. How much lower depends on the share of tasks an older model answered:
How much of Opus 5.5's score an older Claude model answered
Opus 5.5 at max, Argon at high. The share is of Opus's tasks (Vals) or attempts (Artificial Analysis) passed to Claude Opus 5 or Opus 4.8 after a refusal
| test | Opus 5.5 | Gemini 4 Argon | answered by an older model |
|---|---|---|---|
| Vals Index | 67.0% | 68.9% | 4.0% |
| Vals, Finance Agent | 58.6% | 65.4% | 1.3% |
| Vals, Terminal-Bench 4.0 | 65.2% | 57.6% | 11.1% |
| Vals, ProgramBench | 18.5% | 2.5% | 12.0% |
| Vals, Vibe Code Bench v1.1 | 90.3% | 91.9% | 8.0% |
| Vals, Code Migration | 66.7% | 68.2% | 81.5% |
| Vals, SRE Bench | 33.6% | 44.3% | 82.8% |
| Vals, IOI | 95.1% | 100.0% | 5.6% |
| Vals, CyberBench v1.1 | 55.4% | 77.9% | 52.6% |
| Vals, Terminal-Bench Science | 47.1% | 44.3% | 2.9% |
| AA agent index, DeepSWE v1.1 | 68.4% | 78.8% | 0.6% |
| AA agent index, SWE-Atlas-QnA | 66.4% | 56.5% | 14.0% |
| AA agent index, Terminal-Bench v4 | 63.1% | 56.1% | 12.1% |
On the terminal tests that share is 11.1% of Vals' tasks and 12.1% of Artificial Analysis's attempts. On Code Migration and SRE Bench it is 81.5% and 82.8%, so those two say more about the older models than about Opus 5.5. Vals' own note shows the size of it on SRE Bench: counted as failures, Opus's score falls from 33.6% to 5.3%. On DeepSWE in the agent index it is 0.6%, and Argon's runs show none.
Opus's lead on Vals' Terminal-Bench 4.0 is about 15 tasks of 198, against 22 that an older model answered. When Vals first published Opus 5.5 on 22 September, its note said that counting the 30 fallback-assisted tasks it had then (of 198) as failures moved Opus's Terminal-Bench 4.0 score from 61.6% to 53.5%, below Argon's 57.6% today. Vals' page now shows Opus at 65.2%, and the matching fallback-free value sits behind a toggle on that page which our saved copy doesn't carry. In the agent index, Opus's 2.2 points lead compares with 8.6% of its attempts falling back. So Opus's lead in a terminal, and its lead on the coding-agent index, aren't settled.
How often it guesses▶ 5:57
Google's launch post doesn't mention it, but the independent numbers show how rarely Argon guesses when it doesn't know. AA-Omniscience is a knowledge test where a model can answer, or say it doesn't know; a right answer adds to the score, a wrong one takes away from it, and not answering does neither.
Opus 5.5 at high gets 64.6% of the questions right and 24.0% wrong. Argon gets 49.9% right and only 7.6% wrong; the remaining 42.5% it doesn't answer, or answers only in part. It knows less, and it bluffs much less, so the test's score comes out about the same: 42.4 for Argon, 40.6 for Opus. We split the shares out of Artificial Analysis's accuracy and score; the split reproduces the hallucination rate Artificial Analysis publishes.
Right, wrong, and not answered
AA-Omniscience, both models at high
Artificial Analysis's hallucination rate is the wrong answers as a share of everything a model didn't get right. Argon's is 15.1%. Among the 34 model-and-setting rows on Artificial Analysis's list that score above forty-five on the index, none is lower; the next is Qwen3.8 Max (0902) at 28.8%. In legal and finance work, where a confident wrong answer does the most damage, that could matter more than the test's score.
How often each model guesses when it doesn't know
Hallucination rate: wrong answers as a share of everything not answered right
Half the price?▶ 6:49
Per token, Argon is half of Opus, for now. Per task, it depends on which Opus you compare it with.
Against Opus at max, Argon at high costs a third as much on Artificial Analysis's index ($1.99 against $5.98) and scores 5.0 points less. In the coding-agent index it's $5.84 a task against $13.04, under half (0.45 of the cost).
Against Opus at the same quality, the gap goes away: Opus at high scores a point more than Argon, for $0.17 less per task. Argon writes more to get there, 61.6K output tokens per index task against 35.6K for Opus at high, and spends more on the input side too, $1.37 against $1.11.
All of those Argon prices are introductory. When that period ends, every token price doubles. Google states the cached-input discount only for the introductory price, so we give a range: with the cached rate doubled too, or left where it is. We price cache writes at the input rate in both, as Artificial Analysis does at the introductory price. By our arithmetic, Argon then costs $3.66 to $3.98 per index task, about twice Opus at high (2.01 to 2.18 times its cost), and the coding agent's $5.84 becomes $10.51 to $11.70, close to Opus at max's $13.04. We checked the method on the agent run first: Argon's mean tokens at the introductory price come to $5.85, which matches Artificial Analysis's figure.
Per task, at the introductory price and after it
Argon at high against the Opus setting each comparison uses
| comparison | Gemini 4 Argon now | Gemini 4 Argon after the introductory price | Opus 5.5 | scores |
|---|---|---|---|---|
| AA index, Opus at max | $1.99 | $3.66 to $3.98 | $5.98 | 52.6 against 57.6 |
| AA index, Opus at high | $1.99 | $3.66 to $3.98 | $1.82 | 52.6 against 53.6 |
| AA coding agent, Opus at max | $5.84 | $10.51 to $11.70 | $13.04 | 63.8 against 66.0 |
| Vals Index, Opus at max | already at the full price | $15.68 | $32.14 | 68.9% against 67.0% |
Vals priced Argon at the full price, $4 in and $20 out, from the start. Its Vals Index costs $15.68 per test for Argon against $32.14 for Opus at max, about half, on a different mix of tasks from Artificial Analysis's. Google describes the Vals Index as spanning finance, coding, legal and tax work, each weighted by its share of US GDP, and Argon scores higher on it (68.9% against 67.0%).
Google's table, checked▶ 7:49
Google's methodology page says where each row of its table comes from. The Vals Index, Finance Agent and Harvey's legal benchmark rows are sourced from Vals AI, Vibe Code Bench from Vals' public leaderboard, AutomationBench from Zapier's public leaderboard, and several more from public leaderboards. Argon's DeepSWE and Terminal-Bench 4.0 scores are Google's own runs.
Google's benchmark table, with where each row comes from
Argon against GPT-6 Astra, Fable 5.1 and Opus 5.5, as Google published it. Marked rows: Opus scores higher
| benchmark | Gemini 4 Argon | GPT-6 Astra | Fable 5.1 | Opus 5.5 | where Google's numbers come from |
|---|---|---|---|---|---|
| Vals Index | 68.9% | 63.1% | 65.8% | 67.0% | Vals AI |
| AutomationBench | 51.3% | 41.4% | 31.4% | 42.5% | Zapier's public leaderboard |
| Vals Finance Agent | 65.4% | 53.5% | 58.9% | 58.6% | Vals AI |
| Harvey's Legal Agent Benchmark | 19.6% | 5.4% | 6.7% | 3.8% | Vals AI |
| DeepSWE v1.1 | 77.9% | 74.1% | 67.4% | 74.2% | Argon: Google's own run; Astra: public leaderboard; Fable and Opus: their system cards |
| FrontierSWE | 55.0% | 65.5% | 56.3% | 62.3% | Proximal's public leaderboard |
| Vibe Code Bench | 91.9% | 89.6% | 90.3% | 90.3% | Vals AI's public leaderboard |
| Terminal-Bench 4.0 | 57.4% | 58.2% | 57.9% | 66.4% | Argon: Google's own run; others: public leaderboard |
| PostTrainBench | 45.3% | 44.3% | 40.2% | 49.3% | Google's runs for every model |
| Terminal-Bench Science | 57.6% | 68.1% | 52.6% | 63.3% | Argon: Google's own run, with a longer verifier timeout; others: public leaderboard |
| LABBench 2 | 88.8% | 85.4% | 68.6% | 73.1% | Google's runs for every model |
| RiemannBench | 76.0% | 72.0% | 65.6% | 69.6% | Surge's public leaderboard |
| GraphWalks, shorter contexts | 99.7% | 98.7% | 91.4% | 90.6% | Google's runs for every model |
| GraphWalks, longest contexts | 84.2% | 71.8% | 65.0% | 66.8% | Google's runs for every model |
| Agent's Last Exam | 39.5% | 34.2% | no score | 38.2% | Argon: Google's own run; others: public leaderboard |
| OSWorld-2.0 | 69.2% | 72.6% | no score | no score | Argon: Google's own run; Astra: OpenAI's post |
| Chartography | 71.6% | 71.0% | 46.2% | 66.3% | Surge's public leaderboard |
| LVBench | 91.7% | 87.5% | 79.7% | 83.7% | Google's runs for every model |
| CWE-bench | 68.0% | 68.0% | 58.0% | 67.0% | public leaderboard |
Argon's Vals numbers in the table match Vals' own pages: the Vals Index 68.9%, Finance Agent 65.4%, Harvey's benchmark 19.6% and Vibe Code Bench v1.1 91.9%, rounded. DeepSWE, the headline number, is Google's own run, and Artificial Analysis's agent run gives a similar score, 78.8% against Google's 77.9%.
On Terminal-Bench Science, Google's figure for Argon, 57.6%, is "self computed, with 6x verifier timeout to address timeout issues with verification", according to the methodology page. Vals runs the same test and has Argon at 44.3% and Opus 5.5 at max at 47.1%, with its own caution that its numbers are "not directly comparable to the official" leaderboard, which pairs each model with its own agent. So the gap between Google's figure and Vals' says little about Argon on its own.
GPT-6.1 Sol isn't in the table, but it came out the day before Argon.
Against the rest▶ 8:12
GPT-6 Astra at max scores 52.7, level with Argon, at $3.26 a task, so for now Argon costs about sixty percent of what Astra does (0.61 times). GPT-6.1 Sol launched at the same price per token as Argon's introductory price ($2 in, $0.10 cached, $10 out). At max it scores 51.8, within a point of Argon, for $0.72 a task, a little over a third of Argon's cost (0.36 times). In Codex at xhigh it scores 62.9 on the coding-agent index for $1.04 a task, against Argon's 63.8 for $5.84. Anthropic's Fable 5.1 at max, which Artificial Analysis also runs with fallback, scores 53.4, for $7.63.
Against the rest of the top tier
Intelligence Index score against cost per task, four models at the settings Artificial Analysis lists; Fable 5.1 is in the text above. The yellow lines join the two pairs the text compares.
The numbers in this chart
| setting | GPT-6.1 Sol score | GPT-6.1 Sol cost per index task | GPT-6 Astra score | GPT-6 Astra cost per index task | Opus 5.5 score | Opus 5.5 cost per index task | Gemini 4 Argon score | Gemini 4 Argon cost per index task |
|---|---|---|---|---|---|---|---|---|
| low | 42.1 | $0.13 | 45.8 | $0.82 | 42.3 | $0.55 | – | – |
| medium | 47.8 | $0.21 | 49.6 | $1.54 | 51.2 | $1.34 | – | – |
| high | 50.2 | $0.32 | 50.9 | $1.73 | 53.6 | $1.82 | 52.6 | $1.99 |
| xhigh | 51.0 | $0.39 | 52.4 | $2.31 | 56.0 | $3.46 | – | – |
| max | 51.8 | $0.72 | 52.7 | $3.26 | 57.6 | $5.98 | – | – |
Why it's gated▶ 8:59
The launch pages we read don't link a system card for Argon. Google's methodology page links Anthropic's Opus 5.5 system card, not one for Argon. So this part is Google's own account.
Google says that "safely releasing frontier capabilities at this level requires a phased approach", starting with cyber defenders, and that it is "actively engaged in the U.S. government's voluntary process for pre-release model access". For trusted defenders and Google's own teams, it is releasing Argon "without cyber guardrails". Before a wider release, Google describes four kinds of safeguard: defending against misuse, defending against prompt injection attacks, monitoring for misalignment (watching Argon's chain of thought and actions and stopping execution when necessary), and hardening the environments it is tested in.
Two numbers here don't come from Google's own runs. On Vals' CyberBench v1.1, Argon scores 77.9% against 55.4% for Opus at max, and Opus refused 52.6% of those tasks (61), which older Claude models then answered. On CWE-bench, which tests fixing security vulnerabilities and whose public leaderboard Google quotes, it is 68.0% against 67.0%.
Should you switch?▶ 9:56
You can't yet. For when Argon opens to paid API customers and Google AI Ultra subscribers, this is what the evidence says so far.
For legal, finance and business-automation agents, it's worth trying first: independent tests put it well ahead of Opus there, and it guesses wrong far less often. For coding it's mixed, so test both on your own work before you move. Budget for the full price, not the introductory one. If you want something close today, GPT-6.1 Sol scores within a point of Argon on Artificial Analysis's index for a little over a third of the cost per task.
Where the evidence points, so far
For when Argon opens to paid API customers and Google AI Ultra subscribers
| work | what the independent tests show |
|---|---|
| legal and finance agents | Argon well ahead on Vals' tests (19.6% against 3.8%; 65.4% against 58.6%) |
| business workflows across apps | Argon ahead on Artificial Analysis's AutomationBench (77.5% against 69.5%) |
| coding | mixed: Argon ahead on DeepSWE, Opus in a terminal and on ProgramBench, and part of Opus's scores are older models' answers |
| finished work: spreadsheets, presentations, memos | Opus at high ahead on both of Artificial Analysis's tests |
| web front ends | Opus ranked #1 on the Arena's board, Argon #8 |
| not measured yet | real use outside Google, Argon at settings other than high on Artificial Analysis, Vals and the Arena, how long the introductory price lasts |
Hold all of this loosely. It's under a day of evidence, and Artificial Analysis, Vals and the Arena each ran Argon at high only. When model-economics' sweep stopped looking, early on 1 October, it hadn't found a verified write-up from anyone outside Google using Argon for real work. Real use, the other settings, and how long the introductory price lasts would settle it.
Just before this video, we looked at GPT-6.1 Sol, OpenAI's release from the day before. That's our previous video.