The short answer
On Artificial Analysis's Intelligence Index, read on 4 October 2026, Claude Opus 5.5 scores higher than Claude Fable 5.1 at every reasoning setting from medium to max, by 2.3 points to 4.3 points, and costs less per task at every setting. Fable leads only at low, where it costs 4.3× as much.
Intelligence Index at each reasoning setting
Artificial Analysis Intelligence Index, Opus 5.5 outlined and Fable 5.1 filled; the column on the right is Fable minus Opus
The numbers in this chart
| Opus 5.5 | Fable 5.1 | Fable minus Opus | |
|---|---|---|---|
| low | 42.3 | 46.8 | +4.5 points |
| medium | 51.2 | 48.9 | −2.3 points |
| high | 53.6 | 51.2 | −2.4 points |
| xhigh | 56.0 | 53.2 | −2.8 points |
| max | 57.6 | 53.4 | −4.3 points |
Fable's leads are in a few places. On Artificial Analysis it is ahead on Harvey LAB-AA, a legal agent test, at every setting. Vals.ai's published scores show eight clear leads, four each, but both models ran there with other Claude models answering when they refused. Two of the eight are shown to stay clear with the most those fallback answers could have added taken off the leader's score, and both are Fable's: US corporate tax questions and SNAP benefits questions. A third, Opus's on hard terminal tasks, holds on the page's own score with both models' fallback answers counted as failures.
Every score and cost here is Artificial Analysis's or Vals.ai's. What we added is the setting-by-setting comparison, the margin test and the arithmetic.
Setting by setting
Claude Fable 5.1 lists at $10 per million input tokens and $50 per million output tokens, against $4 and $20 for Claude Opus 5.5, so both prices are 2.5× as high. Artificial Analysis runs both models at each of their five reasoning settings, which lets us put them side by side at the same setting.
At low, Fable scores 4.5 points higher than Opus, and it costs 4.3× as much per task, because at that setting it also writes 2.1× as many output tokens. From medium up the order reverses. Opus is ahead by 2.3 points at medium and by 4.3 points at max, and Fable still costs more per task at every setting.
The gap in cost narrows as the setting rises, from 2.2× at medium to 1.3× at max. Opus's output grows faster as the setting rises: at max it writes 119K output tokens per task against Fable's 78K, so Fable's higher price per token is partly made up by writing less.
Setting by setting on Artificial Analysis
Score and cost per task on the Intelligence Index, and what Fable costs and writes per task against Opus
| setting | Opus score | Fable score | Opus cost | Fable cost | Fable's cost | Fable's output tokens |
|---|---|---|---|---|---|---|
| low | 42.3 | 46.8 | $0.55 | $2.37 | 4.3× | 2.1× |
| medium | 51.2 | 48.9 | $1.34 | $2.98 | 2.2× | 1.1× |
| high | 53.6 | 51.2 | $1.82 | $3.91 | 2.1× | 1.1× |
| xhigh | 56.0 | 53.2 | $3.46 | $5.98 | 1.7× | 0.9× |
| max | 57.6 | 53.4 | $5.98 | $7.63 | 1.3× | 0.7× |
Test by test
The Intelligence Index averages ten tests, and an average can hide a test where the other model leads. Artificial Analysis publishes the score on each one, and on a few more tests outside the index, so we lined the two models up on every test its page scores both of them on at each setting.
Test by test on Artificial Analysis: Fable minus Opus
Each test both models have at a setting; a plus means Fable is ahead. Percentage points, except Elo points for AA-Briefcase and GDPval-AA and index points for AA-Omniscience
| test | low | medium | high | xhigh | max |
|---|---|---|---|---|---|
| AA-Briefcase (Elo) | +202 | −98 | −108 | −111 | −132 |
| GDPval-AA (Elo) | +233 | −36 | −72 | −102 | −109 |
| AutomationBench-AA | −0.7 | −6.6 | −7.9 | −7.2 | −10.2 |
| Terminal-Bench 4.0 | +9.1 | −7.6 | −4.5 | −4.5 | −7.6 |
| SciCode | −1.9 | −2.9 | −1.7 | −4.2 | −3.8 |
| Humanity's Last Exam | +0.6 | −0.9 | +0.4 | +1.2 | −2.2 |
| GDP.pdf | +2.4 | +1.2 | −2.0 | −0.4 | 0.0 |
| CritPt | +10.0 | +1.4 | −0.6 | −0.6 | −2.0 |
| AA-Omniscience (index) | −4.7 | −2.7 | +0.2 | −0.3 | −3.0 |
| AA-LCR | +1.7 | +0.3 | +1.0 | −1.7 | +0.7 |
| Harvey LAB-AA | +3.2 | +2.4 | +2.1 | +2.0 | +1.8 |
| Terminal-Bench Science | – | – | – | – | −15.7 |
| ITBench-AA | – | – | – | – | +11.3 |
| AA-AnalystAgent | – | – | – | – | +1.2 |
| MLCR-AA | – | – | – | – | +4.4 |
At low, Fable is ahead on 8 of the 11 tests both models have. From medium up it leads on between 2 and 5 at each setting; at max the two are level on GDP.pdf, and Opus leads on the rest.
Harvey LAB-AA, which Artificial Analysis describes as legal agentic work, is the one test where Fable is ahead at every setting, by between 1.8 points and 3.2 points. At max it scores 93.0% against Opus's 91.2%. AutomationBench-AA runs the other way: Opus is ahead at every setting, by 10.2 points at max.
The largest gaps at max are on tests the page scores Fable on only at that setting. Fable leads on ITBench-AA, Kubernetes incident root-cause analysis, by 11.3 points, and Opus leads on Terminal-Bench Science by 15.7 points.
Artificial Analysis also groups its tests into domain scores such as legal, finance and engineering. Fable is ahead on all 5 of them at low. Above low it is ahead on none, except legal at xhigh, where the two are level to one decimal at 60.8.
On Vals.ai
Vals.ai runs a suite of benchmarks, some built by its own team and some by others, and its pages for both models give the compute effort as max. Both models have results on 24 of them. Most results come with an error margin, so we counted a lead as clear only when the gap is wider than the margins shown, added together. On the published scores that gives 4 clear leads for Fable and 4 for Opus.
Vals.ai: the clear leads on the published scores
Share of each benchmark passed, Opus 5.5 outlined and Fable 5.1 filled; the column on the right is Fable minus Opus in points
The numbers in this chart
| Opus 5.5 | Fable 5.1 | Fable minus Opus | |
|---|---|---|---|
| CyberBench v1.1 | 55.36% | 70.42% | +15.1 points |
| Harvey legal agent | 3.75% | 6.67% | +2.9 points |
| Public Benefits Bench | 70.64% | 74.90% | +4.3 points |
| Tax Agent Bench | 70.50% | 77.64% | +7.1 points |
| Code Migration | 66.65% | 54.61% | −12.0 points |
| ProgramBench | 18.50% | 7.00% | −11.5 points |
| SRE Bench | 33.59% | 22.90% | −10.7 points |
| Terminal-Bench 4.0 | 65.15% | 58.08% | −7.1 points |
Fable leads clearly on CyberBench, where an agent has to find security bugs and patch them, by 15.1 points; on Tax Agent Bench, research-grade US corporate tax questions, by 7.1 points; on Public Benefits Bench, which is about helping people with SNAP benefits, by 4.3 points; and on Harvey's Legal Agent Benchmark, legal work with documents, spreadsheets and file tools, by 2.9 points, though both models pass very little of it.
Opus leads clearly on Code Migration, reimplementing working programs in another language, by 12.0 points; on ProgramBench, which asks a model to rebuild programs from scratch, by 11.5 points; on SRE Bench, working out what a binary does without its source code, by 10.7 points; and on Terminal-Bench 4.0, a set of hard terminal tasks, by 7.1 points.
The fallback models
Both Vals.ai pages say the models ran with Claude Opus 5 and Claude Opus 4.8 as server-side fallbacks for refusals, and the published scores count a task the fallback model passed as passed. Across the Vals Index the pages give a fallback rate of 3.99% for Opus and 2.10% for Fable, but on some benchmarks the share is far higher, and each row's tooltip says how high.
Three of the eight clear leads sit on benchmarks where most or half of the leader's tasks were fallback-assisted: 81.5% of Opus's tasks on Code Migration, 82.8% on SRE Bench, and on CyberBench 53.5% of Fable's and 52.6% of Opus's. So for each clear lead we took the leader's fallback share off its score, the most those answers could have added, and asked whether the lead was still wider than the margins. That is a strict test: a lead that fails it may still hold. Only 2 pass it: Fable's on Tax Agent Bench and Public Benefits Bench, where its fallback share is 0.0% and 0.4%. Of the other 6, the pages give a score with the fallbacks counted as failures for three. They give such scores for three benchmarks inside the margins too.
On SRE Bench, counting the fallback-assisted tasks as failures takes Opus from 33.59% to 5.34% and Fable from 22.90% to 10.69%, so Fable comes out ahead by 5.3 points: Opus's published lead is mostly the fallback models' work. On Harvey's Legal Agent Benchmark, Opus's tooltip lists no fallbacks, and Fable's score falls from 6.67% to 5.83%, which leaves it 2.1 points ahead, inside the margins. On Terminal-Bench 4.0, Opus's lead holds: counted the same way, Opus scores 58.08% and Fable 50.00%. On Vibe Code Bench v1.1, level on the published scores, Opus falls to 83.34% while Fable shows no fallbacks, which puts Fable 6.9 points ahead. The other two, MysteryMechanism and Legal Research Bench, stay inside the margins counted either way.
Where Vals.ai gives a score with fallbacks counted as failures
Published score, and the score with fallback-assisted tasks counted as failures, as Vals.ai's pages give them
| benchmark | Opus, published | Opus, fallbacks as failures | Fable, published | Fable, fallbacks as failures |
|---|---|---|---|---|
| SRE Bench | 33.59% | 5.34% | 22.90% | 10.69% |
| Harvey's Legal Agent Benchmark | 3.75% | the same (no fallbacks) | 6.67% | 5.83% |
| Terminal-Bench 4.0 | 65.15% | 58.08% | 58.08% | 50.00% |
| Vibe Code Bench v1.1 | 90.29% | 83.34% | 90.26% | the same (no fallbacks) |
| MysteryMechanism | 49.55% | 49.10% | 47.75% | the same (no fallbacks) |
| Legal Research Bench | 50.48% | not given | 55.29% | 54.33% |
Of the rest, the gap is inside the margins on 12 benchmarks, the Vals Index among them. On the other 4 the page gives no margin to test against: on ProofBench both score 100.00%, and on Vals RSI Index and CUA-bench Opus is about a point ahead. On Time Horizon Index: KSP, Opus scores 91.33% against Fable's 63.33%, the largest gap on the page either way, but with both margins shown as zero there is nothing to test it against.
Every benchmark both models have on Vals.ai
Vals.ai gives both models' compute effort as max. Scores in per cent as published, which count a task a fallback model passed as passed; gap and margins in percentage points; the share of tasks a fallback model helped with, from each row's tooltip
| benchmark | Opus 5.5 | Fable 5.1 | Fable minus Opus | margins added | fallback-assisted, Opus / Fable | who leads |
|---|---|---|---|---|---|---|
| Public Benefits Bench v1.1 | 70.64% | 74.90% | +4.3 | 2.3 | 1.3% / 0.4% | Fable, clear either way |
| Tax Agent Bench | 70.50% | 77.64% | +7.1 | 6.0 | 0.5% / 0.0% | Fable, clear either way |
| Code Migration | 66.65% | 54.61% | −12.0 | 9.1 | 81.5% / 1.5% | Opus, clear; could rest on fallbacks |
| CyberBench v1.1 | 55.36% | 70.42% | +15.1 | 10.6 | 52.6% / 53.5% | Fable, clear; could rest on fallbacks |
| Harvey's Legal Agent Benchmark | 3.75% | 6.67% | +2.9 | 2.6 | 0.0% / 1.7% | Fable, clear; inside the margins with fallbacks as failures |
| ProgramBench | 18.50% | 7.00% | −11.5 | 4.6 | 12.0% / 16.5% | Opus, clear; could rest on fallbacks |
| SRE Bench | 33.59% | 22.90% | −10.7 | 5.5 | 82.8% / 75.2% | Opus, clear; Fable ahead with fallbacks as failures |
| Terminal-Bench 4.0 | 65.15% | 58.08% | −7.1 | 3.3 | 11.1% / 11.1% | Opus, clear; holds on the page's fallback figure |
| EMB | 75.94% | 76.67% | +0.7 | 4.5 | 0.0% / 0.0% | inside the margins |
| Finance Agent (v2) | 58.59% | 58.88% | +0.3 | 2.2 | 1.3% / 0.0% | inside the margins |
| IOI | 95.06% | 90.78% | −4.3 | 9.6 | 5.6% / 22.2% | inside the margins |
| Legal Research Bench | 50.48% | 55.29% | +4.8 | 6.9 | 2.4% / 1.4% | inside the margins |
| MedCode | 49.80% | 53.51% | +3.7 | 4.4 | 0.0% / 0.0% | inside the margins |
| MedScribe | 91.43% | 91.29% | −0.1 | 3.9 | 0.0% / 0.0% | inside the margins |
| MysteryMechanism | 49.55% | 47.75% | −1.8 | 6.7 | 0.5% / 0.0% | inside the margins |
| SAGE | 45.83% | 48.53% | +2.7 | 6.7 | 0.0% / 0.0% | inside the margins |
| Terminal-Bench Science | 47.14% | 40.00% | −7.1 | 11.9 | 2.9% / 5.7% | inside the margins |
| Vals Index | 66.97% | 65.83% | −1.1 | 2.0 | 4.0% / 2.1% | inside the margins |
| Vibe Code Bench 1-100 | 30.36% | 28.00% | −2.4 | 9.3 | 16.0% / 28.0% | inside the margins |
| Vibe Code Bench v1.1 | 90.29% | 90.26% | 0.0 | 3.1 | 8.0% / 0.0% | inside the margins; Fable ahead with fallbacks as failures |
| CUA-bench | 14.00% | 13.17% | −0.8 | – | 0.0% / 0.0% | no margin given |
| ProofBench v1.1 | 100.00% | 100.00% | 0.0 | 0.0 | 0.0% / 0.0% | no margin given |
| Time Horizon Index: KSP | 91.33% | 63.33% | −28.0 | 0.0 | 0.0% / 0.0% | no margin given |
| Vals RSI Index | 37.31% | 36.09% | −1.2 | – | 0.0% / 0.0% | no margin given |
Which to pick
On Artificial Analysis's figures, Opus 5.5 is the cheaper and higher-scoring choice for most work from medium up, and at max Fable costs 1.3× as much per task, the smallest premium of any setting. Fable 5.1 leads at low, on Artificial Analysis's legal agent test at every setting, and on Vals.ai's tax and benefits benchmarks, the two leads there that hold however the fallback answers are counted. On Vals.ai's hard terminal tasks Opus's lead holds on the page's own figure. Fable's CyberBench lead, and Opus's leads on rebuilding and porting programs, may hold too, but on those three the leader's fallback share is larger than its lead, and the pages don't give a score without the fallbacks.
What this doesn't tell you
Every figure here is a benchmark run by someone else, on tasks that evaluator chose. The Terminal-Bench 4.0 page says its scores are the mean of three runs; the other pages we saved don't say how many runs are behind each score. A lead of a few points on a tax benchmark says how the models did on those tasks, not on your returns.
Artificial Analysis names its rows for both models "Default Fallback", but its page doesn't say what that is or how many tasks it affected, so its figures may include fallback answers of the kind Vals.ai reports.
The costs are Artificial Analysis's cost per task in dollars at list prices. Neither evaluator's page says anything about use on a subscription.
Vals.ai's pages give one setting, max, for both models, so the Vals comparison says nothing about the lower settings.
Both evaluators update their figures, and every figure is dated in the numbers table below.