model fatıgue
Comparison, updated in place

Opus 5.5 or Fable 5.1: where the dearer model still leads

Published 4 Oct 2026Data checked 4 Oct 2026Updated 4 Oct 2026Cite this readingEvery reading

The short answer

On Artificial Analysis's Intelligence Index, read on 4 October 2026, Claude Opus 5.5 scores higher than Claude Fable 5.1 at every reasoning setting from medium to max, by 2.3 points to 4.3 points, and costs less per task at every setting. Fable leads only at low, where it costs 4.3× as much.

Intelligence Index at each reasoning setting

Artificial Analysis Intelligence Index, Opus 5.5 outlined and Fable 5.1 filled; the column on the right is Fable minus Opus

Opus 5.5Fable 5.1
4045505560scoreFable minus Opuslowlow: Opus 5.5 42.3 (read 4 Oct 2026, 07:30 CEST), Fable 5.1 46.8 (read 4 Oct 2026, 07:30 CEST) · Artificial Analysis+4.5mediummedium: Opus 5.5 51.2 (read 4 Oct 2026, 07:30 CEST), Fable 5.1 48.9 (read 4 Oct 2026, 07:30 CEST) · Artificial Analysis−2.3highhigh: Opus 5.5 53.6 (read 4 Oct 2026, 07:30 CEST), Fable 5.1 51.2 (read 4 Oct 2026, 07:30 CEST) · Artificial Analysis−2.4xhighxhigh: Opus 5.5 56.0 (read 4 Oct 2026, 07:30 CEST), Fable 5.1 53.2 (read 4 Oct 2026, 07:30 CEST) · Artificial Analysis−2.8maxmax: Opus 5.5 57.6 (read 4 Oct 2026, 07:30 CEST), Fable 5.1 53.4 (read 4 Oct 2026, 07:30 CEST) · Artificial Analysis−4.3
4045505560scoreFable minus Opuslowlow: Opus 5.5 42.3 (read 4 Oct 2026, 07:30 CEST), Fable 5.1 46.8 (read 4 Oct 2026, 07:30 CEST) · Artificial Analysis+4.5mediummedium: Opus 5.5 51.2 (read 4 Oct 2026, 07:30 CEST), Fable 5.1 48.9 (read 4 Oct 2026, 07:30 CEST) · Artificial Analysis−2.3highhigh: Opus 5.5 53.6 (read 4 Oct 2026, 07:30 CEST), Fable 5.1 51.2 (read 4 Oct 2026, 07:30 CEST) · Artificial Analysis−2.4xhighxhigh: Opus 5.5 56.0 (read 4 Oct 2026, 07:30 CEST), Fable 5.1 53.2 (read 4 Oct 2026, 07:30 CEST) · Artificial Analysis−2.8maxmax: Opus 5.5 57.6 (read 4 Oct 2026, 07:30 CEST), Fable 5.1 53.4 (read 4 Oct 2026, 07:30 CEST) · Artificial Analysis−4.3
Artificial Analysis, Claude Opus 5.5 model page, read 4 October 2026. The rows Artificial Analysis names "Default Fallback" for both models.
The numbers in this chart

Fable's leads are in a few places. On Artificial Analysis it is ahead on Harvey LAB-AA, a legal agent test, at every setting. Vals.ai's published scores show eight clear leads, four each, but both models ran there with other Claude models answering when they refused. Two of the eight are shown to stay clear with the most those fallback answers could have added taken off the leader's score, and both are Fable's: US corporate tax questions and SNAP benefits questions. A third, Opus's on hard terminal tasks, holds on the page's own score with both models' fallback answers counted as failures.

Every score and cost here is Artificial Analysis's or Vals.ai's. What we added is the setting-by-setting comparison, the margin test and the arithmetic.

Setting by setting

Claude Fable 5.1 lists at $10 per million input tokens and $50 per million output tokens, against $4 and $20 for Claude Opus 5.5, so both prices are 2.5× as high. Artificial Analysis runs both models at each of their five reasoning settings, which lets us put them side by side at the same setting.

At low, Fable scores 4.5 points higher than Opus, and it costs 4.3× as much per task, because at that setting it also writes 2.1× as many output tokens. From medium up the order reverses. Opus is ahead by 2.3 points at medium and by 4.3 points at max, and Fable still costs more per task at every setting.

The gap in cost narrows as the setting rises, from 2.2× at medium to 1.3× at max. Opus's output grows faster as the setting rises: at max it writes 119K output tokens per task against Fable's 78K, so Fable's higher price per token is partly made up by writing less.

Setting by setting on Artificial Analysis

Score and cost per task on the Intelligence Index, and what Fable costs and writes per task against Opus

settingOpus scoreFable scoreOpus costFable costFable's costFable's output tokens
low42.346.8$0.55$2.374.3×2.1×
medium51.248.9$1.34$2.982.2×1.1×
high53.651.2$1.82$3.912.1×1.1×
xhigh56.053.2$3.46$5.981.7×0.9×
max57.653.4$5.98$7.631.3×0.7×
Artificial Analysis's figures; the last two columns are our arithmetic, Fable's cost and output tokens per task as a multiple of Opus's at the same setting. Marked: the one setting where Fable scores higher.

Test by test

The Intelligence Index averages ten tests, and an average can hide a test where the other model leads. Artificial Analysis publishes the score on each one, and on a few more tests outside the index, so we lined the two models up on every test its page scores both of them on at each setting.

Test by test on Artificial Analysis: Fable minus Opus

Each test both models have at a setting; a plus means Fable is ahead. Percentage points, except Elo points for AA-Briefcase and GDPval-AA and index points for AA-Omniscience

testlowmediumhighxhighmax
AA-Briefcase (Elo)+202−98−108−111−132
GDPval-AA (Elo)+233−36−72−102−109
AutomationBench-AA−0.7−6.6−7.9−7.2−10.2
Terminal-Bench 4.0+9.1−7.6−4.5−4.5−7.6
SciCode−1.9−2.9−1.7−4.2−3.8
Humanity's Last Exam+0.6−0.9+0.4+1.2−2.2
GDP.pdf+2.4+1.2−2.0−0.40.0
CritPt+10.0+1.4−0.6−0.6−2.0
AA-Omniscience (index)−4.7−2.7+0.2−0.3−3.0
AA-LCR+1.7+0.3+1.0−1.7+0.7
Harvey LAB-AA+3.2+2.4+2.1+2.0+1.8
Terminal-Bench Science––––−15.7
ITBench-AA––––+11.3
AA-AnalystAgent––––+1.2
MLCR-AA––––+4.4
Our arithmetic on Artificial Analysis's per-test scores: the ten tests of its Intelligence Index, then the other tests its page scores both models on. A dash means the page has no score for one of the two models at that setting. Marked: the one test where Fable is ahead at every setting.

At low, Fable is ahead on 8 of the 11 tests both models have. From medium up it leads on between 2 and 5 at each setting; at max the two are level on GDP.pdf, and Opus leads on the rest.

Harvey LAB-AA, which Artificial Analysis describes as legal agentic work, is the one test where Fable is ahead at every setting, by between 1.8 points and 3.2 points. At max it scores 93.0% against Opus's 91.2%. AutomationBench-AA runs the other way: Opus is ahead at every setting, by 10.2 points at max.

The largest gaps at max are on tests the page scores Fable on only at that setting. Fable leads on ITBench-AA, Kubernetes incident root-cause analysis, by 11.3 points, and Opus leads on Terminal-Bench Science by 15.7 points.

Artificial Analysis also groups its tests into domain scores such as legal, finance and engineering. Fable is ahead on all 5 of them at low. Above low it is ahead on none, except legal at xhigh, where the two are level to one decimal at 60.8.

On Vals.ai

Vals.ai runs a suite of benchmarks, some built by its own team and some by others, and its pages for both models give the compute effort as max. Both models have results on 24 of them. Most results come with an error margin, so we counted a lead as clear only when the gap is wider than the margins shown, added together. On the published scores that gives 4 clear leads for Fable and 4 for Opus.

Vals.ai: the clear leads on the published scores

Share of each benchmark passed, Opus 5.5 outlined and Fable 5.1 filled; the column on the right is Fable minus Opus in points

Opus 5.5Fable 5.1
0%20%40%60%80%100%share passedFable minus OpusCyberBench v1.1CyberBench v1.1: Opus 5.5 55.36% (read 4 Oct 2026, 07:30 CEST), Fable 5.1 70.42% (read 4 Oct 2026, 07:30 CEST) · Vals.ai+15.1Harvey legal agentHarvey legal agent: Opus 5.5 3.75% (read 4 Oct 2026, 07:30 CEST), Fable 5.1 6.67% (read 4 Oct 2026, 07:30 CEST) · Vals.ai+2.9Public Benefits BenchPublic Benefits Bench: Opus 5.5 70.64% (read 4 Oct 2026, 07:30 CEST), Fable 5.1 74.90% (read 4 Oct 2026, 07:30 CEST) · Vals.ai+4.3Tax Agent BenchTax Agent Bench: Opus 5.5 70.50% (read 4 Oct 2026, 07:30 CEST), Fable 5.1 77.64% (read 4 Oct 2026, 07:30 CEST) · Vals.ai+7.1Code MigrationCode Migration: Opus 5.5 66.65% (read 4 Oct 2026, 07:30 CEST), Fable 5.1 54.61% (read 4 Oct 2026, 07:30 CEST) · Vals.ai−12.0ProgramBenchProgramBench: Opus 5.5 18.50% (read 4 Oct 2026, 07:30 CEST), Fable 5.1 7.00% (read 4 Oct 2026, 07:30 CEST) · Vals.ai−11.5SRE BenchSRE Bench: Opus 5.5 33.59% (read 4 Oct 2026, 07:30 CEST), Fable 5.1 22.90% (read 4 Oct 2026, 07:30 CEST) · Vals.ai−10.7Terminal-Bench 4.0Terminal-Bench 4.0: Opus 5.5 65.15% (read 4 Oct 2026, 07:30 CEST), Fable 5.1 58.08% (read 4 Oct 2026, 07:30 CEST) · Vals.ai−7.1
0%20%40%60%80%100%share passedFable minus OpusCyberBench v1.1CyberBench v1.1: Opus 5.5 55.36% (read 4 Oct 2026, 07:30 CEST), Fable 5.1 70.42% (read 4 Oct 2026, 07:30 CEST) · Vals.ai+15.1Harvey legal agentHarvey legal agent: Opus 5.5 3.75% (read 4 Oct 2026, 07:30 CEST), Fable 5.1 6.67% (read 4 Oct 2026, 07:30 CEST) · Vals.ai+2.9Public Benefits BenchPublic Benefits Bench: Opus 5.5 70.64% (read 4 Oct 2026, 07:30 CEST), Fable 5.1 74.90% (read 4 Oct 2026, 07:30 CEST) · Vals.ai+4.3Tax Agent BenchTax Agent Bench: Opus 5.5 70.50% (read 4 Oct 2026, 07:30 CEST), Fable 5.1 77.64% (read 4 Oct 2026, 07:30 CEST) · Vals.ai+7.1Code MigrationCode Migration: Opus 5.5 66.65% (read 4 Oct 2026, 07:30 CEST), Fable 5.1 54.61% (read 4 Oct 2026, 07:30 CEST) · Vals.ai−12.0ProgramBenchProgramBench: Opus 5.5 18.50% (read 4 Oct 2026, 07:30 CEST), Fable 5.1 7.00% (read 4 Oct 2026, 07:30 CEST) · Vals.ai−11.5SRE BenchSRE Bench: Opus 5.5 33.59% (read 4 Oct 2026, 07:30 CEST), Fable 5.1 22.90% (read 4 Oct 2026, 07:30 CEST) · Vals.ai−10.7Terminal-Bench 4.0Terminal-Bench 4.0: Opus 5.5 65.15% (read 4 Oct 2026, 07:30 CEST), Fable 5.1 58.08% (read 4 Oct 2026, 07:30 CEST) · Vals.ai−7.1
Vals.ai model pages for both models, read 4 October 2026; both give the compute effort as max. A lead counts as clear when it is wider than the error margins shown, added. Highlighted: the two that stay clear with the leader's fallback share taken off its score. The choice of benchmarks and the test are ours.
The numbers in this chart
Opus 5.5Fable 5.1Fable minus Opus
CyberBench v1.155.36%70.42%+15.1 points
Harvey legal agent3.75%6.67%+2.9 points
Public Benefits Bench70.64%74.90%+4.3 points
Tax Agent Bench70.50%77.64%+7.1 points
Code Migration66.65%54.61%−12.0 points
ProgramBench18.50%7.00%−11.5 points
SRE Bench33.59%22.90%−10.7 points
Terminal-Bench 4.065.15%58.08%−7.1 points

Fable leads clearly on CyberBench, where an agent has to find security bugs and patch them, by 15.1 points; on Tax Agent Bench, research-grade US corporate tax questions, by 7.1 points; on Public Benefits Bench, which is about helping people with SNAP benefits, by 4.3 points; and on Harvey's Legal Agent Benchmark, legal work with documents, spreadsheets and file tools, by 2.9 points, though both models pass very little of it.

Opus leads clearly on Code Migration, reimplementing working programs in another language, by 12.0 points; on ProgramBench, which asks a model to rebuild programs from scratch, by 11.5 points; on SRE Bench, working out what a binary does without its source code, by 10.7 points; and on Terminal-Bench 4.0, a set of hard terminal tasks, by 7.1 points.

The fallback models

Both Vals.ai pages say the models ran with Claude Opus 5 and Claude Opus 4.8 as server-side fallbacks for refusals, and the published scores count a task the fallback model passed as passed. Across the Vals Index the pages give a fallback rate of 3.99% for Opus and 2.10% for Fable, but on some benchmarks the share is far higher, and each row's tooltip says how high.

Three of the eight clear leads sit on benchmarks where most or half of the leader's tasks were fallback-assisted: 81.5% of Opus's tasks on Code Migration, 82.8% on SRE Bench, and on CyberBench 53.5% of Fable's and 52.6% of Opus's. So for each clear lead we took the leader's fallback share off its score, the most those answers could have added, and asked whether the lead was still wider than the margins. That is a strict test: a lead that fails it may still hold. Only 2 pass it: Fable's on Tax Agent Bench and Public Benefits Bench, where its fallback share is 0.0% and 0.4%. Of the other 6, the pages give a score with the fallbacks counted as failures for three. They give such scores for three benchmarks inside the margins too.

On SRE Bench, counting the fallback-assisted tasks as failures takes Opus from 33.59% to 5.34% and Fable from 22.90% to 10.69%, so Fable comes out ahead by 5.3 points: Opus's published lead is mostly the fallback models' work. On Harvey's Legal Agent Benchmark, Opus's tooltip lists no fallbacks, and Fable's score falls from 6.67% to 5.83%, which leaves it 2.1 points ahead, inside the margins. On Terminal-Bench 4.0, Opus's lead holds: counted the same way, Opus scores 58.08% and Fable 50.00%. On Vibe Code Bench v1.1, level on the published scores, Opus falls to 83.34% while Fable shows no fallbacks, which puts Fable 6.9 points ahead. The other two, MysteryMechanism and Legal Research Bench, stay inside the margins counted either way.

Where Vals.ai gives a score with fallbacks counted as failures

Published score, and the score with fallback-assisted tasks counted as failures, as Vals.ai's pages give them

benchmarkOpus, publishedOpus, fallbacks as failuresFable, publishedFable, fallbacks as failures
SRE Bench33.59%5.34%22.90%10.69%
Harvey's Legal Agent Benchmark3.75%the same (no fallbacks)6.67%5.83%
Terminal-Bench 4.065.15%58.08%58.08%50.00%
Vibe Code Bench v1.190.29%83.34%90.26%the same (no fallbacks)
MysteryMechanism49.55%49.10%47.75%the same (no fallbacks)
Legal Research Bench50.48%not given55.29%54.33%
From the two Vals.ai model pages' launch notes and the Terminal-Bench 4.0 page, wherever the score the figure starts from matches today's. "No fallbacks": the page's tooltip lists none for that model there. "Not given": the pages don't give that model's score without its fallbacks. The notes give a few more figures from scores that have since changed, which we leave out.

Of the rest, the gap is inside the margins on 12 benchmarks, the Vals Index among them. On the other 4 the page gives no margin to test against: on ProofBench both score 100.00%, and on Vals RSI Index and CUA-bench Opus is about a point ahead. On Time Horizon Index: KSP, Opus scores 91.33% against Fable's 63.33%, the largest gap on the page either way, but with both margins shown as zero there is nothing to test it against.

Every benchmark both models have on Vals.ai

Vals.ai gives both models' compute effort as max. Scores in per cent as published, which count a task a fallback model passed as passed; gap and margins in percentage points; the share of tasks a fallback model helped with, from each row's tooltip

benchmarkOpus 5.5Fable 5.1Fable minus Opusmargins addedfallback-assisted, Opus / Fablewho leads
Public Benefits Bench v1.170.64%74.90%+4.32.31.3% / 0.4%Fable, clear either way
Tax Agent Bench70.50%77.64%+7.16.00.5% / 0.0%Fable, clear either way
Code Migration66.65%54.61%−12.09.181.5% / 1.5%Opus, clear; could rest on fallbacks
CyberBench v1.155.36%70.42%+15.110.652.6% / 53.5%Fable, clear; could rest on fallbacks
Harvey's Legal Agent Benchmark3.75%6.67%+2.92.60.0% / 1.7%Fable, clear; inside the margins with fallbacks as failures
ProgramBench18.50%7.00%−11.54.612.0% / 16.5%Opus, clear; could rest on fallbacks
SRE Bench33.59%22.90%−10.75.582.8% / 75.2%Opus, clear; Fable ahead with fallbacks as failures
Terminal-Bench 4.065.15%58.08%−7.13.311.1% / 11.1%Opus, clear; holds on the page's fallback figure
EMB75.94%76.67%+0.74.50.0% / 0.0%inside the margins
Finance Agent (v2)58.59%58.88%+0.32.21.3% / 0.0%inside the margins
IOI95.06%90.78%−4.39.65.6% / 22.2%inside the margins
Legal Research Bench50.48%55.29%+4.86.92.4% / 1.4%inside the margins
MedCode49.80%53.51%+3.74.40.0% / 0.0%inside the margins
MedScribe91.43%91.29%−0.13.90.0% / 0.0%inside the margins
MysteryMechanism49.55%47.75%−1.86.70.5% / 0.0%inside the margins
SAGE45.83%48.53%+2.76.70.0% / 0.0%inside the margins
Terminal-Bench Science47.14%40.00%−7.111.92.9% / 5.7%inside the margins
Vals Index66.97%65.83%−1.12.04.0% / 2.1%inside the margins
Vibe Code Bench 1-10030.36%28.00%−2.49.316.0% / 28.0%inside the margins
Vibe Code Bench v1.190.29%90.26%0.03.18.0% / 0.0%inside the margins; Fable ahead with fallbacks as failures
CUA-bench14.00%13.17%−0.8–0.0% / 0.0%no margin given
ProofBench v1.1100.00%100.00%0.00.00.0% / 0.0%no margin given
Time Horizon Index: KSP91.33%63.33%−28.00.00.0% / 0.0%no margin given
Vals RSI Index37.31%36.09%−1.2–0.0% / 0.0%no margin given
Vals.ai's scores, error margins and fallback shares; the gap, the sum of the margins and the verdict are ours. "Clear either way": the lead stays wider than the margins with the leader's fallback share taken off its score, the most the fallback answers could have added. "Could rest on fallbacks": it doesn't pass that test, and the pages give no score without the fallbacks. A dash means the page shows no margin for either model; Opus's Terminal-Bench 4.0 margin is shown as zero. Marked: the leads that are clear either way.

Which to pick

On Artificial Analysis's figures, Opus 5.5 is the cheaper and higher-scoring choice for most work from medium up, and at max Fable costs 1.3× as much per task, the smallest premium of any setting. Fable 5.1 leads at low, on Artificial Analysis's legal agent test at every setting, and on Vals.ai's tax and benefits benchmarks, the two leads there that hold however the fallback answers are counted. On Vals.ai's hard terminal tasks Opus's lead holds on the page's own figure. Fable's CyberBench lead, and Opus's leads on rebuilding and porting programs, may hold too, but on those three the leader's fallback share is larger than its lead, and the pages don't give a score without the fallbacks.

What this doesn't tell you

Every figure here is a benchmark run by someone else, on tasks that evaluator chose. The Terminal-Bench 4.0 page says its scores are the mean of three runs; the other pages we saved don't say how many runs are behind each score. A lead of a few points on a tax benchmark says how the models did on those tasks, not on your returns.

Artificial Analysis names its rows for both models "Default Fallback", but its page doesn't say what that is or how many tasks it affected, so its figures may include fallback answers of the kind Vals.ai reports.

The costs are Artificial Analysis's cost per task in dollars at list prices. Neither evaluator's page says anything about use on a subscription.

Vals.ai's pages give one setting, max, for both models, so the Vals comparison says nothing about the lower settings.

Both evaluators update their figures, and every figure is dated in the numbers table below.

Every number

These are all 612 figures behind this article, grouped by whose they are, with the page each came from and when we read it. Figures marked ⟳ can move. When a re-read finds a change, the new value shows next to the one we first published.

Artificial Analysis

Claude Opus 5.5, list price per million input tokens$4
Claude Opus 5.5, list price per million output tokens$20
Artificial Analysis Intelligence Index, Claude Opus 5.5 at low42.3 ⟳
Cost per task on the Intelligence Index (USD), Claude Opus 5.5 at low$0.55 ⟳
Output tokens per task on the Intelligence Index, Claude Opus 5.5 at low10K ⟳
AA-Briefcase score (Elo), Claude Opus 5.5 at low1280 ⟳
GDPval-AA score (Elo), Claude Opus 5.5 at low1236 ⟳
AutomationBench-AA score, Claude Opus 5.5 at low52.9% ⟳
Terminal-Bench 4.0 score, Claude Opus 5.5 at low31.3% ⟳
SciCode score, Claude Opus 5.5 at low58.6% ⟳
Humanity's Last Exam score, Claude Opus 5.5 at low48.3% ⟳
GDP.pdf score, Claude Opus 5.5 at low25.6% ⟳
CritPt score, Claude Opus 5.5 at low17.7% ⟳
AA-Omniscience score, Claude Opus 5.5 at low38.9 ⟳
AA-LCR score, Claude Opus 5.5 at low80.7% ⟳
Harvey LAB-AA score, Claude Opus 5.5 at low89.1% ⟳
Terminal-Bench Science score, Claude Opus 5.5 at low24.3% ⟳
Artificial Analysis Economics score, Claude Opus 5.5 at low52.5 ⟳
Artificial Analysis Engineering score, Claude Opus 5.5 at low44.6 ⟳
Artificial Analysis Finance and accounting score, Claude Opus 5.5 at low45.9 ⟳
Artificial Analysis Strategy and ops score, Claude Opus 5.5 at low48.5 ⟳
Artificial Analysis Intelligence Index, Claude Opus 5.5 at medium51.2 ⟳
Cost per task on the Intelligence Index (USD), Claude Opus 5.5 at medium$1.34 ⟳
Output tokens per task on the Intelligence Index, Claude Opus 5.5 at medium26K ⟳
AA-Briefcase score (Elo), Claude Opus 5.5 at medium1628 ⟳
GDPval-AA score (Elo), Claude Opus 5.5 at medium1586 ⟳
AutomationBench-AA score, Claude Opus 5.5 at medium61.2% ⟳
Terminal-Bench 4.0 score, Claude Opus 5.5 at medium52.5% ⟳
SciCode score, Claude Opus 5.5 at medium59.3% ⟳
Humanity's Last Exam score, Claude Opus 5.5 at medium54.7% ⟳
GDP.pdf score, Claude Opus 5.5 at medium25.6% ⟳
CritPt score, Claude Opus 5.5 at medium27.7% ⟳
AA-Omniscience score, Claude Opus 5.5 at medium40.3 ⟳
AA-LCR score, Claude Opus 5.5 at medium84.3% ⟳
Harvey LAB-AA score, Claude Opus 5.5 at medium90.3% ⟳
Terminal-Bench Science score, Claude Opus 5.5 at medium43.3% ⟳
Artificial Analysis Economics score, Claude Opus 5.5 at medium59.7 ⟳
Artificial Analysis Engineering score, Claude Opus 5.5 at medium53.7 ⟳
Artificial Analysis Finance and accounting score, Claude Opus 5.5 at medium54.1 ⟳
Artificial Analysis Strategy and ops score, Claude Opus 5.5 at medium57.5 ⟳
Artificial Analysis Intelligence Index, Claude Opus 5.5 at high53.6 ⟳
Cost per task on the Intelligence Index (USD), Claude Opus 5.5 at high$1.82 ⟳
Output tokens per task on the Intelligence Index, Claude Opus 5.5 at high36K ⟳
AA-Briefcase score (Elo), Claude Opus 5.5 at high1689 ⟳
GDPval-AA score (Elo), Claude Opus 5.5 at high1707 ⟳
AutomationBench-AA score, Claude Opus 5.5 at high63.2% ⟳
Terminal-Bench 4.0 score, Claude Opus 5.5 at high56.6% ⟳
SciCode score, Claude Opus 5.5 at high60.4% ⟳
Humanity's Last Exam score, Claude Opus 5.5 at high55.6% ⟳
GDP.pdf score, Claude Opus 5.5 at high28.8% ⟳
CritPt score, Claude Opus 5.5 at high30.9% ⟳
AA-Omniscience score, Claude Opus 5.5 at high40.6 ⟳
AA-LCR score, Claude Opus 5.5 at high82.7% ⟳
Harvey LAB-AA score, Claude Opus 5.5 at high90.9% ⟳
Terminal-Bench Science score, Claude Opus 5.5 at high49.0% ⟳
Artificial Analysis Economics score, Claude Opus 5.5 at high60.6 ⟳
Artificial Analysis Engineering score, Claude Opus 5.5 at high56.2 ⟳
Artificial Analysis Finance and accounting score, Claude Opus 5.5 at high55.9 ⟳
Artificial Analysis Strategy and ops score, Claude Opus 5.5 at high59.2 ⟳
Artificial Analysis Intelligence Index, Claude Opus 5.5 at xhigh56.0 ⟳
Cost per task on the Intelligence Index (USD), Claude Opus 5.5 at xhigh$3.46 ⟳
Output tokens per task on the Intelligence Index, Claude Opus 5.5 at xhigh66K ⟳
AA-Briefcase score (Elo), Claude Opus 5.5 at xhigh1768 ⟳
GDPval-AA score (Elo), Claude Opus 5.5 at xhigh1837 ⟳
AutomationBench-AA score, Claude Opus 5.5 at xhigh65.0% ⟳
Terminal-Bench 4.0 score, Claude Opus 5.5 at xhigh59.6% ⟳
SciCode score, Claude Opus 5.5 at xhigh65.0% ⟳
Humanity's Last Exam score, Claude Opus 5.5 at xhigh57.5% ⟳
GDP.pdf score, Claude Opus 5.5 at xhigh26.6% ⟳
CritPt score, Claude Opus 5.5 at xhigh31.7% ⟳
AA-Omniscience score, Claude Opus 5.5 at xhigh42.6 ⟳
AA-LCR score, Claude Opus 5.5 at xhigh84.7% ⟳
Harvey LAB-AA score, Claude Opus 5.5 at xhigh91.2% ⟳
Terminal-Bench Science score, Claude Opus 5.5 at xhigh61.9% ⟳
Artificial Analysis Economics score, Claude Opus 5.5 at xhigh63.0 ⟳
Artificial Analysis Engineering score, Claude Opus 5.5 at xhigh58.7 ⟳
Artificial Analysis Finance and accounting score, Claude Opus 5.5 at xhigh58.4 ⟳
Artificial Analysis Strategy and ops score, Claude Opus 5.5 at xhigh61.6 ⟳
Artificial Analysis Intelligence Index, Claude Opus 5.5 at max57.6 ⟳
Cost per task on the Intelligence Index (USD), Claude Opus 5.5 at max$5.98 ⟳
Output tokens per task on the Intelligence Index, Claude Opus 5.5 at max119K ⟳
AA-Briefcase score (Elo), Claude Opus 5.5 at max1808 ⟳
GDPval-AA score (Elo), Claude Opus 5.5 at max1867 ⟳
AutomationBench-AA score, Claude Opus 5.5 at max69.5% ⟳
Terminal-Bench 4.0 score, Claude Opus 5.5 at max59.6% ⟳
SciCode score, Claude Opus 5.5 at max66.9% ⟳
Humanity's Last Exam score, Claude Opus 5.5 at max61.4% ⟳
GDP.pdf score, Claude Opus 5.5 at max26.2% ⟳
CritPt score, Claude Opus 5.5 at max31.7% ⟳
AA-Omniscience score, Claude Opus 5.5 at max46.4 ⟳
AA-LCR score, Claude Opus 5.5 at max84.7% ⟳
Harvey LAB-AA score, Claude Opus 5.5 at max91.2% ⟳
Terminal-Bench Science score, Claude Opus 5.5 at max59.0% ⟳
ITBench-AA score, Claude Opus 5.5 at max38.2% ⟳
AA-AnalystAgent score, Claude Opus 5.5 at max56.2% ⟳
MLCR-AA score, Claude Opus 5.5 at max66.7% ⟳
Artificial Analysis Economics score, Claude Opus 5.5 at max65.6 ⟳
Artificial Analysis Engineering score, Claude Opus 5.5 at max60.4 ⟳
Artificial Analysis Finance and accounting score, Claude Opus 5.5 at max60.7 ⟳
Artificial Analysis Strategy and ops score, Claude Opus 5.5 at max63.7 ⟳
Artificial Analysis Healthcare and medical score, Claude Opus 5.5 at max60.5 ⟳
Claude Fable 5.1, list price per million input tokens$10
Claude Fable 5.1, list price per million output tokens$50
Artificial Analysis Intelligence Index, Claude Fable 5.1 at low46.8 ⟳
Cost per task on the Intelligence Index (USD), Claude Fable 5.1 at low$2.37 ⟳
Output tokens per task on the Intelligence Index, Claude Fable 5.1 at low22K ⟳
AA-Briefcase score (Elo), Claude Fable 5.1 at low1482 ⟳
GDPval-AA score (Elo), Claude Fable 5.1 at low1469 ⟳
AutomationBench-AA score, Claude Fable 5.1 at low52.2% ⟳
Terminal-Bench 4.0 score, Claude Fable 5.1 at low40.4% ⟳
SciCode score, Claude Fable 5.1 at low56.7% ⟳
Humanity's Last Exam score, Claude Fable 5.1 at low48.9% ⟳
GDP.pdf score, Claude Fable 5.1 at low28.0% ⟳
CritPt score, Claude Fable 5.1 at low27.7% ⟳
AA-Omniscience score, Claude Fable 5.1 at low34.1 ⟳
AA-LCR score, Claude Fable 5.1 at low82.3% ⟳
Harvey LAB-AA score, Claude Fable 5.1 at low92.3% ⟳
Artificial Analysis Economics score, Claude Fable 5.1 at low55.1 ⟳
Artificial Analysis Engineering score, Claude Fable 5.1 at low48.7 ⟳
Artificial Analysis Finance and accounting score, Claude Fable 5.1 at low48.7 ⟳
Artificial Analysis Strategy and ops score, Claude Fable 5.1 at low51.2 ⟳
Artificial Analysis Intelligence Index, Claude Fable 5.1 at medium48.9 ⟳
Cost per task on the Intelligence Index (USD), Claude Fable 5.1 at medium$2.98 ⟳
Output tokens per task on the Intelligence Index, Claude Fable 5.1 at medium28K ⟳
AA-Briefcase score (Elo), Claude Fable 5.1 at medium1529 ⟳
GDPval-AA score (Elo), Claude Fable 5.1 at medium1550 ⟳
AutomationBench-AA score, Claude Fable 5.1 at medium54.7% ⟳
Terminal-Bench 4.0 score, Claude Fable 5.1 at medium44.9% ⟳
SciCode score, Claude Fable 5.1 at medium56.4% ⟳
Humanity's Last Exam score, Claude Fable 5.1 at medium53.8% ⟳
GDP.pdf score, Claude Fable 5.1 at medium26.8% ⟳
CritPt score, Claude Fable 5.1 at medium29.1% ⟳
AA-Omniscience score, Claude Fable 5.1 at medium37.6 ⟳
AA-LCR score, Claude Fable 5.1 at medium84.7% ⟳
Harvey LAB-AA score, Claude Fable 5.1 at medium92.6% ⟳
Artificial Analysis Economics score, Claude Fable 5.1 at medium57.8 ⟳
Artificial Analysis Engineering score, Claude Fable 5.1 at medium51.5 ⟳
Artificial Analysis Finance and accounting score, Claude Fable 5.1 at medium51.3 ⟳
Artificial Analysis Strategy and ops score, Claude Fable 5.1 at medium53.9 ⟳
Artificial Analysis Intelligence Index, Claude Fable 5.1 at high51.2 ⟳
Cost per task on the Intelligence Index (USD), Claude Fable 5.1 at high$3.91 ⟳
Output tokens per task on the Intelligence Index, Claude Fable 5.1 at high38K ⟳
AA-Briefcase score (Elo), Claude Fable 5.1 at high1581 ⟳
GDPval-AA score (Elo), Claude Fable 5.1 at high1635 ⟳
AutomationBench-AA score, Claude Fable 5.1 at high55.3% ⟳
Terminal-Bench 4.0 score, Claude Fable 5.1 at high52.0% ⟳
SciCode score, Claude Fable 5.1 at high58.7% ⟳
Humanity's Last Exam score, Claude Fable 5.1 at high55.9% ⟳
GDP.pdf score, Claude Fable 5.1 at high26.8% ⟳
CritPt score, Claude Fable 5.1 at high30.3% ⟳
AA-Omniscience score, Claude Fable 5.1 at high40.8 ⟳
AA-LCR score, Claude Fable 5.1 at high83.7% ⟳
Harvey LAB-AA score, Claude Fable 5.1 at high93.0% ⟳
Artificial Analysis Economics score, Claude Fable 5.1 at high60.2 ⟳
Artificial Analysis Engineering score, Claude Fable 5.1 at high54.5 ⟳
Artificial Analysis Finance and accounting score, Claude Fable 5.1 at high53.8 ⟳
Artificial Analysis Strategy and ops score, Claude Fable 5.1 at high56.1 ⟳
Artificial Analysis Intelligence Index, Claude Fable 5.1 at xhigh53.2 ⟳
Cost per task on the Intelligence Index (USD), Claude Fable 5.1 at xhigh$5.98 ⟳
Output tokens per task on the Intelligence Index, Claude Fable 5.1 at xhigh61K ⟳
AA-Briefcase score (Elo), Claude Fable 5.1 at xhigh1657 ⟳
GDPval-AA score (Elo), Claude Fable 5.1 at xhigh1735 ⟳
AutomationBench-AA score, Claude Fable 5.1 at xhigh57.8% ⟳
Terminal-Bench 4.0 score, Claude Fable 5.1 at xhigh55.1% ⟳
SciCode score, Claude Fable 5.1 at xhigh60.9% ⟳
Humanity's Last Exam score, Claude Fable 5.1 at xhigh58.7% ⟳
GDP.pdf score, Claude Fable 5.1 at xhigh26.2% ⟳
CritPt score, Claude Fable 5.1 at xhigh31.1% ⟳
AA-Omniscience score, Claude Fable 5.1 at xhigh42.4 ⟳
AA-LCR score, Claude Fable 5.1 at xhigh83.0% ⟳
Harvey LAB-AA score, Claude Fable 5.1 at xhigh93.3% ⟳
Artificial Analysis Economics score, Claude Fable 5.1 at xhigh62.4 ⟳
Artificial Analysis Engineering score, Claude Fable 5.1 at xhigh56.8 ⟳
Artificial Analysis Finance and accounting score, Claude Fable 5.1 at xhigh56.2 ⟳
Artificial Analysis Strategy and ops score, Claude Fable 5.1 at xhigh58.4 ⟳
Artificial Analysis Intelligence Index, Claude Fable 5.1 at max53.4 ⟳
Cost per task on the Intelligence Index (USD), Claude Fable 5.1 at max$7.63 ⟳
Output tokens per task on the Intelligence Index, Claude Fable 5.1 at max78K ⟳
AA-Briefcase score (Elo), Claude Fable 5.1 at max1676 ⟳
GDPval-AA score (Elo), Claude Fable 5.1 at max1758 ⟳
AutomationBench-AA score, Claude Fable 5.1 at max59.4% ⟳
Terminal-Bench 4.0 score, Claude Fable 5.1 at max52.0% ⟳
SciCode score, Claude Fable 5.1 at max63.1% ⟳
Humanity's Last Exam score, Claude Fable 5.1 at max59.1% ⟳
GDP.pdf score, Claude Fable 5.1 at max26.2% ⟳
CritPt score, Claude Fable 5.1 at max29.7% ⟳
AA-Omniscience score, Claude Fable 5.1 at max43.5 ⟳
AA-LCR score, Claude Fable 5.1 at max85.3% ⟳
Harvey LAB-AA score, Claude Fable 5.1 at max93.0% ⟳
Terminal-Bench Science score, Claude Fable 5.1 at max43.3% ⟳
ITBench-AA score, Claude Fable 5.1 at max49.5% ⟳
AA-AnalystAgent score, Claude Fable 5.1 at max57.5% ⟳
MLCR-AA score, Claude Fable 5.1 at max71.1% ⟳
Artificial Analysis Economics score, Claude Fable 5.1 at max62.7 ⟳
Artificial Analysis Engineering score, Claude Fable 5.1 at max56.5 ⟳
Artificial Analysis Finance and accounting score, Claude Fable 5.1 at max56.4 ⟳
Artificial Analysis Strategy and ops score, Claude Fable 5.1 at max59.7 ⟳
Artificial Analysis Healthcare and medical score, Claude Fable 5.1 at max58.0 ⟳
Intelligence Index, Claude Opus 5.5 at low (for the chart)42.3
Intelligence Index, Claude Fable 5.1 at low (for the chart)46.8
Intelligence Index, Claude Opus 5.5 at medium (for the chart)51.2
Intelligence Index, Claude Fable 5.1 at medium (for the chart)48.9
Intelligence Index, Claude Opus 5.5 at high (for the chart)53.6
Intelligence Index, Claude Fable 5.1 at high (for the chart)51.2
Intelligence Index, Claude Opus 5.5 at xhigh (for the chart)56.0
Intelligence Index, Claude Fable 5.1 at xhigh (for the chart)53.2
Intelligence Index, Claude Opus 5.5 at max (for the chart)57.6
Intelligence Index, Claude Fable 5.1 at max (for the chart)53.4

Model Fatigue, from Artificial Analysis's and Vals.ai's figures

Our comparison of the two models (derived/compare.py) · from the read of 4 Oct 2026, 07:30 CEST
Per-test scores where Fable 5.1 is ahead of Opus 5.5, both at low (our count)8
Per-test scores both models have at low on the page (our count)11
Per-test scores where Opus 5.5 is ahead of Fable 5.1, both at low (our count)3
Domain scores where Fable 5.1 is ahead of Opus 5.5, both at low (our count)5
Domain scores both models have at low (our count)5
Per-test scores where Fable 5.1 is ahead of Opus 5.5, both at medium (our count)4
Per-test scores both models have at medium on the page (our count)11
Per-test scores where Opus 5.5 is ahead of Fable 5.1, both at medium (our count)7
Domain scores where Fable 5.1 is ahead of Opus 5.5, both at medium (our count)0
Domain scores both models have at medium (our count)5
Per-test scores where Fable 5.1 is ahead of Opus 5.5, both at high (our count)4
Per-test scores both models have at high on the page (our count)11
Per-test scores where Opus 5.5 is ahead of Fable 5.1, both at high (our count)7
Domain scores where Fable 5.1 is ahead of Opus 5.5, both at high (our count)0
Domain scores both models have at high (our count)5
Per-test scores where Fable 5.1 is ahead of Opus 5.5, both at xhigh (our count)2
Per-test scores both models have at xhigh on the page (our count)11
Per-test scores where Opus 5.5 is ahead of Fable 5.1, both at xhigh (our count)9
Domain scores where Fable 5.1 is ahead of Opus 5.5, both at xhigh (our count)1
Domain scores both models have at xhigh (our count)5
Per-test scores where Fable 5.1 is ahead of Opus 5.5, both at max (our count)5
Per-test scores both models have at max on the page (our count)15
Per-test scores where Opus 5.5 is ahead of Fable 5.1, both at max (our count)9
Domain scores where Fable 5.1 is ahead of Opus 5.5, both at max (our count)0
Domain scores both models have at max (our count)6
Vals.ai benchmarks both models have where Fable 5.1 is ahead by more than the two margins added (our count)4
Vals.ai benchmarks both models have where Opus 5.5 is ahead by more than the two margins added (our count)4
Vals.ai benchmarks both models have where the gap is inside the two margins added (our count)12
Vals.ai benchmarks both models have where the page gives no margin, or both as ± 0.00 (our count)4
Vals.ai benchmarks both models have a result on (our count)24
Vals.ai Code Migration: the leader's score minus its fallback share, minus the other model's published score (the lead with the most the fallback answers could have added to it taken off; our arithmetic)−69.5 points
Vals.ai CyberBench v1.1: the leader's score minus its fallback share, minus the other model's published score (the lead with the most the fallback answers could have added to it taken off; our arithmetic)−38.4 points
Vals.ai ProgramBench: the leader's score minus its fallback share, minus the other model's published score (the lead with the most the fallback answers could have added to it taken off; our arithmetic)−0.5 points
Vals.ai Public Benefits Bench v1.1: the leader's score minus its fallback share, minus the other model's published score (the lead with the most the fallback answers could have added to it taken off; our arithmetic)+3.8 points
Vals.ai SRE Bench: the leader's score minus its fallback share, minus the other model's published score (the lead with the most the fallback answers could have added to it taken off; our arithmetic)−72.1 points
Vals.ai Tax Agent Bench: the leader's score minus its fallback share, minus the other model's published score (the lead with the most the fallback answers could have added to it taken off; our arithmetic)+7.1 points
Vals.ai Terminal-Bench 4.0: the leader's score minus its fallback share, minus the other model's published score (the lead with the most the fallback answers could have added to it taken off; our arithmetic)−4.0 points
Clear leads on Vals.ai that stay clear with the leader's fallback share taken off its score (our count)2
Clear leads on Vals.ai that aren't shown to stay clear with the leader's fallback share taken off its score (our count)6
Fable 5.1's cost per task as a multiple of Opus 5.5's, both at low (our arithmetic)4.3×
Intelligence Index, Fable 5.1 minus Opus 5.5, both at low (our arithmetic)+4.5 points
Fable 5.1's output tokens per task as a multiple of Opus 5.5's, both at low (our arithmetic)2.1×
AA-Briefcase, Fable 5.1 minus Opus 5.5 in Elo points, both at low (our arithmetic)+202
GDPval-AA, Fable 5.1 minus Opus 5.5 in Elo points, both at low (our arithmetic)+233
AutomationBench-AA, Fable 5.1 minus Opus 5.5 in percentage points, both at low (our arithmetic)−0.7 points
Terminal-Bench 4.0, Fable 5.1 minus Opus 5.5 in percentage points, both at low (our arithmetic)+9.1 points
SciCode, Fable 5.1 minus Opus 5.5 in percentage points, both at low (our arithmetic)−1.9 points
Humanity's Last Exam, Fable 5.1 minus Opus 5.5 in percentage points, both at low (our arithmetic)+0.6 points
GDP.pdf, Fable 5.1 minus Opus 5.5 in percentage points, both at low (our arithmetic)+2.4 points
CritPt, Fable 5.1 minus Opus 5.5 in percentage points, both at low (our arithmetic)+10.0 points
AA-Omniscience, Fable 5.1 minus Opus 5.5 in index points, both at low (our arithmetic)−4.7
AA-LCR, Fable 5.1 minus Opus 5.5 in percentage points, both at low (our arithmetic)+1.7 points
Harvey LAB-AA, Fable 5.1 minus Opus 5.5 in percentage points, both at low (our arithmetic)+3.2 points
Economics score, Fable 5.1 minus Opus 5.5, both at low (our arithmetic)+2.6
Engineering score, Fable 5.1 minus Opus 5.5, both at low (our arithmetic)+4.1
Finance and accounting score, Fable 5.1 minus Opus 5.5, both at low (our arithmetic)+2.8
Strategy and ops score, Fable 5.1 minus Opus 5.5, both at low (our arithmetic)+2.7
Fable 5.1's cost per task as a multiple of Opus 5.5's, both at medium (our arithmetic)2.2×
Intelligence Index, Fable 5.1 minus Opus 5.5, both at medium (our arithmetic)−2.3 points
Fable 5.1's output tokens per task as a multiple of Opus 5.5's, both at medium (our arithmetic)1.1×
AA-Briefcase, Fable 5.1 minus Opus 5.5 in Elo points, both at medium (our arithmetic)−98
GDPval-AA, Fable 5.1 minus Opus 5.5 in Elo points, both at medium (our arithmetic)−36
AutomationBench-AA, Fable 5.1 minus Opus 5.5 in percentage points, both at medium (our arithmetic)−6.6 points
Terminal-Bench 4.0, Fable 5.1 minus Opus 5.5 in percentage points, both at medium (our arithmetic)−7.6 points
SciCode, Fable 5.1 minus Opus 5.5 in percentage points, both at medium (our arithmetic)−2.9 points
Humanity's Last Exam, Fable 5.1 minus Opus 5.5 in percentage points, both at medium (our arithmetic)−0.9 points
GDP.pdf, Fable 5.1 minus Opus 5.5 in percentage points, both at medium (our arithmetic)+1.2 points
CritPt, Fable 5.1 minus Opus 5.5 in percentage points, both at medium (our arithmetic)+1.4 points
AA-Omniscience, Fable 5.1 minus Opus 5.5 in index points, both at medium (our arithmetic)−2.7
AA-LCR, Fable 5.1 minus Opus 5.5 in percentage points, both at medium (our arithmetic)+0.3 points
Harvey LAB-AA, Fable 5.1 minus Opus 5.5 in percentage points, both at medium (our arithmetic)+2.4 points
Economics score, Fable 5.1 minus Opus 5.5, both at medium (our arithmetic)−1.9
Engineering score, Fable 5.1 minus Opus 5.5, both at medium (our arithmetic)−2.3
Finance and accounting score, Fable 5.1 minus Opus 5.5, both at medium (our arithmetic)−2.8
Strategy and ops score, Fable 5.1 minus Opus 5.5, both at medium (our arithmetic)−3.6
Fable 5.1's cost per task as a multiple of Opus 5.5's, both at high (our arithmetic)2.1×
Intelligence Index, Fable 5.1 minus Opus 5.5, both at high (our arithmetic)−2.4 points
Fable 5.1's output tokens per task as a multiple of Opus 5.5's, both at high (our arithmetic)1.1×
AA-Briefcase, Fable 5.1 minus Opus 5.5 in Elo points, both at high (our arithmetic)−108
GDPval-AA, Fable 5.1 minus Opus 5.5 in Elo points, both at high (our arithmetic)−72
AutomationBench-AA, Fable 5.1 minus Opus 5.5 in percentage points, both at high (our arithmetic)−7.9 points
Terminal-Bench 4.0, Fable 5.1 minus Opus 5.5 in percentage points, both at high (our arithmetic)−4.5 points
SciCode, Fable 5.1 minus Opus 5.5 in percentage points, both at high (our arithmetic)−1.7 points
Humanity's Last Exam, Fable 5.1 minus Opus 5.5 in percentage points, both at high (our arithmetic)+0.4 points
GDP.pdf, Fable 5.1 minus Opus 5.5 in percentage points, both at high (our arithmetic)−2.0 points
CritPt, Fable 5.1 minus Opus 5.5 in percentage points, both at high (our arithmetic)−0.6 points
AA-Omniscience, Fable 5.1 minus Opus 5.5 in index points, both at high (our arithmetic)+0.2
AA-LCR, Fable 5.1 minus Opus 5.5 in percentage points, both at high (our arithmetic)+1.0 points
Harvey LAB-AA, Fable 5.1 minus Opus 5.5 in percentage points, both at high (our arithmetic)+2.1 points
Economics score, Fable 5.1 minus Opus 5.5, both at high (our arithmetic)−0.4
Engineering score, Fable 5.1 minus Opus 5.5, both at high (our arithmetic)−1.7
Finance and accounting score, Fable 5.1 minus Opus 5.5, both at high (our arithmetic)−2.2
Strategy and ops score, Fable 5.1 minus Opus 5.5, both at high (our arithmetic)−3.1
Fable 5.1's cost per task as a multiple of Opus 5.5's, both at xhigh (our arithmetic)1.7×
Intelligence Index, Fable 5.1 minus Opus 5.5, both at xhigh (our arithmetic)−2.8 points
Fable 5.1's output tokens per task as a multiple of Opus 5.5's, both at xhigh (our arithmetic)0.9×
AA-Briefcase, Fable 5.1 minus Opus 5.5 in Elo points, both at xhigh (our arithmetic)−111
GDPval-AA, Fable 5.1 minus Opus 5.5 in Elo points, both at xhigh (our arithmetic)−102
AutomationBench-AA, Fable 5.1 minus Opus 5.5 in percentage points, both at xhigh (our arithmetic)−7.2 points
Terminal-Bench 4.0, Fable 5.1 minus Opus 5.5 in percentage points, both at xhigh (our arithmetic)−4.5 points
SciCode, Fable 5.1 minus Opus 5.5 in percentage points, both at xhigh (our arithmetic)−4.2 points
Humanity's Last Exam, Fable 5.1 minus Opus 5.5 in percentage points, both at xhigh (our arithmetic)+1.2 points
GDP.pdf, Fable 5.1 minus Opus 5.5 in percentage points, both at xhigh (our arithmetic)−0.4 points
CritPt, Fable 5.1 minus Opus 5.5 in percentage points, both at xhigh (our arithmetic)−0.6 points
AA-Omniscience, Fable 5.1 minus Opus 5.5 in index points, both at xhigh (our arithmetic)−0.3
AA-LCR, Fable 5.1 minus Opus 5.5 in percentage points, both at xhigh (our arithmetic)−1.7 points
Harvey LAB-AA, Fable 5.1 minus Opus 5.5 in percentage points, both at xhigh (our arithmetic)+2.0 points
Economics score, Fable 5.1 minus Opus 5.5, both at xhigh (our arithmetic)−0.6
Engineering score, Fable 5.1 minus Opus 5.5, both at xhigh (our arithmetic)−2.0
Finance and accounting score, Fable 5.1 minus Opus 5.5, both at xhigh (our arithmetic)−2.2
Strategy and ops score, Fable 5.1 minus Opus 5.5, both at xhigh (our arithmetic)−3.2
Fable 5.1's cost per task as a multiple of Opus 5.5's, both at max (our arithmetic)1.3×
Intelligence Index, Fable 5.1 minus Opus 5.5, both at max (our arithmetic)−4.3 points
Fable 5.1's output tokens per task as a multiple of Opus 5.5's, both at max (our arithmetic)0.7×
AA-Briefcase, Fable 5.1 minus Opus 5.5 in Elo points, both at max (our arithmetic)−132
GDPval-AA, Fable 5.1 minus Opus 5.5 in Elo points, both at max (our arithmetic)−109
AutomationBench-AA, Fable 5.1 minus Opus 5.5 in percentage points, both at max (our arithmetic)−10.2 points
Terminal-Bench 4.0, Fable 5.1 minus Opus 5.5 in percentage points, both at max (our arithmetic)−7.6 points
SciCode, Fable 5.1 minus Opus 5.5 in percentage points, both at max (our arithmetic)−3.8 points
Humanity's Last Exam, Fable 5.1 minus Opus 5.5 in percentage points, both at max (our arithmetic)−2.2 points
GDP.pdf, Fable 5.1 minus Opus 5.5 in percentage points, both at max (our arithmetic)0.0 points
CritPt, Fable 5.1 minus Opus 5.5 in percentage points, both at max (our arithmetic)−2.0 points
AA-Omniscience, Fable 5.1 minus Opus 5.5 in index points, both at max (our arithmetic)−3.0
AA-LCR, Fable 5.1 minus Opus 5.5 in percentage points, both at max (our arithmetic)+0.7 points
Harvey LAB-AA, Fable 5.1 minus Opus 5.5 in percentage points, both at max (our arithmetic)+1.8 points
Terminal-Bench Science, Fable 5.1 minus Opus 5.5 in percentage points, both at max (our arithmetic)−15.7 points
ITBench-AA, Fable 5.1 minus Opus 5.5 in percentage points, both at max (our arithmetic)+11.3 points
AA-AnalystAgent, Fable 5.1 minus Opus 5.5 in percentage points, both at max (our arithmetic)+1.2 points
MLCR-AA, Fable 5.1 minus Opus 5.5 in percentage points, both at max (our arithmetic)+4.4 points
Economics score, Fable 5.1 minus Opus 5.5, both at max (our arithmetic)−2.9
Engineering score, Fable 5.1 minus Opus 5.5, both at max (our arithmetic)−3.9
Finance and accounting score, Fable 5.1 minus Opus 5.5, both at max (our arithmetic)−4.3
Strategy and ops score, Fable 5.1 minus Opus 5.5, both at max (our arithmetic)−4.0
Healthcare and medical score, Fable 5.1 minus Opus 5.5, both at max (our arithmetic)−2.5
Fable 5.1's list price per output token as a multiple of Opus 5.5's (input is the same multiple; our arithmetic)2.5×
Vals.ai CUA-bench, Fable 5.1 minus Opus 5.5 in percentage points (our arithmetic)−0.8 points
Vals.ai Code Migration, Fable 5.1 minus Opus 5.5 in percentage points (our arithmetic)−12.0 points
Vals.ai Code Migration, the error margins shown added (our arithmetic)9.1
Vals.ai CyberBench v1.1, Fable 5.1 minus Opus 5.5 in percentage points (our arithmetic)+15.1 points
Vals.ai CyberBench v1.1, the error margins shown added (our arithmetic)10.6
Vals.ai EMB, Fable 5.1 minus Opus 5.5 in percentage points (our arithmetic)+0.7 points
Vals.ai EMB, the error margins shown added (our arithmetic)4.5
Vals.ai Finance Agent (v2), Fable 5.1 minus Opus 5.5 in percentage points (our arithmetic)+0.3 points
Vals.ai Finance Agent (v2), the error margins shown added (our arithmetic)2.2
Vals.ai IOI, Fable 5.1 minus Opus 5.5 in percentage points (our arithmetic)−4.3 points
Vals.ai IOI, the error margins shown added (our arithmetic)9.6
Vals.ai MedCode, Fable 5.1 minus Opus 5.5 in percentage points (our arithmetic)+3.7 points
Vals.ai MedCode, the error margins shown added (our arithmetic)4.4
Vals.ai MedScribe, Fable 5.1 minus Opus 5.5 in percentage points (our arithmetic)−0.1 points
Vals.ai MedScribe, the error margins shown added (our arithmetic)3.9
Vals.ai MysteryMechanism, Fable 5.1 minus Opus 5.5 in percentage points (our arithmetic)−1.8 points
Vals.ai MysteryMechanism, the error margins shown added (our arithmetic)6.7
Vals.ai ProgramBench, Fable 5.1 minus Opus 5.5 in percentage points (our arithmetic)−11.5 points
Vals.ai ProgramBench, the error margins shown added (our arithmetic)4.6
Vals.ai ProofBench v1.1, Fable 5.1 minus Opus 5.5 in percentage points (our arithmetic)0.0 points
Vals.ai ProofBench v1.1, the error margins shown added (our arithmetic)0.0
Vals.ai Public Benefits Bench v1.1, Fable 5.1 minus Opus 5.5 in percentage points (our arithmetic)+4.3 points
Vals.ai Public Benefits Bench v1.1, the error margins shown added (our arithmetic)2.3
Vals.ai SAGE, Fable 5.1 minus Opus 5.5 in percentage points (our arithmetic)+2.7 points
Vals.ai SAGE, the error margins shown added (our arithmetic)6.7
Vals.ai SRE Bench, Fable 5.1 minus Opus 5.5 in percentage points (our arithmetic)−10.7 points
Vals.ai SRE Bench, the error margins shown added (our arithmetic)5.5
Vals.ai Tax Agent Bench, Fable 5.1 minus Opus 5.5 in percentage points (our arithmetic)+7.1 points
Vals.ai Tax Agent Bench, the error margins shown added (our arithmetic)6.0
Vals.ai Terminal-Bench 4.0, Fable 5.1 minus Opus 5.5 in percentage points (our arithmetic)−7.1 points
Vals.ai Terminal-Bench 4.0, the error margins shown added (our arithmetic)3.3
Vals.ai Terminal-Bench Science, Fable 5.1 minus Opus 5.5 in percentage points (our arithmetic)−7.1 points
Vals.ai Terminal-Bench Science, the error margins shown added (our arithmetic)11.9
Vals.ai Time Horizon Index: KSP, Fable 5.1 minus Opus 5.5 in percentage points (our arithmetic)−28.0 points
Vals.ai Time Horizon Index: KSP, the error margins shown added (our arithmetic)0.0
Vals.ai Vals Index, Fable 5.1 minus Opus 5.5 in percentage points (our arithmetic)−1.1 points
Vals.ai Vals Index, the error margins shown added (our arithmetic)2.0
Vals.ai Vals RSI Index, Fable 5.1 minus Opus 5.5 in percentage points (our arithmetic)−1.2 points
Vals.ai Vibe Code Bench 1-100, Fable 5.1 minus Opus 5.5 in percentage points (our arithmetic)−2.4 points
Vals.ai Vibe Code Bench 1-100, the error margins shown added (our arithmetic)9.3
Vals.ai Vibe Code Bench v1.1, Fable 5.1 minus Opus 5.5 in percentage points (our arithmetic)0.0 points
Vals.ai Vibe Code Bench v1.1, the error margins shown added (our arithmetic)3.1
SRE Bench with fallbacks counted as failures, Fable 5.1 minus Opus 5.5 (our arithmetic)+5.3 points
Harvey's Legal Agent Benchmark, Fable 5.1 with fallbacks counted as failures minus Opus 5.5's published score (our arithmetic)+2.1 points
Terminal-Bench 4.0 with fallbacks counted as failures, Fable 5.1 minus Opus 5.5 (our arithmetic)−8.1 points
Vibe Code Bench v1.1, Fable 5.1's published score (no fallbacks shown) minus Opus 5.5's with fallbacks counted as failures (our arithmetic)+6.9 points

Vals.ai

Vals.ai, Claude Opus 5.5 · read 4 Oct 2026, 07:30 CEST
Vals.ai CUA-bench, Claude Opus 5.5 (max)14.00% ⟳
Vals.ai Code Migration, Claude Opus 5.5 (max)66.65% ⟳
Vals.ai Code Migration, Claude Opus 5.5: the page's error margin (percentage points)±4.33 ⟳
Vals.ai CyberBench v1.1, Claude Opus 5.5 (max)55.36% ⟳
Vals.ai CyberBench v1.1, Claude Opus 5.5: the page's error margin (percentage points)±5.13 ⟳
Vals.ai EMB, Claude Opus 5.5 (max)75.94% ⟳
Vals.ai EMB, Claude Opus 5.5: the page's error margin (percentage points)±2.38 ⟳
Vals.ai Finance Agent (v2), Claude Opus 5.5 (max)58.59% ⟳
Vals.ai Finance Agent (v2), Claude Opus 5.5: the page's error margin (percentage points)±0.17 ⟳
Vals.ai IOI, Claude Opus 5.5 (max)95.06% ⟳
Vals.ai IOI, Claude Opus 5.5: the page's error margin (percentage points)±4.94 ⟳
Vals.ai MedCode, Claude Opus 5.5 (max)49.80% ⟳
Vals.ai MedCode, Claude Opus 5.5: the page's error margin (percentage points)±2.27 ⟳
Vals.ai MedScribe, Claude Opus 5.5 (max)91.43% ⟳
Vals.ai MedScribe, Claude Opus 5.5: the page's error margin (percentage points)±1.93 ⟳
Vals.ai MysteryMechanism, Claude Opus 5.5 (max)49.55% ⟳
Vals.ai MysteryMechanism, Claude Opus 5.5: the page's error margin (percentage points)±3.36 ⟳
Vals.ai ProgramBench, Claude Opus 5.5 (max)18.50% ⟳
Vals.ai ProgramBench, Claude Opus 5.5: the page's error margin (percentage points)±2.75 ⟳
Vals.ai ProofBench v1.1, Claude Opus 5.5 (max)100.00% ⟳
Vals.ai ProofBench v1.1, Claude Opus 5.5: the page's error margin (percentage points)±0.00 ⟳
Vals.ai Public Benefits Bench v1.1, Claude Opus 5.5 (max)70.64% ⟳
Vals.ai Public Benefits Bench v1.1, Claude Opus 5.5: the page's error margin (percentage points)±1.19 ⟳
Vals.ai SAGE, Claude Opus 5.5 (max)45.83% ⟳
Vals.ai SAGE, Claude Opus 5.5: the page's error margin (percentage points)±3.36 ⟳
Vals.ai SRE Bench, Claude Opus 5.5 (max)33.59% ⟳
Vals.ai SRE Bench, Claude Opus 5.5: the page's error margin (percentage points)±2.92 ⟳
Vals.ai Tax Agent Bench, Claude Opus 5.5 (max)70.50% ⟳
Vals.ai Tax Agent Bench, Claude Opus 5.5: the page's error margin (percentage points)±3.15 ⟳
Vals.ai Terminal-Bench 4.0, Claude Opus 5.5 (max)65.15% ⟳
Vals.ai Terminal-Bench 4.0, Claude Opus 5.5: the page's error margin (percentage points)±0.00 ⟳
Vals.ai Terminal-Bench Science, Claude Opus 5.5 (max)47.14% ⟳
Vals.ai Terminal-Bench Science, Claude Opus 5.5: the page's error margin (percentage points)±6.01 ⟳
Vals.ai Time Horizon Index: KSP, Claude Opus 5.5 (max)91.33% ⟳
Vals.ai Time Horizon Index: KSP, Claude Opus 5.5: the page's error margin (percentage points)±0.00 ⟳
Vals.ai Vals Index, Claude Opus 5.5 (max)66.97% ⟳
Vals.ai Vals Index, Claude Opus 5.5: the page's error margin (percentage points)±0.89 ⟳
Vals.ai Vals RSI Index, Claude Opus 5.5 (max)37.31% ⟳
Vals.ai Vibe Code Bench 1-100, Claude Opus 5.5 (max)30.36% ⟳
Vals.ai Vibe Code Bench 1-100, Claude Opus 5.5: the page's error margin (percentage points)±4.83 ⟳
Vals.ai Vibe Code Bench v1.1, Claude Opus 5.5 (max)90.29% ⟳
Vals.ai Vibe Code Bench v1.1, Claude Opus 5.5: the page's error margin (percentage points)±1.53 ⟳
Vals.ai's Fallback Rate for Claude Opus 5.5, as its page shows it3.99% ⟳
Vals.ai's Refusal Rate for Claude Opus 5.5, as its page shows it0.79% ⟳
Vals.ai CUA-bench, Claude Opus 5.5: share of tasks a fallback model helped with, from the row's tooltip (0 where the page shows none)0.0% ⟳
Vals.ai Code Migration, Claude Opus 5.5: share of tasks a fallback model helped with, from the row's tooltip (0 where the page shows none)81.5% ⟳
Vals.ai CyberBench v1.1, Claude Opus 5.5: share of tasks a fallback model helped with, from the row's tooltip (0 where the page shows none)52.6% ⟳
Vals.ai EMB, Claude Opus 5.5: share of tasks a fallback model helped with, from the row's tooltip (0 where the page shows none)0.0% ⟳
Vals.ai Finance Agent (v2), Claude Opus 5.5: share of tasks a fallback model helped with, from the row's tooltip (0 where the page shows none)1.3% ⟳
Vals.ai IOI, Claude Opus 5.5: share of tasks a fallback model helped with, from the row's tooltip (0 where the page shows none)5.6% ⟳
Vals.ai MedCode, Claude Opus 5.5: share of tasks a fallback model helped with, from the row's tooltip (0 where the page shows none)0.0% ⟳
Vals.ai MedScribe, Claude Opus 5.5: share of tasks a fallback model helped with, from the row's tooltip (0 where the page shows none)0.0% ⟳
Vals.ai MysteryMechanism, Claude Opus 5.5: share of tasks a fallback model helped with, from the row's tooltip (0 where the page shows none)0.5% ⟳
Vals.ai ProgramBench, Claude Opus 5.5: share of tasks a fallback model helped with, from the row's tooltip (0 where the page shows none)12.0% ⟳
Vals.ai ProofBench v1.1, Claude Opus 5.5: share of tasks a fallback model helped with, from the row's tooltip (0 where the page shows none)0.0% ⟳
Vals.ai Public Benefits Bench v1.1, Claude Opus 5.5: share of tasks a fallback model helped with, from the row's tooltip (0 where the page shows none)1.3% ⟳
Vals.ai SAGE, Claude Opus 5.5: share of tasks a fallback model helped with, from the row's tooltip (0 where the page shows none)0.0% ⟳
Vals.ai SRE Bench, Claude Opus 5.5: share of tasks a fallback model helped with, from the row's tooltip (0 where the page shows none)82.8% ⟳
Vals.ai Tax Agent Bench, Claude Opus 5.5: share of tasks a fallback model helped with, from the row's tooltip (0 where the page shows none)0.5% ⟳
Vals.ai Terminal-Bench 4.0, Claude Opus 5.5: share of tasks a fallback model helped with, from the row's tooltip (0 where the page shows none)11.1% ⟳
Vals.ai Terminal-Bench Science, Claude Opus 5.5: share of tasks a fallback model helped with, from the row's tooltip (0 where the page shows none)2.9% ⟳
Vals.ai Time Horizon Index: KSP, Claude Opus 5.5: share of tasks a fallback model helped with, from the row's tooltip (0 where the page shows none)0.0% ⟳
Vals.ai Vals Index, Claude Opus 5.5: share of tasks a fallback model helped with, from the row's tooltip (0 where the page shows none)4.0% ⟳
Vals.ai Vals RSI Index, Claude Opus 5.5: share of tasks a fallback model helped with, from the row's tooltip (0 where the page shows none)0.0% ⟳
Vals.ai Vibe Code Bench 1-100, Claude Opus 5.5: share of tasks a fallback model helped with, from the row's tooltip (0 where the page shows none)16.0% ⟳
Vals.ai Vibe Code Bench v1.1, Claude Opus 5.5: share of tasks a fallback model helped with, from the row's tooltip (0 where the page shows none)8.0% ⟳
Vals.ai CUA-bench, Claude Opus 5.5 (for the chart)14.00%
Vals.ai Code Migration, Claude Opus 5.5 (for the chart)66.65%
Vals.ai CyberBench v1.1, Claude Opus 5.5 (for the chart)55.36%
Vals.ai EMB, Claude Opus 5.5 (for the chart)75.94%
Vals.ai Finance Agent (v2), Claude Opus 5.5 (for the chart)58.59%
Vals.ai IOI, Claude Opus 5.5 (for the chart)95.06%
Vals.ai MedCode, Claude Opus 5.5 (for the chart)49.80%
Vals.ai MedScribe, Claude Opus 5.5 (for the chart)91.43%
Vals.ai MysteryMechanism, Claude Opus 5.5 (for the chart)49.55%
Vals.ai ProgramBench, Claude Opus 5.5 (for the chart)18.50%
Vals.ai ProofBench v1.1, Claude Opus 5.5 (for the chart)100.00%
Vals.ai Public Benefits Bench v1.1, Claude Opus 5.5 (for the chart)70.64%
Vals.ai SAGE, Claude Opus 5.5 (for the chart)45.83%
Vals.ai SRE Bench, Claude Opus 5.5 (for the chart)33.59%
Vals.ai Tax Agent Bench, Claude Opus 5.5 (for the chart)70.50%
Vals.ai Terminal-Bench 4.0, Claude Opus 5.5 (for the chart)65.15%
Vals.ai Terminal-Bench Science, Claude Opus 5.5 (for the chart)47.14%
Vals.ai Time Horizon Index: KSP, Claude Opus 5.5 (for the chart)91.33%
Vals.ai Vals Index, Claude Opus 5.5 (for the chart)66.97%
Vals.ai Vals RSI Index, Claude Opus 5.5 (for the chart)37.31%
Vals.ai Vibe Code Bench 1-100, Claude Opus 5.5 (for the chart)30.36%
Vals.ai Vibe Code Bench v1.1, Claude Opus 5.5 (for the chart)90.29%

Vals.ai

Vals.ai, Claude Fable 5.1 · read 4 Oct 2026, 07:30 CEST
Vals.ai CUA-bench, Claude Fable 5.1 (max)13.17% ⟳
Vals.ai Code Migration, Claude Fable 5.1 (max)54.61% ⟳
Vals.ai Code Migration, Claude Fable 5.1: the page's error margin (percentage points)±4.81 ⟳
Vals.ai CyberBench v1.1, Claude Fable 5.1 (max)70.42% ⟳
Vals.ai CyberBench v1.1, Claude Fable 5.1: the page's error margin (percentage points)±5.43 ⟳
Vals.ai EMB, Claude Fable 5.1 (max)76.67% ⟳
Vals.ai EMB, Claude Fable 5.1: the page's error margin (percentage points)±2.08 ⟳
Vals.ai Finance Agent (v2), Claude Fable 5.1 (max)58.88% ⟳
Vals.ai Finance Agent (v2), Claude Fable 5.1: the page's error margin (percentage points)±2.06 ⟳
Vals.ai IOI, Claude Fable 5.1 (max)90.78% ⟳
Vals.ai IOI, Claude Fable 5.1: the page's error margin (percentage points)±4.65 ⟳
Vals.ai MedCode, Claude Fable 5.1 (max)53.51% ⟳
Vals.ai MedCode, Claude Fable 5.1: the page's error margin (percentage points)±2.17 ⟳
Vals.ai MedScribe, Claude Fable 5.1 (max)91.29% ⟳
Vals.ai MedScribe, Claude Fable 5.1: the page's error margin (percentage points)±1.95 ⟳
Vals.ai MysteryMechanism, Claude Fable 5.1 (max)47.75% ⟳
Vals.ai MysteryMechanism, Claude Fable 5.1: the page's error margin (percentage points)±3.36 ⟳
Vals.ai ProgramBench, Claude Fable 5.1 (max)7.00% ⟳
Vals.ai ProgramBench, Claude Fable 5.1: the page's error margin (percentage points)±1.81 ⟳
Vals.ai ProofBench v1.1, Claude Fable 5.1 (max)100.00% ⟳
Vals.ai ProofBench v1.1, Claude Fable 5.1: the page's error margin (percentage points)±0.00 ⟳
Vals.ai Public Benefits Bench v1.1, Claude Fable 5.1 (max)74.90% ⟳
Vals.ai Public Benefits Bench v1.1, Claude Fable 5.1: the page's error margin (percentage points)±1.13 ⟳
Vals.ai SAGE, Claude Fable 5.1 (max)48.53% ⟳
Vals.ai SAGE, Claude Fable 5.1: the page's error margin (percentage points)±3.34 ⟳
Vals.ai SRE Bench, Claude Fable 5.1 (max)22.90% ⟳
Vals.ai SRE Bench, Claude Fable 5.1: the page's error margin (percentage points)±2.60 ⟳
Vals.ai Tax Agent Bench, Claude Fable 5.1 (max)77.64% ⟳
Vals.ai Tax Agent Bench, Claude Fable 5.1: the page's error margin (percentage points)±2.83 ⟳
Vals.ai Terminal-Bench 4.0, Claude Fable 5.1 (max)58.08% ⟳
Vals.ai Terminal-Bench 4.0, Claude Fable 5.1: the page's error margin (percentage points)±3.31 ⟳
Vals.ai Terminal-Bench Science, Claude Fable 5.1 (max)40.00% ⟳
Vals.ai Terminal-Bench Science, Claude Fable 5.1: the page's error margin (percentage points)±5.90 ⟳
Vals.ai Time Horizon Index: KSP, Claude Fable 5.1 (max)63.33% ⟳
Vals.ai Time Horizon Index: KSP, Claude Fable 5.1: the page's error margin (percentage points)±0.00 ⟳
Vals.ai Vals Index, Claude Fable 5.1 (max)65.83% ⟳
Vals.ai Vals Index, Claude Fable 5.1: the page's error margin (percentage points)±1.11 ⟳
Vals.ai Vals RSI Index, Claude Fable 5.1 (max)36.09% ⟳
Vals.ai Vibe Code Bench 1-100, Claude Fable 5.1 (max)28.00% ⟳
Vals.ai Vibe Code Bench 1-100, Claude Fable 5.1: the page's error margin (percentage points)±4.49 ⟳
Vals.ai Vibe Code Bench v1.1, Claude Fable 5.1 (max)90.26% ⟳
Vals.ai Vibe Code Bench v1.1, Claude Fable 5.1: the page's error margin (percentage points)±1.57 ⟳
Vals.ai's Fallback Rate for Claude Fable 5.1, as its page shows it2.10% ⟳
Vals.ai's Refusal Rate for Claude Fable 5.1, as its page shows it0.22% ⟳
Vals.ai CUA-bench, Claude Fable 5.1: share of tasks a fallback model helped with, from the row's tooltip (0 where the page shows none)0.0% ⟳
Vals.ai Code Migration, Claude Fable 5.1: share of tasks a fallback model helped with, from the row's tooltip (0 where the page shows none)1.5% ⟳
Vals.ai CyberBench v1.1, Claude Fable 5.1: share of tasks a fallback model helped with, from the row's tooltip (0 where the page shows none)53.5% ⟳
Vals.ai EMB, Claude Fable 5.1: share of tasks a fallback model helped with, from the row's tooltip (0 where the page shows none)0.0% ⟳
Vals.ai Finance Agent (v2), Claude Fable 5.1: share of tasks a fallback model helped with, from the row's tooltip (0 where the page shows none)0.0% ⟳
Vals.ai IOI, Claude Fable 5.1: share of tasks a fallback model helped with, from the row's tooltip (0 where the page shows none)22.2% ⟳
Vals.ai MedCode, Claude Fable 5.1: share of tasks a fallback model helped with, from the row's tooltip (0 where the page shows none)0.0% ⟳
Vals.ai MedScribe, Claude Fable 5.1: share of tasks a fallback model helped with, from the row's tooltip (0 where the page shows none)0.0% ⟳
Vals.ai MysteryMechanism, Claude Fable 5.1: share of tasks a fallback model helped with, from the row's tooltip (0 where the page shows none)0.0% ⟳
Vals.ai ProgramBench, Claude Fable 5.1: share of tasks a fallback model helped with, from the row's tooltip (0 where the page shows none)16.5% ⟳
Vals.ai ProofBench v1.1, Claude Fable 5.1: share of tasks a fallback model helped with, from the row's tooltip (0 where the page shows none)0.0% ⟳
Vals.ai Public Benefits Bench v1.1, Claude Fable 5.1: share of tasks a fallback model helped with, from the row's tooltip (0 where the page shows none)0.4% ⟳
Vals.ai SAGE, Claude Fable 5.1: share of tasks a fallback model helped with, from the row's tooltip (0 where the page shows none)0.0% ⟳
Vals.ai SRE Bench, Claude Fable 5.1: share of tasks a fallback model helped with, from the row's tooltip (0 where the page shows none)75.2% ⟳
Vals.ai Tax Agent Bench, Claude Fable 5.1: share of tasks a fallback model helped with, from the row's tooltip (0 where the page shows none)0.0% ⟳
Vals.ai Terminal-Bench 4.0, Claude Fable 5.1: share of tasks a fallback model helped with, from the row's tooltip (0 where the page shows none)11.1% ⟳
Vals.ai Terminal-Bench Science, Claude Fable 5.1: share of tasks a fallback model helped with, from the row's tooltip (0 where the page shows none)5.7% ⟳
Vals.ai Time Horizon Index: KSP, Claude Fable 5.1: share of tasks a fallback model helped with, from the row's tooltip (0 where the page shows none)0.0% ⟳
Vals.ai Vals Index, Claude Fable 5.1: share of tasks a fallback model helped with, from the row's tooltip (0 where the page shows none)2.1% ⟳
Vals.ai Vals RSI Index, Claude Fable 5.1: share of tasks a fallback model helped with, from the row's tooltip (0 where the page shows none)0.0% ⟳
Vals.ai Vibe Code Bench 1-100, Claude Fable 5.1: share of tasks a fallback model helped with, from the row's tooltip (0 where the page shows none)28.0% ⟳
Vals.ai Vibe Code Bench v1.1, Claude Fable 5.1: share of tasks a fallback model helped with, from the row's tooltip (0 where the page shows none)0.0% ⟳
Vals.ai CUA-bench, Claude Fable 5.1 (for the chart)13.17%
Vals.ai Code Migration, Claude Fable 5.1 (for the chart)54.61%
Vals.ai CyberBench v1.1, Claude Fable 5.1 (for the chart)70.42%
Vals.ai EMB, Claude Fable 5.1 (for the chart)76.67%
Vals.ai Finance Agent (v2), Claude Fable 5.1 (for the chart)58.88%
Vals.ai IOI, Claude Fable 5.1 (for the chart)90.78%
Vals.ai MedCode, Claude Fable 5.1 (for the chart)53.51%
Vals.ai MedScribe, Claude Fable 5.1 (for the chart)91.29%
Vals.ai MysteryMechanism, Claude Fable 5.1 (for the chart)47.75%
Vals.ai ProgramBench, Claude Fable 5.1 (for the chart)7.00%
Vals.ai ProofBench v1.1, Claude Fable 5.1 (for the chart)100.00%
Vals.ai Public Benefits Bench v1.1, Claude Fable 5.1 (for the chart)74.90%
Vals.ai SAGE, Claude Fable 5.1 (for the chart)48.53%
Vals.ai SRE Bench, Claude Fable 5.1 (for the chart)22.90%
Vals.ai Tax Agent Bench, Claude Fable 5.1 (for the chart)77.64%
Vals.ai Terminal-Bench 4.0, Claude Fable 5.1 (for the chart)58.08%
Vals.ai Terminal-Bench Science, Claude Fable 5.1 (for the chart)40.00%
Vals.ai Time Horizon Index: KSP, Claude Fable 5.1 (for the chart)63.33%
Vals.ai Vals Index, Claude Fable 5.1 (for the chart)65.83%
Vals.ai Vals RSI Index, Claude Fable 5.1 (for the chart)36.09%
Vals.ai Vibe Code Bench 1-100, Claude Fable 5.1 (for the chart)28.00%
Vals.ai Vibe Code Bench v1.1, Claude Fable 5.1 (for the chart)90.26%

Vals.ai

Vals.ai, Claude Opus 5.5 (the page's launch notes) · read 4 Oct 2026, 07:30 CEST
Vals.ai SRE Bench, Opus 5.5 with fallback-assisted tasks counted as failures5.34%
Vals.ai SRE Bench, Opus 5.5: tasks a fallback model helped with217
Vals.ai SRE Bench: tasks262
Vals.ai Vibe Code Bench, Opus 5.5 with fallback-assisted tasks counted as failures (the launch note's figure; the current score matches its 90.29%)83.34%
Vals.ai MysteryMechanism, Opus 5.5 with fallback-assisted tasks counted as failures (the launch note's figure; the current score matches its 49.55%)49.10%

Vals.ai

Vals.ai, Claude Fable 5.1 (the page's launch notes) · read 4 Oct 2026, 07:30 CEST
Vals.ai SRE Bench, Fable 5.1 with fallback-assisted tasks counted as failures10.69%
Vals.ai SRE Bench, Fable 5.1: tasks a fallback model helped with195
Vals.ai Harvey's Legal Agent Benchmark, Fable 5.1 with fallback-assisted tasks counted as failures5.83%

Vals.ai

Vals.ai, Terminal-Bench 4.0 · read 4 Oct 2026, 07:48 CEST
Vals.ai Terminal-Bench 4.0, Opus 5.5 with fallback-served attempts counted as failures58.08%
Vals.ai Terminal-Bench 4.0, Fable 5.1 with fallback-served attempts counted as failures50.00%

Sources

These are the pages this article draws on. We keep a copy of each page as we read it, so a figure can be checked against what the page said at the time.