We gave two models the same prompt and the same picture, each on the effort setting Claude Code gives it by default: high for the older Sonnet 5, medium for Sonnet 5.5. Sonnet 5 wrote 66,874 output tokens to paint it. Sonnet 5.5 wrote 10,660. That's one run each, on a test we built, and it fits what Anthropic said at launch on 28 September, that Sonnet 5.5 "typically needs far fewer tokens to do the same work."
About half an hour after Anthropic's post, Artificial Analysis published a number pointing the other way. By its count, Sonnet 5.5 at max effort uses more output tokens per task than any model it has measured, ~193k, and costs 49% more per task than Sonnet 5. Both numbers are right. Which one applies to you depends on one setting, called effort.
Sonnet 5.5 is the middle model of Anthropic's 5.5 family. Opus 5.5 came out the week before, and Anthropic says a Haiku 5.5 will join "in the coming weeks". You can use Sonnet 5.5 in the Claude apps, in Claude Code and on the API, including through Amazon's, Google's and Microsoft's clouds.
It costs the same per token as Sonnet 5: $2 per million input tokens and $10 per million output tokens, half of Opus 5.5's $4 and $20. Anthropic says it writes its output 30%+ faster than Sonnet 5.
Its headline score on Terminal-Bench 4.0, an agentic coding test, is 70.6%, where Sonnet 5 scored 10.3%. On GDPval-AA, a test of office work that is scored like a chess rating, it rates 1844 against Opus 5.5's 1846. Two independent evaluators, both running it at max, rank it second behind Opus 5.5: Artificial Analysis on its Intelligence Index, and Vals AI, where it scores 69.22% against Opus 5.5's 69.69%, a gap inside Vals's own ±0.96 margin, and places #2 of 66. Anthropic itself says "Opus 5.5 remains clearly stronger at complex, open-ended work requiring sustained judgment."
Like other recent Claude models, Sonnet 5.5 has five effort levels: low, medium, high, xhigh (extra high) and max. Effort controls how much the model reasons and how thoroughly it works, and you can change it. The Claude apps and Claude Code start it on medium, and the API starts it on high.
Anthropic ran its headline scores at max. One step down, at xhigh, the system card puts the GDPval-AA rating at 1725, with about 67% fewer output tokens than at max. On Terminal-Bench 4.0, Anthropic's table puts Sonnet 5.5 at max 4.2 points ahead of Opus 5.5 at xhigh, 70.6% against 66.4%. The system card gives those scores standard errors of ±2.5 and ±2.6 points, which makes a gap of that size too small to call.
In 10% of Opus 5.5's trials on that test, a request flagged by Anthropic's safeguards was answered by an older model instead. For Sonnet 5.5 it was 1.5% of trials. Anthropic says this likely lowered Opus 5.5's score.
On launch day, Anthropic's staff posted a demo in which models paint a picture by writing their own brush engine in Python, looking at the result and revising. We rebuilt that test and gave it Hokusai's Great Wave, a woodblock print from c. 1830–32. Five models got the same prompt, one run each: Sonnet 5, Sonnet 5.5 and Opus 5.5 in Claude Code, and OpenAI's GPT-6 Sol and GPT-6 Astra in Codex, each on its tool's defaults.
Five models, one prompt, the Great Wave
The printKatsushika Hokusai, Under the Wave off Kanagawa (The Great Wave), c. 1830–32. The Metropolitan Museum of Art, H. O. Havemeyer Collection, Bequest of Mrs. H. O. Havemeyer, 1929 (JP1847). Public domain.Sonnet 5Claude Code · high · 1 run 66,874 output tokensSonnet 5.5Claude Code · medium · 1 run 10,660 output tokensOpus 5.5Claude Code · medium · 1 run 28,600 output tokensGPT-6 SolCodex default · 1 run tokens not comparable (different tool)GPT-6 AstraCodex default · 1 run tokens not comparable (different tool)
One run per model on 28 September 2026, with the same prompt for all five. We give output tokens only for the Claude runs, because the OpenAI runs used a different tool and their counts aren't comparable.
Sonnet 5 painted one big, smooth wave. Sonnet 5.5 got the print's blue stripes and its spray. Opus 5.5 got closer to the print, with the foam curling into claws all along the crest, and Astra's painting puts the print's curling foam on every wave.
One run is too few to count tokens by, so we ran each Claude model 5 times, on the Great Wave and on two other pictures. Sonnet 5.5's median run wrote 9,481 output tokens. Sonnet 5's median was 57,847, 6.1× as many, and Opus 5.5's was 21,606, 2.3× as many. The OpenAI runs used a different tool, so their counts aren't comparable and we leave them out.
Output tokens per run, our painting test
Five runs per Claude model: three on one photo, one on the Great Wave, one on a chart. Each model at its Claude Code default
Our runs, Claude Code 2.1.284, each model at its default (Sonnet 5 high, Sonnet 5.5 and Opus 5.5 medium), 28 September 2026. The line marks each model's median.The numbers in this chart
This says what we saw on one small job, not how the models compare in general. The full painting test gets its own video, and its prompts, runs and numbers will be published with it.
Every setting on Artificial Analysis's index▶ 4:38
One painting job is small. Artificial Analysis ran both Sonnets across its whole Intelligence Index at all five settings, on a pre-release version of Sonnet 5.5, and counted the tokens.
What each setting costs
Cost per task on Artificial Analysis's index, Sonnet 5 against Sonnet 5.5, at each effort setting
Sonnet 5Sonnet 5.5
Artificial Analysis Intelligence Index, average per task, list prices including caching. Pre-release deployment of Sonnet 5.5; Artificial Analysis says it will re-run. Read 28 Sep 2026, 21:29 CEST. Re-read 29 Sep 2026, 12:28 CEST: none of the values in this chart has moved.The numbers in this chart
At low, medium and high, Sonnet 5.5 writes fewer output tokens than Sonnet 5 at the same setting and costs less per task. At medium, the Claude Code default, it costs 41% less per task, $0.59 against $1.00, and it scores 40.7, higher than Sonnet 5 scores at max (38.2). So Anthropic's claim holds on someone else's test.
Above high, that changes. At xhigh the two cost about the same, $2.74 against $2.87. At max, Sonnet 5.5 writes 192.8K output tokens per task, 64% more than Sonnet 5 and the most Artificial Analysis has measured for any model. It costs 49% more per task, $7.60 against $5.09. And max is the setting Anthropic's headline scores came from.
What each setting writes
Output tokens per task on Artificial Analysis's index, Sonnet 5 against Sonnet 5.5
Sonnet 5Sonnet 5.5
Artificial Analysis Intelligence Index, average per task, list prices including caching. Pre-release deployment of Sonnet 5.5; Artificial Analysis says it will re-run. Read 28 Sep 2026, 21:29 CEST. Re-read 29 Sep 2026, 12:28 CEST: none of the values in this chart has moved.The numbers in this chart
Anthropic's launch page has four charts of score against cost (per attempt on Terminal-Bench, per task on the other three), and on all four, Sonnet 5.5 at max costs more than Sonnet 5 at max.
Anthropic's own charts, both Sonnets at max
Cost at max effort on the four charts of Anthropic's launch page: per attempt on Terminal-Bench, per task on the other three
Sonnet 5Sonnet 5.5
Redrawn from Anthropic's launch page; values read off the chart, so approximate. Page read 28 Sep 2026, 21:17 CEST.The numbers in this chart
On FrontierCode, a test of whether a code change would get merged, Sonnet 5.5 scores lower at max than at xhigh, 46.2% against 52.1% on Anthropic's table, while costing about 13× as much, $20.73 a task against $1.59. The costs are read off Anthropic's chart, so they are approximate. Anthropic's footnote says that at max it more often ran a code-review skill that splits the work across many subagents. Cognition, which ran that test, looked at two cases where this led to a timeout or to edits outside the task, and so to a lower score.
Simon Willison, on Hacker News. Read 28 Sep 2026, 21:27 CEST.
Low took 10s and medium 11s, high 17s and xhigh 41s. At max it thought for 15m 40s, hit the limit for a single response, 128,000 output tokens, and returned no pelican at all.
At medium, Sonnet 5.5 is a straight upgrade on Sonnet 5. It scores higher and costs less per task on Artificial Analysis's index and on all four of Anthropic's charts. At high, the API's default, the same is true on all but one chart, AA-Briefcase, where it costs a little more: $3.96 a task against $3.77.
Even at the Claude Code default, Opus 5.5 is worth comparing on Artificial Analysis's index. Opus 5.5 at low scores slightly higher than Sonnet 5.5 at medium for slightly less: 42.3 against 40.7, at $0.55 a task against $0.59. At high, no Opus 5.5 setting scores higher than Sonnet 5.5 for less.
Above high, the gap in Opus 5.5's favour grows. On Artificial Analysis's index, Opus 5.5 at medium scores about the same as Sonnet 5.5 at xhigh, 51.2 against 51.9, for about half the cost: $1.34 a task against $2.74. Opus 5.5 at xhigh matches Sonnet 5.5 at max, 56.0 against 56.0, for less than half: $3.46 against $7.60.
What to run it at
Intelligence Index score against cost per task, each model at its five effort settings
Sonnet 5.5Sonnet 5Opus 5.5GPT-6 Sol
Artificial Analysis Intelligence Index, average per task, list prices including caching. Pre-release deployment of Sonnet 5.5; Artificial Analysis says it will re-run. Read 28 Sep 2026, 21:29 CEST. Re-read 29 Sep 2026, 12:28 CEST: 3 values in this chart have moved since, all GPT-6 Sol's cost per index task; the numbers table below has them.The numbers in this chart
Most of the difference is cache reads. At those settings Sonnet 5.5 reads 3.6× and 3.2× as much from its cache as Opus 5.5, and cache reads cost $0.20 per million tokens on both, so Opus's higher price doesn't apply there. Of the $1.41 gap per task between Sonnet 5.5 at xhigh and Opus 5.5 at medium, $1.07 is cache reads and $0.22 is output; of the $4.14 gap between Sonnet 5.5 at max and Opus 5.5 at xhigh, $3.04 is cache reads and $0.62 is output. Opus 5.5 costs twice as much per token it writes, but Sonnet 5.5 writes about three times as many (2.9× and 2.9×), which cancels most of that. So if a job needs Sonnet 5.5 at xhigh or max, try Opus 5.5 a step or two lower first.
If you're not tied to Anthropic, OpenAI's GPT-6 Sol has, on the same index, a setting that scores higher for less than each of Sonnet 5.5's settings from low to high. At low and medium it is much cheaper: Sol at medium scores 39.8 for $0.25 a task against Sonnet 5.5 at low with 35.8 for $0.41 (40% less), and Sol at high scores 42.8 for $0.37now $0.38 against Sonnet 5.5 at medium with 40.7 for $0.59 (36% less). At high the two cost about the same: Sol at max scores 47.5 for $1.06now $1.05, against Sonnet 5.5 at high with 46.7 for $1.08.
If you do security work, Anthropic says "higher-risk cybersecurity tasks will visibly fall back to Sonnet 5." On Anthropic's own API that fallback is a beta option you switch on. Without it, those requests are refused.
All of that is in API dollars. If you pay for a Claude plan instead, Anthropic's ClaudeDevs account on X says Sonnet 5.5 means "your Claude Code usage goes further". None of the Anthropic pages we saved on launch day (the launch page, the Sonnet product page and the migration guide, read 28 September) says how plan usage is counted for Sonnet against Opus, so from those pages, nobody can say how much further.
Most of the launch-day posts you'll see aren't measurements. The biggest are Anthropic's own and its staff's. Then come people who had the model early: Matthew Berman and two editors from his channel, and Alex Finn, who wrote that at first he couldn't tell it apart from Opus 5.5. Then come companies quoted on Anthropic's launch page, such as Box and Every, posting their own results.
One of those posts is useful. Kieran Klaassen, who works at Every, says to run it on low or medium, and to move to Opus if a job needs more. Artificial Analysis's numbers agree with him at medium, by a small margin: Opus 5.5 at low scores 42.3 against 40.7, at $0.55 a task against $0.59. At high they don't: no Opus 5.5 setting scores higher than Sonnet 5.5 for less.
None of this tells you how Sonnet 5.5 does on your own work, at the setting you use. Artificial Analysis ran a pre-release version that Anthropic found had a bug with structured outputs, and it says it will re-run. Our painting test is one small job. And when we made the video, we hadn't found anyone outside Anthropic who had published the comparison that would settle it: everyday work, like a long coding session, at each setting, run more than once, with the tokens counted.
Five days before Sonnet 5.5, Fireworks released a model called Ember-1 that makes the same promise: "half the tokens, same answers". That's our next video.
Every number
These are all 304 figures behind the video and this page, grouped by whose they are, with the page each came from and when we read it. Figures marked ⟳ can move. When a re-read finds a change, the new value shows next to the one from the video.
Artificial Analysis Intelligence Index, by effort setting
Our Great Wave runs and token runs (Claude Code 2.1.284 and Codex, each tool's defaults) · run 28 Sep 2026, 21:15–22:41 CEST (start times of the fifteen runs)
output tokens, Sonnet 5 painting the Great Wave, 1 run, Claude Code default (high)
These are the pages the video and this page draw on. We keep a copy of each page as we read it, so a figure can be checked against what the page said at the time.