The short answer
On Artificial Analysis's leaderboard as we read it on 4 October 2026, 102 current model settings have a score and a cost per task above zero. Only 15 of them are beaten by no other setting on score and cost at once: GPT-6 Luna at every reasoning setting from low to max, Xiaomi's MiMo-V2.6-Flash and MiMo-V2.6-Pro, GPT-6.1 Sol at every setting, and Claude Opus 5.5 at high, xhigh and max.
Score against cost per task, every current setting with a cost per task
Artificial Analysis Intelligence Index up, cost per task across on a log scale. The line is the best score on the list at or under each cost.
The numbers in this chart
| score | cost per task | |
|---|---|---|
| Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) | 57.6 | $5.98 |
| Claude Sonnet 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) | 56.0 | $7.67 |
| Claude Opus 5.5 (Adaptive Reasoning, Xhigh Effort, Default Fallback) | 56.0 | $3.46 |
| Claude Opus 5.5 (Adaptive Reasoning, High Effort, Default Fallback) | 53.6 | $1.82 |
| Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) | 53.4 | $7.63 |
| Claude Fable 5.1 (Adaptive Reasoning, Xhigh Effort, Default Fallback) | 53.2 | $5.98 |
| GPT-6 Astra (Max) | 52.7 | $3.26 |
| Gemini 4 Argon (High) | 52.6 | $1.99 |
| GPT-6 Astra (Xhigh) | 52.4 | $2.31 |
| Claude Sonnet 5.5 (Adaptive Reasoning, Xhigh Effort, Default Fallback) | 51.9 | $2.75 |
| GPT-6.1 Sol (Max) | 51.8 | $0.72 |
| Claude Opus 5.5 (Adaptive Reasoning, Medium Effort, Default Fallback) | 51.2 | $1.34 |
| Claude Fable 5.1 (Adaptive Reasoning, High Effort, Default Fallback) | 51.2 | $3.91 |
| GPT-6.1 Sol (Xhigh) | 51.0 | $0.39 |
| GPT-6 Astra (High) | 50.9 | $1.73 |
| GPT-6.1 Sol (High) | 50.2 | $0.32 |
| GPT-6 Astra (Medium) | 49.6 | $1.54 |
| Claude Fable 5.1 (Adaptive Reasoning, Medium Effort, Default Fallback) | 48.9 | $2.98 |
| Muse Spark 1.3 (Max) | 48.1 | $1.60 |
| GPT-6.1 Sol (Medium) | 47.8 | $0.21 |
| Claude Fable 5.1 (Adaptive Reasoning, Low Effort, Default Fallback) | 46.8 | $2.37 |
| Claude Sonnet 5.5 (Adaptive Reasoning, High Effort, Default Fallback) | 46.8 | $1.12 |
| Grok 4.7 (Xhigh) | 46.4 | $3.74 |
| Grok 4.7 (High) | 46.3 | $2.73 |
| MiMo-V2.6-Pro | 46.3 | $0.13 |
| GPT-6 Astra (Low) | 45.8 | $0.82 |
| Qwen3.8 Max (0902) | 45.4 | $5.41 |
| Muse Spark 1.3 (Xhigh) | 45.1 | $1.37 |
| GLM-5.3 (Max) | 44.8 | $2.01 |
| Step 5 Preview | 43.7 | $0.72 |
| Kimi K3 (Max) | 43.6 | $2.00 |
| Claude Opus 5.5 (Adaptive Reasoning, Low Effort, Default Fallback) | 42.3 | $0.55 |
| Grok 4.7 (Low) | 42.2 | $1.25 |
| GPT-6.1 Sol (Low) | 42.1 | $0.13 |
| GPT-5.6 Terra (Max) | 42.1 | $1.40 |
| GLM 5.3 Flash | 41.8 | $0.25 |
| Gemini 3.8 Flash (High) | 40.9 | $1.24 |
| Claude Sonnet 5.5 (Adaptive Reasoning, Medium Effort, Default Fallback) | 40.8 | $0.59 |
| Qwen3.8 2.4T A95B | 39.9 | $2.16 |
| Qwen3.8-Flash-Next | 39.8 | $0.37 |
| Gemini 3.8 Flash (Medium) | 39.8 | $0.93 |
| DeepSeek V4.1 Flash (Max) | 39.5 | $0.27 |
| GPT-6 Luna (Max) | 38.1 | $0.068 |
| GPT-5.6 Terra (Xhigh) | 38.0 | $0.63 |
| MiMo-V2.6-Flash | 37.9 | $0.062 |
| DeepSeek V4 Pro 0813 (Max) | 36.0 | $0.67 |
| Claude Sonnet 5.5 (Adaptive Reasoning, Low Effort, Default Fallback) | 35.9 | $0.42 |
| DeepSeek V4 Flash Vision (Max) | 34.8 | $0.31 |
| GPT-6 Luna (Xhigh) | 34.6 | $0.042 |
| GLM-5.3 (Low) | 34.3 | $0.85 |
| GPT-5.6 Terra (High) | 34.2 | $0.34 |
| Qwen3.8 27B (Xhigh) | 33.7 | $1.01 |
| GPT-6 Luna (High) | 32.9 | $0.029 |
| GPT-5.6 Terra (Medium) | 30.1 | $0.18 |
| Kimi K3 (Low) | 30.1 | $1.15 |
| GPT-6 Luna (Medium) | 29.9 | $0.017 |
| Gemini 3.1 Pro Preview | 29.7 | $0.67 |
| MiniMax-M3 | 29.2 | $0.51 |
| Qwen3.8 27B (Medium) | 27.6 | $1.13 |
| GPT-5.6 Terra (Low) | 27.5 | $0.14 |
| Quasar 438B (Max, Based on GLM-5.2) | 26.7 | $2.02 |
| Apodex 1.1 | 26.4 | $0.46 |
| Qwen3.8 27B (Low) | 26.2 | $1.05 |
| GPT-5.5 Instant (June 2026) | 26.0 | $0.69 |
| Kimi K2.7 Code | 25.8 | $0.54 |
| Inkling Small | 25.7 | $0.089 |
| Hy3 | 25.3 | $0.072 |
| Qwen3.7 Plus | 25.2 | $0.22 |
| DeepSeek V4.1 Flash (Non-reasoning) | 24.7 | $0.15 |
| Solar Mini 4 | 24.1 | $0.36 |
| Nemotron 3 Ultra 550B A55B (Reasoning) | 22.9 | $0.60 |
| Gemini 3.5 Flash-Lite | 22.2 | $0.12 |
| GPT-6 Luna (Low) | 21.5 | $0.0045 |
| GPT-5.6 Terra (Non-reasoning) | 20.8 | $0.14 |
| DeepSeek V4 Pro 0813 (Non-reasoning) | 20.4 | $0.48 |
| Qwen3.8 27B (Non-reasoning) | 20.2 | $2.49 |
| LongCat 2.0 | 19.1 | $0.059 |
| GPT-6 Luna (Non-reasoning) | 18.5 | $0.011 |
| Qwen3.5 397B A17B (Reasoning) | 18.4 | $0.47 |
| Qwen3.6 35B A3B (Reasoning) | 18.2 | $0.48 |
| Muse Glimmer (High) | 17.5 | $0.057 |
| Claude 4.5 Haiku (Reasoning) | 16.9 | $0.28 |
| Ring-2.6-1T | 16.6 | $0.29 |
| Qwen3.5 122B A10B (Reasoning) | 15.6 | $0.32 |
| Mistral Medium 3.5 | 14.2 | $0.50 |
| Nemotron 3.5 Lightning | 12.9 | $0.093 |
| Nemotron 3 Super 120B A12B (Reasoning) | 12.8 | $1.64 |
| Mercury 2.5 | 12.3 | $0.12 |
| gpt-oss-120b (High) | 11.6 | $0.11 |
| Mistral Small 4 (Reasoning) | 11.3 | $0.015 |
| Qwen3.5 9B (Reasoning) | 11.2 | $0.21 |
| Granite 4.2 8B | 11.1 | $0.024 |
| Trinity Large Thinking | 10.8 | $0.12 |
| Mistral Large 3 | 9.3 | $0.031 |
| Qwen3 Coder Next | 9.2 | $0.55 |
| Granite 4.2 3B | 9.1 | $0.0060 |
| gpt-oss-20b (High) | 9.0 | $0.012 |
| NVIDIA Nemotron 3 Nano 30B A3B (Reasoning) | 8.9 | $0.017 |
| Celeris-1 | 6.3 | $0.050 |
| Ministral 3 14B | 6.0 | $0.019 |
| Ministral 3 8B | 5.5 | $0.011 |
| Ministral 3 3B | 4.8 | $0.0078 |
Marked rows are the ones no other row beats on both.
Gemini 4 Argon, new on the list since our first read on 30 September, isn't among them: it scores 52.6 for $1.99 a task at Google's introductory price, and Opus 5.5 at high scores 1.0 points more for 8% less. No setting of Sonnet 5.5, Fable 5.1 or GPT-6 Astra is on the line either.
Every score and cost here is Artificial Analysis's. What we added is the selection and the arithmetic, and both rest on one aggregate score read on one morning.
How the line is drawn
Artificial Analysis publishes many figures for each model, and this article uses two of them: its Intelligence Index, one score across its tests, and a cost per task in US dollars. Its leaderboard lists each reasoning setting as a row of its own, so GPT-6.1 Sol at low and GPT-6.1 Sol at max are two entries.
The page's data holds 688 rows. We took the ones it doesn't mark as deprecated and that have a score and a cost per task above zero, which leaves 102. The page gives a cost of zero for 5 more current rows, and we left those out, since a zero can't be compared with a price. Another 152 current rows have a score but no cost per task on the page, so they can't be placed on the line either. Then we kept every row that scores higher than all the rows costing the same or less. That leaves 15.
Read the chart at the top of the page from left to right. At any cost along the bottom, the line's height is the best score on the list that costs that much or less. The hollow points are the settings you could pick instead, each beaten by something on or under the line.
Where the extra points get expensive
The bottom of the line belongs to GPT-6 Luna, whose five settings from low to max go from $0.0045 to $0.068 a task, with Xiaomi's MiMo-V2.6-Flash between its xhigh and max settings. Then the line moves through GPT-6.1 Sol and MiMo-V2.6-Pro. MiMo-V2.6-Pro costs 2% more per task than GPT-6.1 Sol at low and scores 4.2 points higher. Both MiMo models are open-weights models, the only ones on the line.
GPT-6.1 Sol covers the middle at every setting it has. From low to max its score rises from 42.1 to 51.8, and its cost goes from $0.13 to $0.72. The last step costs the most for the least, since going from xhigh to max adds 0.8 points for 84% more per task.
Above Sol's best score, the only rows on the line are Claude Opus 5.5 at high, xhigh and max. The first of them adds 1.7 points over Sol at max for 2.5× the cost. Opus at max scores 5.8 points more than Sol at max and costs 8.3× as much. Opus's own climb costs a lot too: high to max adds 4.0 points at 3.3× the cost.
The rows on the line, cheapest first
Each step is against the row above it.
| setting | score | cost per task | points more | cost multiple |
|---|---|---|---|---|
| GPT-6 Luna, low | 21.5 | $0.0045 | – | – |
| GPT-6 Luna, medium | 29.9 | $0.017 | +8.4 | 3.87× |
| GPT-6 Luna, high | 32.9 | $0.029 | +3.0 | 1.66× |
| GPT-6 Luna, xhigh | 34.6 | $0.042 | +1.6 | 1.45× |
| MiMo-V2.6-Flash | 37.9 | $0.062 | +3.3 | 1.47× |
| GPT-6 Luna, max | 38.1 | $0.068 | +0.2 | 1.09× |
| GPT-6.1 Sol, low | 42.1 | $0.13 | +4.0 | 1.93× |
| MiMo-V2.6-Pro | 46.3 | $0.13 | +4.2 | 1.02× |
| GPT-6.1 Sol, medium | 47.8 | $0.21 | +1.5 | 1.60× |
| GPT-6.1 Sol, high | 50.2 | $0.32 | +2.5 | 1.49× |
| GPT-6.1 Sol, xhigh | 51.0 | $0.39 | +0.8 | 1.23× |
| GPT-6.1 Sol, max | 51.8 | $0.72 | +0.8 | 1.84× |
| Opus 5.5, high | 53.6 | $1.82 | +1.7 | 2.52× |
| Opus 5.5, xhigh | 56.0 | $3.46 | +2.4 | 1.90× |
| Opus 5.5, max | 57.6 | $5.98 | +1.6 | 1.73× |
Settings that cost more for less
Every row off the line has another that costs no more and scores at least as high. The table puts some of the familiar ones next to theirs.
Beaten settings, next to the cheapest row that scores as high
The best current row from several labs that isn't on the line, plus Fable 5.1, Opus 5.5 at medium and Gemini 3.8 Flash at high, each against the cheapest row that scores at least as high.
| setting | score | cost per task | cheapest row scoring as high | its score | its cost | its cost as a share |
|---|---|---|---|---|---|---|
| Sonnet 5.5, max | 56.0 | $7.67 | Opus 5.5, max | 57.6 | $5.98 | 78% |
| Fable 5.1, max | 53.4 | $7.63 | Opus 5.5, high | 53.6 | $1.82 | 24% |
| GPT-6 Astra, max | 52.7 | $3.26 | Opus 5.5, high | 53.6 | $1.82 | 56% |
| Gemini 4 Argon, high | 52.6 | $1.99 | Opus 5.5, high | 53.6 | $1.82 | 92% |
| Opus 5.5, medium | 51.2 | $1.34 | GPT-6.1 Sol, max | 51.8 | $0.72 | 54% |
| Muse Spark 1.3, max | 48.1 | $1.60 | GPT-6.1 Sol, high | 50.2 | $0.32 | 20% |
| Grok 4.7, xhigh | 46.4 | $3.74 | GPT-6.1 Sol, medium | 47.8 | $0.21 | 6% |
| Qwen3.8 Max | 45.4 | $5.41 | MiMo-V2.6-Pro | 46.3 | $0.13 | 2% |
| GLM-5.3, max | 44.8 | $2.01 | MiMo-V2.6-Pro | 46.3 | $0.13 | 7% |
| Kimi K3, max | 43.6 | $2.00 | MiMo-V2.6-Pro | 46.3 | $0.13 | 7% |
| Gemini 3.8 Flash, high | 40.9 | $1.24 | GPT-6.1 Sol, low | 42.1 | $0.13 | 11% |
| DeepSeek V4.1 Flash, max | 39.5 | $0.27 | GPT-6.1 Sol, low | 42.1 | $0.13 | 49% |
The Claude rows are the ones a Claude user chooses between. Fable 5.1 at max scores 53.4 for $7.63 a task, and Opus 5.5 at high scores more for 24% of that. Opus 5.5 at medium costs more than GPT-6.1 Sol at max, which scores higher.
Sonnet 5.5 at max and Opus 5.5 at xhigh show the same score to one decimal, 56.0. In the page's unrounded data Sonnet's is a little higher, so the cheapest row that scores at least as high as Sonnet at max is Opus at max, at 78% of Sonnet's cost. Opus at xhigh costs $3.46 a task against Sonnet's $7.67.
What changed since 30 September
We first read the leaderboard for this article on 30 September and didn't publish that version. Between the two reads, 4 rows became current with a score and a price, 6 left the current list, and 9 that are on both lists changed their score by at least 0.05 points or their cost per task by at least 1%. Sonnet 5.5 at max and at xhigh moved by less than that, so we don't count them.
The new rows include Gemini 4 Argon. It scores 52.6 for $1.99 a task, 0.7 points above GPT-6.1 Sol at max for 2.7× its cost. Opus 5.5 at high scores 1.0 points more than Argon and costs 8% less, so Argon arrived off the line. The other new rows are Sonnet 5.5 at low, Grok 4.7 at low and Solar Mini 4, none of them on the line either.
The rows that left are all 6 settings of GPT-6 Sol, which the page now marks as deprecated. None of them was on the line.
GPT-6 Luna's scores rose at every reasoning setting, from +0.5 points at medium to +0.9 points at max, with its costs about where they were. That rise is what put Luna at max on the line, ahead of MiMo-V2.6-Flash, so the line has 15 rows where it had 14. No row left the line.
What this doesn't tell you
It is one aggregate score. A model off this line can still be the right one for a particular job, and the ranking can change test by test. Our GPT-6.1 Sol video goes through Sol and Opus 5.5 test by test, and our Gemini 4 Argon video does the same for Argon.
We wouldn't choose between two rows on a gap of about a point, which is the size of the gap between Argon and Opus at high. What separates them in this comparison is the cost, and for Argon that is an introductory price: Artificial Analysis prices it at Google's $2 and $10 per million input and output tokens, and Google says $4 and $20 will apply after the introductory period. With every part of the price doubled, Argon would cost 2.2× as much per task as Opus at high. Google doesn't say what cached input will cost then, and our Gemini 4 Argon video works through that.
The cost is Artificial Analysis's cost per task on its own tests, in dollars. On a Claude or ChatGPT subscription you pay in quota instead, and this list doesn't measure that.
An older GPT-5.6 Luna row would sit on the line if we had counted rows the page marks as deprecated. We left it out because the question is what to pick today.
Artificial Analysis updates its figures, and Luna's rise between our two reads shows how a few days can move a row on or off the line. Every figure here is dated in the numbers table below.