model fatıgue
First look

Where GPT-6.1 Sol beats Opus 5.5, and where Opus is still well ahead

Published 30 Sep 2026Our test: one painting job

Pressing play loads the video from YouTube (Google), in privacy-enhanced mode. Privacy

This is the video written out, with every figure in full, each linked to its source in the table below.

A new Sol, a week after the last one▶ 0:00

OpenAI released GPT-6 Sol on 22 September. Seven days later, on 29 September, it released GPT-6.1 Sol, which its launch post calls an upgrade to GPT-6 Sol.

OpenAI pitches it against its own top model, GPT-6 Astra, as "near-Astra intelligence for a fifth of the price". For a lot of people, though, the model to compare it with is Anthropic's Opus 5.5. On Artificial Analysis's Intelligence Index, an independent score averaged over ten tests, the two score about the same at their lowest effort setting: 42.1 for Sol and 42.3 for Opus. Sol gets there for $0.13 a task against $0.55 for Opus, which is 0.24 of the cost, about a quarter. So what happens when you turn them both up?

The basics▶ 0:44

GPT-6.1 Sol costs the same per token as GPT-6 Sol for input and output, $2 per million input tokens and $10 per million output tokens, and cached input drops from $0.20 to $0.10 per million. Each of those is half what Opus 5.5 charges.

Price per million tokens

API list prices

GPT-6.1 SolGPT-6 SolOpus 5.5
input$2$2$4
cached input$0.10$0.20$0.20
output$10$10$20
OpenAI's launch posts for GPT-6.1 Sol and GPT-6 Sol, and Anthropic's Opus 5.5 page. GPT-6 Sol's cached price is our arithmetic from OpenAI's line that GPT-6.1 Sol's is 50% less. For prompts over 272K input tokens, GPT-6.1 Sol charges 2× the input and cached-input rates and 1.5× the output rate, according to OpenAI's model page.

You can use it through the API, and in Codex and ChatGPT Work on the Plus, Pro, Business, Enterprise and Edu plans. OpenAI says it isn't yet available in the regular ChatGPT chat. Its context window is 1,050,000 tokens, and it has five effort settings, low, medium, high, xhigh and max, with medium as the API's default.

On Artificial Analysis's main list, which puts each model at its top setting, four models score higher than GPT-6.1 Sol: Opus 5.5, Sonnet 5.5, Fable 5.1 and GPT-6 Astra. Sol costs less per task than any of them. Against the cheapest of the four, Astra at $3.26, Sol's $0.72 is 0.22 of the cost.

The top five on Artificial Analysis's main list

Each model at its top setting

modelIntelligence Indexcost per task
Opus 5.557.6$5.98
Sonnet 5.556.0$7.60
Fable 5.153.4$7.63
GPT-6 Astra52.7$3.26
GPT-6.1 Sol51.8$0.72
Artificial Analysis's main list, read 30 September 2026. Cost per task at API list prices.

If you use GPT-6 Sol, this is the one to try first. At every setting it scores four to eight points higher, from +4.3 points at max to +8.2 points at low, for the same money or less.

GPT-6.1 Sol against GPT-6 Sol

Intelligence Index score and cost per task at each setting

settingGPT-6.1 SolGPT-6 Soldifferencecost, GPT-6.1 Sol against GPT-6 Sol
low42.1 at $0.1333.9 at $0.13+8.20.99
medium47.8 at $0.2139.8 at $0.25+8.00.86
high50.2 at $0.3242.8 at $0.38+7.40.85
xhigh51.0 at $0.3944.1 at $0.52+6.90.75
max51.8 at $0.7247.5 at $1.05+4.30.69
Artificial Analysis. GPT-6 Sol's rows are from its own release page, saved on 29 September. Differences and cost ratios are our arithmetic.

Every setting against Opus 5.5▶ 1:52

Every setting against Opus 5.5

Intelligence Index score against cost per task, each model at its five effort settings. The yellow lines join the two pairs the text compares.

GPT-6.1 SolOpus 5.5
405060$0.10$0.30$1$3$10GPT-6.1 Sol at low: score 42.1, $0.13 per task · Artificial Analysis, read 30 Sep 2026, 02:16 CESTlowGPT-6.1 Sol at medium: score 47.8, $0.21 per task · Artificial Analysis, read 30 Sep 2026, 02:16 CESTmediumGPT-6.1 Sol at high: score 50.2, $0.32 per task · Artificial Analysis, read 30 Sep 2026, 02:16 CESThighGPT-6.1 Sol at xhigh: score 51.0, $0.39 per task · Artificial Analysis, read 30 Sep 2026, 02:16 CESTxhighGPT-6.1 Sol at max: score 51.8, $0.72 per task · Artificial Analysis, read 30 Sep 2026, 02:16 CESTmaxGPT-6.1 SolOpus 5.5 at low: score 42.3, $0.55 per task · Artificial Analysis, read 29 Sep 2026, 12:28 CESTOpus 5.5 at medium: score 51.2, $1.34 per task · Artificial Analysis, read 30 Sep 2026, 02:16 CESTOpus 5.5 at high: score 53.6, $1.82 per task · Artificial Analysis, read 30 Sep 2026, 02:16 CESTOpus 5.5 at xhigh: score 56.0, $3.46 per task · Artificial Analysis, read 30 Sep 2026, 02:16 CESTOpus 5.5 at max: score 57.6, $5.98 per task · Artificial Analysis, read 30 Sep 2026, 02:16 CESTOpus 5.5scorecost per task, log scale
405060$0.10$0.30$1$3$10GPT-6.1 Sol at low: score 42.1, $0.13 per task · Artificial Analysis, read 30 Sep 2026, 02:16 CESTGPT-6.1 Sol at medium: score 47.8, $0.21 per task · Artificial Analysis, read 30 Sep 2026, 02:16 CESTGPT-6.1 Sol at high: score 50.2, $0.32 per task · Artificial Analysis, read 30 Sep 2026, 02:16 CESTGPT-6.1 Sol at xhigh: score 51.0, $0.39 per task · Artificial Analysis, read 30 Sep 2026, 02:16 CESTGPT-6.1 Sol at max: score 51.8, $0.72 per task · Artificial Analysis, read 30 Sep 2026, 02:16 CESTGPT-6.1 SolOpus 5.5 at low: score 42.3, $0.55 per task · Artificial Analysis, read 29 Sep 2026, 12:28 CESTOpus 5.5 at medium: score 51.2, $1.34 per task · Artificial Analysis, read 30 Sep 2026, 02:16 CESTOpus 5.5 at high: score 53.6, $1.82 per task · Artificial Analysis, read 30 Sep 2026, 02:16 CESTOpus 5.5 at xhigh: score 56.0, $3.46 per task · Artificial Analysis, read 30 Sep 2026, 02:16 CESTOpus 5.5 at max: score 57.6, $5.98 per task · Artificial Analysis, read 30 Sep 2026, 02:16 CESTOpus 5.5scorecost per task, log scale
Artificial Analysis Intelligence Index v4.3.2, average per task, API list prices including caching. Read 30 September 2026 (Opus 5.5 at low: 29 September).
The numbers in this chart
settingGPT-6.1 Sol scoreGPT-6.1 Sol cost per index taskOpus 5.5 scoreOpus 5.5 cost per index task
low42.1$0.1342.3$0.55
medium47.8$0.2151.2$1.34
high50.2$0.3253.6$1.82
xhigh51.0$0.3956.0$3.46
max51.8$0.7257.6$5.98

At low, the two are level. At medium, Opus pulls ahead, 51.2 against 47.8. The closer match is Sol at high against Opus at medium. Sol scores 50.2, 1.0 points lower, and costs $0.32 a task against $1.34, which is 0.24 of the cost.

Part of that gap is the price per token, which is half. The rest is how much each model writes and reads. At that pair, Sol writes 13.2K output tokens per task against 25.7K for Opus, 0.51 as many. It also reads about 541K tokens from cache per task against about 2.05M, 0.26 as many. We worked those two cache figures out from Artificial Analysis's cost split and each model's cache price.

Turn Sol all the way up to max and it edges past Opus at medium, 51.8 against 51.2, for 0.54 of the cost. But that's as high as Sol goes. Opus keeps climbing, to 56.0 at xhigh and 57.6 at max. So a quarter of the price gets you to about Opus at medium, and not beyond it.

For Sol, high is where turning it up stops paying. Going from high to max adds +1.6 points for 2.27× the cost per task.

Sol at high is slower to start answering than Opus at medium. On Artificial Analysis's speed test, Sol at high takes 57.6 snow 58.0 s to its first answer token, against 21.9 snow 23.0 s for Opus at medium. On the index tasks themselves, the two take about the same time on average, 205 snow 206 s and 216 snow 219 s. These speed figures are rolling measurements that move from day to day, so where a re-read found a new value, it's shown next to the one from the video.

Where Sol comes out ahead▶ 3:15

The index combines ten tests, and that hides where the two differ. Split by the kind of work, Sol leads on two kinds at most settings, and Opus leads on most of the rest.

Answering questions from long documents full of tables and fine print is Sol's best case. On GDP.pdf, a document test that Artificial Analysis runs, Sol is ahead of Opus at every setting from medium up. Its best score is 32.0%, at high, against 28.8% for Opus at its best, also at high. OpenAI's launch chart for GDP.pdf shows the same values.

The other is business workflows with tools, on AutomationBench as Artificial Analysis runs it. Sol is slightly ahead at medium, high and xhigh, and only at max does Opus pull ahead, 69.5% against Sol's best of 66.6%, which Sol reached at xhigh.

Artificial Analysis's release page has no test scores for Opus at low, so the table starts at medium. Its Sonnet 5.5 page, read on 29 September, has two: at low, Opus scores 31.3% on Terminal-Bench 4.0 against 30.8% for Sol, and 1224 on GDPval-AA against 1297, so at low Sol leads on GDPval-AA.

By kind of work

Scores on the tests inside the index, GPT-6.1 Sol and Opus 5.5 at each setting from medium up

testmodelmediumhighxhighmax
GDP.pdf (questions over long documents)GPT-6.1 Sol30.0%32.0%31.8%31.0%
Opus 5.525.6%28.8%26.6%26.2%
AutomationBench (business workflows with tools)GPT-6.1 Sol62.6%64.5%66.6%64.9%
Opus 5.561.2%63.2%65.0%69.5%
Terminal-Bench 4.0 (work in a terminal)GPT-6.1 Sol48.0%51.5%54.0%56.1%
Opus 5.552.5%56.6%59.6%59.6%
GDPval-AA, Elo (finished work, judged by an AI judge)GPT-6.1 Sol1433148615101575
Opus 5.51576169218201846
AA-Briefcase, Elo (multi-week agent projects)GPT-6.1 Sol1365147115071564
Opus 5.51642170417801822
Humanity's Last Exam (very hard questions)GPT-6.1 Sol49.9%51.4%52.6%52.9%
Opus 5.554.7%55.6%57.5%61.4%
CritPt (research physics)GPT-6.1 Sol27.7%30.0%31.7%31.7%
Opus 5.527.7%30.9%31.7%31.7%
Artificial Analysis's release page, read 30 September 2026. It has no test scores for Opus 5.5 at low.

Where Opus stays ahead▶ 4:03

Opus leads on most of the other tests. On work in a terminal, Terminal-Bench 4.0, it's ahead at every setting from medium up, 59.6% against 56.1% at max. On GDPval-AA, where an AI judge compares finished work like documents, spreadsheets and slides, Opus is well ahead from medium up (at low, Sol is ahead), and Opus at medium, 1576, already matches the best Sol does, 1575 at max. Opus is also well ahead on AA-Briefcase, Artificial Analysis's test of agent work on multi-week business projects, 1822 against 1564 at max.

On Humanity's Last Exam, a set of very hard questions, Opus at max is 8.4 points ahead of Sol at max. On CritPt, a research physics test, they tie at max, 31.7% each.

At max, test by test

Score on each test inside the index, both models at max

Opus 5.5GPT-6.1 Sol
scoreSol minus OpusGDP.pdfGDP.pdf, Opus 5.5: 26.2% · Artificial Analysis, read 30 Sep 2026, 02:16 CEST26.2%GDP.pdf, GPT-6.1 Sol: 31.0% · Artificial Analysis, read 30 Sep 2026, 02:16 CEST31.0%+4.8AutomationBenchAutomationBench, Opus 5.5: 69.5% · Artificial Analysis, read 30 Sep 2026, 02:16 CEST69.5%AutomationBench, GPT-6.1 Sol: 64.9% · Artificial Analysis, read 30 Sep 2026, 02:16 CEST64.9%−4.7Terminal-Bench 4.0Terminal-Bench 4.0, Opus 5.5: 59.6% · Artificial Analysis, read 30 Sep 2026, 02:16 CEST59.6%Terminal-Bench 4.0, GPT-6.1 Sol: 56.1% · Artificial Analysis, read 30 Sep 2026, 02:16 CEST56.1%−3.5Humanity's Last ExamHumanity's Last Exam, Opus 5.5: 61.4% · Artificial Analysis, read 30 Sep 2026, 02:16 CEST61.4%Humanity's Last Exam, GPT-6.1 Sol: 52.9% · Artificial Analysis, read 30 Sep 2026, 02:16 CEST52.9%−8.4CritPtCritPt, Opus 5.5: 31.7% · Artificial Analysis, read 30 Sep 2026, 02:16 CEST31.7%CritPt, GPT-6.1 Sol: 31.7% · Artificial Analysis, read 30 Sep 2026, 02:16 CEST31.7%0.0
scoreSol minus OpusGDP.pdfGDP.pdf, Opus 5.5: 26.2% · Artificial Analysis, read 30 Sep 2026, 02:16 CEST26.2%GDP.pdf, GPT-6.1 Sol: 31.0% · Artificial Analysis, read 30 Sep 2026, 02:16 CEST31.0%+4.8AutomationBenchAutomationBench, Opus 5.5: 69.5% · Artificial Analysis, read 30 Sep 2026, 02:16 CEST69.5%AutomationBench, GPT-6.1 Sol: 64.9% · Artificial Analysis, read 30 Sep 2026, 02:16 CEST64.9%−4.7Terminal-Bench 4.0Terminal-Bench 4.0, Opus 5.5: 59.6% · Artificial Analysis, read 30 Sep 2026, 02:16 CEST59.6%Terminal-Bench 4.0, GPT-6.1 Sol: 56.1% · Artificial Analysis, read 30 Sep 2026, 02:16 CEST56.1%−3.5Humanity's Last ExamHumanity's Last Exam, Opus 5.5: 61.4% · Artificial Analysis, read 30 Sep 2026, 02:16 CEST61.4%Humanity's Last Exam, GPT-6.1 Sol: 52.9% · Artificial Analysis, read 30 Sep 2026, 02:16 CEST52.9%−8.4CritPtCritPt, Opus 5.5: 31.7% · Artificial Analysis, read 30 Sep 2026, 02:16 CEST31.7%CritPt, GPT-6.1 Sol: 31.7% · Artificial Analysis, read 30 Sep 2026, 02:16 CEST31.7%0.0
Artificial Analysis, read 30 September 2026. Differences in percentage points, our arithmetic. The two Elo tests are in the table above.
The numbers in this chart
Opus 5.5GPT-6.1 SolSol minus Opus
GDP.pdf26.2%31.0%+4.8 points
AutomationBench69.5%64.9%−4.7 points
Terminal-Bench 4.059.6%56.1%−3.5 points
Humanity's Last Exam61.4%52.9%−8.4 points
CritPt31.7%31.7%0.0 points

OpenAI's own launch charts agree at the top. On its Terminal-Bench Science chart, a test of scientific work in a terminal, OpenAI plots Opus at max at 63.3%, higher than GPT-6.1 Sol at any setting. Sol's best there is 57.0%, at max. OpenAI says it took its competitor points from "publicly available reports", and Anthropic's own figure for Opus 5.5 on that test is 58.7%, which is also above Sol's best.

OpenAI's Terminal-Bench Science chart

Score at each setting, from the values embedded in OpenAI's launch post

lowmediumhighxhighmax
GPT-6.1 Sol43.7%47.6%51.1%53.7%57.0%
Opus 5.5, as OpenAI plotted it63.3%
Opus 5.5, Anthropic's own figure58.7%
OpenAI's runs of GPT-6.1 Sol. OpenAI says its competitor points were "taken from publicly available reports". At max, OpenAI puts GPT-6.1 Sol's cost at $5.47 a task and Opus 5.5's at $23.21. Anthropic's figure is from its Opus 5.5 page.

In a coding agent▶ 4:54

One more measurement arrived after launch, on coding agents. Artificial Analysis now runs whole agents, each with its own model, through three sets of coding tasks. Claude Code with Opus 5.5 at max scores 66.0, at $13.04 a task, and takes 64 min on each. Codex with GPT-6.1 Sol at xhigh scores 62.9, at $1.04 a task, 0.08 of the cost or about a twelfth, in 16 min. That also puts it ahead of Codex with GPT-6 Astra at max, 61.6, at 0.14 of Astra's cost.

Opus 5.5 has only been run at max so far, so there's no cheaper Opus setting to set against it yet. The top of the list, for now, is Claude Code with Sonnet 5.5 at max, at 68.4.

Artificial Analysis's Coding Agent Index

Each agent running one model at one setting, on DeepSWE v1.1, SWE-Atlas-QnA and Terminal-Bench v4

agent and modelindexcost per tasktime per task
Claude Code, Sonnet 5.5 at max68.4$14.1987 min
Claude Code, Opus 5.5 at max66.0$13.0464 min
Codex, GPT-6.1 Sol at xhigh62.9$1.0416 min
Claude Code, Fable 5.1 at max62.2$12.3935 min
Codex, GPT-6 Astra at max61.6$7.4729 min
Codex, GPT-6.1 Sol at medium61.4$0.7011 min
Codex, GPT-6.1 Sol at high60.1$0.8913 min
Codex, GPT-6.1 Sol at max60.1$1.5524 min
Codex, GPT-6.1 Sol at low57.2$0.509 min
Codex, GPT-6 Sol at max56.7$2.9922 min
Artificial Analysis, read 30 September 2026. Other entries on its list (other settings and other agents) are left out. Cost per task at API list prices; time is the agent's wall time per task.

Our test: the Great Wave▶ 5:33

Each model gets the same job: look at Hokusai's Great Wave and paint it by writing its own painting program, with no image model and no image library. We ran it once per model at each tool's default setting, GPT-6.1 Sol, GPT-6 Sol and GPT-6 Astra in Codex, and Opus 5.5 in Claude Code.

Four models, one prompt, the Great Wave

Hokusai's print, the Great Wave
The printKatsushika Hokusai, Under the Wave off Kanagawa (The Great Wave), c. 1830–32. The Metropolitan Museum of Art, H. O. Havemeyer Collection, Bequest of Mrs. H. O. Havemeyer, 1929 (JP1847). Public domain.
GPT-6.1 Sol's painting of the Great Wave
GPT-6.1 SolCodex · default · 1 run
$0.15 at API prices
Opus 5.5's painting of the Great Wave
Opus 5.5Claude Code · default · 1 run
$1.03 at API prices
GPT-6 Sol's painting of the Great Wave
GPT-6 SolCodex · default · 1 run
$0.25 at API prices
GPT-6 Astra's painting of the Great Wave
GPT-6 AstraCodex · default · 1 run
$1.60 at API prices
Our runs, one per model, on 28 September 2026 and, for GPT-6.1 Sol, 29 September 2026. Costs are the runs' tokens at API list prices; for Opus 5.5 it is Claude Code's own list-price figure.

To our eye, GPT-6.1 Sol's wave has what the print has, the boats, Mount Fuji and the title box, and it's closer to the print than GPT-6 Sol's. Astra's is still the closest. Opus's is more stylized, with its own colours and shapes. Priced at API rates, Sol's painting cost $0.15 and Opus's $1.03, so on this one job Sol cost 0.15 of what Opus did, a smaller share than the quarter on the index.

We also ran Sol's painting at low, high, xhigh and max, as well as the default. The high and max runs took 22 min and 29 min and came to 17.3 and 18.8 Codex credits at the rate card's prices, against 3.8 credits at the default, for a fuller painting closer to the print.

GPT-6.1 Sol's painting at four of its settings

GPT-6.1 Sol's painting of the Great Wave at low
GPT-6.1 Sollow · 1 run
5 min · 3.2 credits
GPT-6.1 Sol's painting of the Great Wave at Codex's default setting
GPT-6.1 Soldefault · 1 run
5 min · 3.8 credits
GPT-6.1 Sol's painting of the Great Wave at high
GPT-6.1 Solhigh · 1 run
22 min · 17.3 credits
GPT-6.1 Sol's painting of the Great Wave at max
GPT-6.1 Solmax · 1 run
29 min · 18.8 credits
Our runs in Codex, 29 September 2026, one at each of low, high and max and one at the default. Credits are the runs' tokens priced at OpenAI's Codex rate card, our arithmetic. The times are with nine runs going at once. The xhigh run (20 min, 14.4 credits) isn't shown.

The xhigh painting isn't shown. Its program has more numbers written into it than our test allows, a check that is there to catch a model copying image data. Reading the program, we found hand-placed curves rather than an encoded image, but we kept it out, as the video does.

These are single runs, so they show what each model can do, and nothing about how often it does it.

Against Astra and Fable▶ 6:44

Against GPT-6 Astra, OpenAI's own top model, Sol at high scores 50.2, within a point of Astra at high, 50.9, for 0.18 of the cost per task, about a fifth. At their top settings Astra stays ahead, 52.7 against 51.8.

On OpenAI's credit rate card for Codex, an Astra token costs 5× as many credits as a Sol token, and 10× as many on cached input, which is most of what a coding agent reads. In our GPT-6.1 Sol Great Wave run, for example, Sol read 168,576 tokens from cache and 25,427 fresh ones, and wrote 8,483.

OpenAI's Codex credit rate card

Credits per million tokens, on credit-based plans

inputcached inputoutput
GPT-6.1 Sol502.5250
GPT-6 Sol505250
GPT-6 Astra250251,250
OpenAI Help Center, read 29 September 2026.

Against Anthropic's Fable 5.1, Sol at max, 51.8, matches Fable at high, 51.2, for 0.19 of the cost, about a fifth. Fable at max goes on to 53.4.

Against Astra and Fable

Intelligence Index score against cost per task, four models at their five settings. The yellow lines join the two pairs the text compares.

GPT-6.1 SolGPT-6 AstraOpus 5.5Fable 5.1
405060$0.10$0.30$1$3$10GPT-6.1 Sol at low: score 42.1, $0.13 per task · Artificial Analysis, read 30 Sep 2026, 02:16 CESTlowGPT-6.1 Sol at medium: score 47.8, $0.21 per task · Artificial Analysis, read 30 Sep 2026, 02:16 CESTmediumGPT-6.1 Sol at high: score 50.2, $0.32 per task · Artificial Analysis, read 30 Sep 2026, 02:16 CESThighGPT-6.1 Sol at xhigh: score 51.0, $0.39 per task · Artificial Analysis, read 30 Sep 2026, 02:16 CESTxhighGPT-6.1 Sol at max: score 51.8, $0.72 per task · Artificial Analysis, read 30 Sep 2026, 02:16 CESTmaxGPT-6.1 SolGPT-6 Astra at low: score 45.8, $0.82 per task · Artificial Analysis, read 29 Sep 2026, 12:28 CESTGPT-6 Astra at medium: score 49.6, $1.54 per task · Artificial Analysis, read 30 Sep 2026, 02:16 CESTGPT-6 Astra at high: score 50.9, $1.73 per task · Artificial Analysis, read 30 Sep 2026, 02:16 CESTGPT-6 Astra at xhigh: score 52.4, $2.31 per task · Artificial Analysis, read 30 Sep 2026, 02:16 CESTGPT-6 Astra at max: score 52.7, $3.26 per task · Artificial Analysis, read 30 Sep 2026, 02:16 CESTGPT-6 AstraOpus 5.5 at low: score 42.3, $0.55 per task · Artificial Analysis, read 29 Sep 2026, 12:28 CESTOpus 5.5 at medium: score 51.2, $1.34 per task · Artificial Analysis, read 30 Sep 2026, 02:16 CESTOpus 5.5 at high: score 53.6, $1.82 per task · Artificial Analysis, read 30 Sep 2026, 02:16 CESTOpus 5.5 at xhigh: score 56.0, $3.46 per task · Artificial Analysis, read 30 Sep 2026, 02:16 CESTOpus 5.5 at max: score 57.6, $5.98 per task · Artificial Analysis, read 30 Sep 2026, 02:16 CESTOpus 5.5Fable 5.1 at low: score 46.8, $2.37 per task · Artificial Analysis, read 29 Sep 2026, 22:10 CESTFable 5.1 at medium: score 48.9, $2.98 per task · Artificial Analysis, read 30 Sep 2026, 02:16 CESTFable 5.1 at high: score 51.2, $3.91 per task · Artificial Analysis, read 30 Sep 2026, 02:16 CESTFable 5.1 at xhigh: score 53.2, $5.98 per task · Artificial Analysis, read 30 Sep 2026, 02:16 CESTFable 5.1 at max: score 53.4, $7.63 per task · Artificial Analysis, read 30 Sep 2026, 02:16 CESTFable 5.1scorecost per task, log scale
405060$0.10$0.30$1$3$10GPT-6.1 Sol at low: score 42.1, $0.13 per task · Artificial Analysis, read 30 Sep 2026, 02:16 CESTGPT-6.1 Sol at medium: score 47.8, $0.21 per task · Artificial Analysis, read 30 Sep 2026, 02:16 CESTGPT-6.1 Sol at high: score 50.2, $0.32 per task · Artificial Analysis, read 30 Sep 2026, 02:16 CESTGPT-6.1 Sol at xhigh: score 51.0, $0.39 per task · Artificial Analysis, read 30 Sep 2026, 02:16 CESTGPT-6.1 Sol at max: score 51.8, $0.72 per task · Artificial Analysis, read 30 Sep 2026, 02:16 CESTGPT-6.1 SolGPT-6 Astra at low: score 45.8, $0.82 per task · Artificial Analysis, read 29 Sep 2026, 12:28 CESTGPT-6 Astra at medium: score 49.6, $1.54 per task · Artificial Analysis, read 30 Sep 2026, 02:16 CESTGPT-6 Astra at high: score 50.9, $1.73 per task · Artificial Analysis, read 30 Sep 2026, 02:16 CESTGPT-6 Astra at xhigh: score 52.4, $2.31 per task · Artificial Analysis, read 30 Sep 2026, 02:16 CESTGPT-6 Astra at max: score 52.7, $3.26 per task · Artificial Analysis, read 30 Sep 2026, 02:16 CESTGPT-6 AstraOpus 5.5 at low: score 42.3, $0.55 per task · Artificial Analysis, read 29 Sep 2026, 12:28 CESTOpus 5.5 at medium: score 51.2, $1.34 per task · Artificial Analysis, read 30 Sep 2026, 02:16 CESTOpus 5.5 at high: score 53.6, $1.82 per task · Artificial Analysis, read 30 Sep 2026, 02:16 CESTOpus 5.5 at xhigh: score 56.0, $3.46 per task · Artificial Analysis, read 30 Sep 2026, 02:16 CESTOpus 5.5 at max: score 57.6, $5.98 per task · Artificial Analysis, read 30 Sep 2026, 02:16 CESTOpus 5.5Fable 5.1 at low: score 46.8, $2.37 per task · Artificial Analysis, read 29 Sep 2026, 22:10 CESTFable 5.1 at medium: score 48.9, $2.98 per task · Artificial Analysis, read 30 Sep 2026, 02:16 CESTFable 5.1 at high: score 51.2, $3.91 per task · Artificial Analysis, read 30 Sep 2026, 02:16 CESTFable 5.1 at xhigh: score 53.2, $5.98 per task · Artificial Analysis, read 30 Sep 2026, 02:16 CESTFable 5.1 at max: score 53.4, $7.63 per task · Artificial Analysis, read 30 Sep 2026, 02:16 CESTFable 5.1scorecost per task, log scale
Artificial Analysis Intelligence Index v4.3.2, API list prices including caching. Read 30 September 2026 (GPT-6 Astra and Opus 5.5 at low: 29 September; Fable 5.1 at low: launch night, 29 September).
The numbers in this chart
settingGPT-6.1 Sol scoreGPT-6.1 Sol cost per index taskGPT-6 Astra scoreGPT-6 Astra cost per index taskOpus 5.5 scoreOpus 5.5 cost per index taskFable 5.1 scoreFable 5.1 cost per index task
low42.1$0.1345.8$0.8242.3$0.5546.8$2.37
medium47.8$0.2149.6$1.5451.2$1.3448.9$2.98
high50.2$0.3250.9$1.7353.6$1.8251.2$3.91
xhigh51.0$0.3952.4$2.3156.0$3.4653.2$5.98
max51.8$0.7252.7$3.2657.6$5.9853.4$7.63

The system card▶ 7:19

OpenAI's system card for GPT-6.1 Sol is an addendum to GPT-6 Astra's. OpenAI treats Sol as Critical in cybersecurity capability and High in biology and chemistry, the same as Astra, so it gets Astra's safeguards, and advanced security work goes through a separate access program.

On OpenAI's coding-deception test, Sol misrepresented its work in 1.50% of cases, a little more often than GPT-6 Sol at 1.30% and about 2.9× as often as Astra at 0.51%. In a test where a tool is blocked with a warning, Sol tried to get around the block in 23.5% of runs, against 17.4% for Astra. OpenAI's launch chart for the same test, all at max, puts GPT-6 Sol at 64.4%. And in OpenAI's simulation of 49,650 internal Codex tasks, GPT-6.1 Sol drew 28 flags at OpenAI's most serious levels, for misaligned behaviour a reasonable user would strongly object to. GPT-6 Sol drew 42, so that's a third fewer (33% by OpenAI's count), and Astra drew 27.

Three of OpenAI's own evaluations

From the system card addendum for GPT-6.1 Sol

evaluationGPT-6.1 SolGPT-6 SolGPT-6 Astra
coding deception, misrepresentation rate1.50%1.30%0.51%
kept going past a warning (a blocked tool)23.5%64.4%17.4%
the most serious flags in a simulation of internal Codex tasks284227
OpenAI, Addendum to GPT-6 Astra System Card: GPT-6.1 Sol, published 29 September 2026. The flags are those OpenAI rates at severity three or higher, over 49,650 simulated tasks. GPT-6 Sol's rate on the persistence test is from OpenAI's launch chart for the same test, at max.

These are all OpenAI's own tests, and they don't say how often any of it happens in normal use.

What it costs you▶ 8:17

If you pay per token, that quarter is what you'd see at matching scores, and it comes from cheaper tokens and fewer of them. Take Sol at high and Opus at medium, $0.32 against $1.34 a task. At Sol's prices per token, Opus's task would cost $0.67. The rest of the gap is Sol using fewer tokens, weighted by what each kind of token costs.

OpenAI's model page adds some small print. With Sol, a prompt over 272K input tokens costs 2× as much for input and cached input and 1.5× as much for output, for the whole request.

If you pay with a plan, the two aren't on the same meter. Opus draws on a Claude plan and Sol on a ChatGPT one, so there's no common unit to compare them in. OpenAI estimates 15–160 Codex messages per five hours on the Plus plan with Sol, against 5–45 with Astra. When we checked on 30 September, we hadn't found a published measurement of how fast Sol actually uses up a plan.

Should you switch?▶ 9:01

If you're on GPT-6 Sol, try this one first on your own work. If you use Astra in Codex, try Sol. And if you're on Opus 5.5, Sol is worth trying on your own tasks for questions over documents, for tool workflows where Opus at medium is good enough, and for coding agents where a few points matter less than the bill. On those three, Artificial Analysis's figures put Sol at less than half the API cost of Opus.

Cost per task on single tests

GPT-6.1 Sol against Opus 5.5, as Artificial Analysis prices each test

testGPT-6.1 SolOpus 5.5Sol's share
GDP.pdf, Sol at high, Opus at medium$0.35$0.800.44
GDP.pdf, both at high$0.35$0.830.42
GDP.pdf, both at medium$0.34$0.800.42
AutomationBench, Sol at high, Opus at medium$0.23$0.640.35
AutomationBench, both at medium$0.20$0.640.31
Coding Agent Index, Sol at xhigh, Opus at max$1.04$13.040.08
Artificial Analysis's release page and Coding Agent Index, read 30 September 2026, at API list prices. The shares are our arithmetic.

For finished work that someone will judge, long projects and the hardest questions, stay on Opus.

Hold all of this loosely, though. Artificial Analysis is still the only independent measurement of GPT-6.1 Sol we've found. When we looked on 30 September, LMArena was still collecting votes and showed no rating, Vals and METR didn't list it, and the SWE-bench and LiveBench pages we saved didn't load their tables, so we can't say either way for those two. GPT-6 Sol scored well on Artificial Analysis too, then spent its week on r/codex being called slow and not very good.

GPT-6 Sol's week on r/codex

Post titles about GPT-6 Sol, the model before this one

r/codex, points as Reddit showed them on 30 September 2026. These are about GPT-6 Sol, not GPT-6.1 Sol.

What would settle it is how fast Sol drains a plan, how it does on your own code, and whether people still like it in a week.

Just before this video, we looked at Fireworks' Ember-1, a retrained Kimi K3 that promises the same work with fewer tokens. That's our previous video.

Every number

These are all 329 figures behind the video and this page, grouped by whose they are, with the page each came from and when we read it. Figures marked ⟳ can move. When a re-read finds a change, the new value shows next to the one from the video.

Artificial Analysis Intelligence Index, by effort setting

Artificial Analysis · Artificial Analysis, GPT-6.1 Sol release page (every effort setting, with the models it's compared with) · read 30 Sep 2026, 02:16 CEST · ⟳ re-read 30 Sep 2026, 14:22 CEST, unchanged · Intelligence Index v4.3.2, average per task; cost at API list prices including caching. · Some cells come from other saved Artificial Analysis pages and were not re-read on 30 September: GPT-6 Sol at every setting and Fable 5.1 at low (launch night, 29 Sep 22:10 to 22:45 CEST), Opus 5.5 and GPT-6 Astra at low (its Sonnet 5.5 page, 29 Sep 12:28 CEST). The downloads (numbers.json, numbers.csv) give each cell's source.

Intelligence index score

lowmediumhighxhighmax
GPT-6.1 Sol42.147.850.251.051.8
Opus 5.542.351.253.656.057.6
GPT-6 Astra45.849.650.952.452.7
Fable 5.146.848.951.253.253.4
GPT-6 Sol33.939.842.844.147.5

Cost per index task, list prices incl. caching

Output tokens per index task

Cost per index task spent on cache reads

lowmediumhighxhighmax
GPT-6.1 Sol$0.016$0.032$0.054$0.068$0.123
Opus 5.5–$0.410$0.604$1.369$2.424

Gdp.pdf score (questions over long documents)

lowmediumhighxhighmax
GPT-6.1 Sol27.0%30.0%32.0%31.8%31.0%
Opus 5.5–25.6%28.8%26.6%26.2%

Automationbench score (business workflows with tools)

lowmediumhighxhighmax
GPT-6.1 Sol52.6%62.6%64.5%66.6%64.9%
Opus 5.5–61.2%63.2%65.0%69.5%

Terminal-bench 4.0 score (work in a terminal)

lowmediumhighxhighmax
GPT-6.1 Sol30.8%48.0%51.5%54.0%56.1%
Opus 5.5–52.5%56.6%59.6%59.6%

Gdpval-aa elo (finished work, judged by an ai judge)

lowmediumhighxhighmax
GPT-6.1 Sol12971433148615101575
Opus 5.5–1576169218201846

Aa-briefcase elo (multi-week agent projects)

lowmediumhighxhighmax
GPT-6.1 Sol11181365147115071564
Opus 5.5–1642170417801822

Humanity's last exam score

lowmediumhighxhighmax
GPT-6.1 Sol47.4%49.9%51.4%52.6%52.9%
Opus 5.5–54.7%55.6%57.5%61.4%

Critpt score (research physics)

lowmediumhighxhighmax
GPT-6.1 Sol24.9%27.7%30.0%31.7%31.7%
Opus 5.5–27.7%30.9%31.7%31.7%

OpenAI

OpenAI, Introducing GPT-6.1 Sol (launch post) · read 29 Sep 2026, 22:10 CEST
GPT-6.1 Sol, API price per million input tokens$2
GPT-6.1 Sol, API price per million cached input tokens$0.10
GPT-6.1 Sol, API price per million output tokens$10
GPT-6.1 Sol's cached input price against GPT-6 Sol's: this much less50%
GPT-6 Sol, API price per million cached input tokens, from OpenAI's '50% less' (our arithmetic)$0.20
Opus 5.5's input price against GPT-6.1 Sol's (our arithmetic)2×

OpenAI

GPT-6 Sol, API price per million input tokens (down from GPT-5.6 Sol's $4)$2
GPT-6 Sol, API price per million output tokens (down from GPT-5.6 Sol's $20)$10

Anthropic

Anthropic, Claude Opus 5.5 (launch page) · read 24 Sep 2026, 18:13 CEST
Opus 5.5, API price per million input tokens$4
Opus 5.5, API price per million output tokens$20
Opus 5.5, API price per million cache-read tokens$0.20
Terminal-Bench Science 0.1, Opus 5.5, Anthropic's own figure58.7%

OpenAI

OpenAI, GPT-6.1 Sol model page for developers · read 29 Sep 2026, 22:10 CEST
GPT-6.1 Sol's context window, tokens1,050,000
prompt length above which GPT-6.1 Sol's long-prompt rates apply, input tokens272K
long prompts: multiplier on input and cached input rates2×
long prompts: multiplier on output rate1.5×

Artificial Analysis

main list, place 1: Opus 5.5 at its top setting, Intelligence Index57.6 ⟳
main list, place 1: Opus 5.5 at its top setting, cost per task$5.98 ⟳
main list, place 2: Sonnet 5.5 at its top setting, Intelligence Index56.0 ⟳
main list, place 2: Sonnet 5.5 at its top setting, cost per task$7.60 ⟳
main list, place 3: Fable 5.1 at its top setting, Intelligence Index53.4 ⟳
main list, place 3: Fable 5.1 at its top setting, cost per task$7.63 ⟳
main list, place 4: GPT-6 Astra at its top setting, Intelligence Index52.7 ⟳
main list, place 4: GPT-6 Astra at its top setting, cost per task$3.26 ⟳
main list, place 5: GPT-6.1 Sol at its top setting, Intelligence Index51.8 ⟳
main list, place 5: GPT-6.1 Sol at its top setting, cost per task$0.72 ⟳
GPT-6.1 Sol's place on the main list5 ⟳
GPT-6.1 Sol's cost per task at max against GPT-6 Astra's, the cheapest of the four above it (our arithmetic)0.22

Artificial Analysis

Terminal-Bench 4.0 score, Opus 5.5 at low31.3%
GDPval-AA Elo, Opus 5.5 at low1224

Artificial Analysis

AA's speed test, time to the first answer token, GPT-6.1 Sol at high57.6 snow 58.0 s ⟳
AA's Intelligence Index, average time per task, GPT-6.1 Sol at high205 snow 206 s ⟳
AA's speed test, time to the first answer token, Opus 5.5 at medium21.9 snow 23.0 s ⟳
AA's Intelligence Index, average time per task, Opus 5.5 at medium216 snow 219 s ⟳
cost per task at low, GPT-6.1 Sol against Opus 5.5 (our arithmetic)0.24
index at low, GPT-6.1 Sol minus Opus 5.5 (our arithmetic)−0.2 points
index at medium, GPT-6.1 Sol minus Opus 5.5 (our arithmetic)−3.5 points
index, GPT-6.1 Sol at high minus Opus 5.5 at medium (our arithmetic)−1.0 points
cost per task, GPT-6.1 Sol at high against Opus 5.5 at medium (our arithmetic)0.24
index, GPT-6.1 Sol at max minus Opus 5.5 at medium (our arithmetic)+0.6 points
cost per task, GPT-6.1 Sol at max against Opus 5.5 at medium (our arithmetic)0.54
GPT-6.1 Sol, index gained from high to max (our arithmetic)+1.6 points
GPT-6.1 Sol, cost per task at max against high (our arithmetic)2.27×
output tokens per task, GPT-6.1 Sol at high against Opus 5.5 at medium (our arithmetic)0.51
cache-read tokens per task, GPT-6.1 Sol at high (AA's cache-read cost divided by the list price) (our arithmetic)541K
cache-read tokens per task, Opus 5.5 at medium (AA's cache-read cost divided by the list price) (our arithmetic)2.05M
cache-read tokens per task, GPT-6.1 Sol at high against Opus 5.5 at medium (our arithmetic)0.26
Opus 5.5 at medium's cost per task if every token cost what GPT-6.1 Sol's does (half) (our arithmetic)$0.67
index at low, GPT-6.1 Sol minus GPT-6 Sol (our arithmetic)+8.2 points
cost per task at low, GPT-6.1 Sol against GPT-6 Sol (our arithmetic)0.99
index at medium, GPT-6.1 Sol minus GPT-6 Sol (our arithmetic)+8.0 points
cost per task at medium, GPT-6.1 Sol against GPT-6 Sol (our arithmetic)0.86
index at high, GPT-6.1 Sol minus GPT-6 Sol (our arithmetic)+7.4 points
cost per task at high, GPT-6.1 Sol against GPT-6 Sol (our arithmetic)0.85
index at xhigh, GPT-6.1 Sol minus GPT-6 Sol (our arithmetic)+6.9 points
cost per task at xhigh, GPT-6.1 Sol against GPT-6 Sol (our arithmetic)0.75
index at max, GPT-6.1 Sol minus GPT-6 Sol (our arithmetic)+4.3 points
cost per task at max, GPT-6.1 Sol against GPT-6 Sol (our arithmetic)0.69
index at high, GPT-6.1 Sol minus GPT-6 Astra (our arithmetic)−0.7 points
cost per task at high, GPT-6.1 Sol against GPT-6 Astra (our arithmetic)0.18
cost per task, GPT-6.1 Sol at max against Fable 5.1 at high (our arithmetic)0.19
Humanity's Last Exam at max, Opus 5.5's lead over GPT-6.1 Sol (our arithmetic)8.4 points
GDP.pdf score at max, GPT-6.1 Sol minus Opus 5.5 (our arithmetic)+4.8 points
AutomationBench score at max, GPT-6.1 Sol minus Opus 5.5 (our arithmetic)−4.7 points
Terminal-Bench 4.0 score at max, GPT-6.1 Sol minus Opus 5.5 (our arithmetic)−3.5 points
Humanity's Last Exam score at max, GPT-6.1 Sol minus Opus 5.5 (our arithmetic)−8.4 points
CritPt score at max, GPT-6.1 Sol minus Opus 5.5 (our arithmetic)0.0 points
GDPval-AA Elo at max, GPT-6.1 Sol minus Opus 5.5 (Elo points) (our arithmetic)−271
AA-Briefcase Elo at max, GPT-6.1 Sol minus Opus 5.5 (Elo points) (our arithmetic)−258

Artificial Analysis

GDP.pdf score, cost per task, GPT-6.1 Sol at medium$0.34 ⟳
AutomationBench score, cost per task, GPT-6.1 Sol at medium$0.20 ⟳
GDP.pdf score, cost per task, GPT-6.1 Sol at high$0.35 ⟳
AutomationBench score, cost per task, GPT-6.1 Sol at high$0.23 ⟳
GDP.pdf score, cost per task, Opus 5.5 at medium$0.80 ⟳
AutomationBench score, cost per task, Opus 5.5 at medium$0.64 ⟳
GDP.pdf score, cost per task, Opus 5.5 at high$0.83 ⟳
AutomationBench score, cost per task, Opus 5.5 at high$0.70 ⟳
GDP.pdf score cost per task, GPT-6.1 Sol at high against Opus 5.5 at medium (our arithmetic)0.44
GDP.pdf score cost per task, GPT-6.1 Sol at medium against Opus 5.5 at medium (our arithmetic)0.42
AutomationBench score cost per task, GPT-6.1 Sol at high against Opus 5.5 at medium (our arithmetic)0.35
AutomationBench score cost per task, GPT-6.1 Sol at medium against Opus 5.5 at medium (our arithmetic)0.31
GDP.pdf cost per task, GPT-6.1 Sol at high against Opus 5.5 at high (our arithmetic)0.42

OpenAI

Terminal-Bench Science 0.1, Opus 5.5 at max, as OpenAI plotted it (from 'publicly available reports')63.3%
Terminal-Bench Science 0.1, Opus 5.5 at max, cost per task, as OpenAI plotted it$23.21
Terminal-Bench Science 0.1, GPT-6.1 Sol at low, as OpenAI plotted it43.7%
Terminal-Bench Science 0.1, GPT-6.1 Sol at medium, as OpenAI plotted it47.6%
Terminal-Bench Science 0.1, GPT-6.1 Sol at high, as OpenAI plotted it51.1%
Terminal-Bench Science 0.1, GPT-6.1 Sol at xhigh, as OpenAI plotted it53.7%
Terminal-Bench Science 0.1, GPT-6.1 Sol at max, as OpenAI plotted it57.0%
Terminal-Bench Science 0.1, GPT-6.1 Sol at max, cost per task, as OpenAI plotted it$5.47
a blocked tool with a warning: runs that tried to get around the block, GPT-6 Sol at max, from OpenAI's launch chart 'Warning circumvention' (the same test: its GPT-6.1 Sol and Astra values equal the system card's)64.4%

Artificial Analysis

Coding Agent Index, Claude Code with Sonnet 5.5 at max68.4 ⟳
Coding Agent Index, Claude Code with Sonnet 5.5 at max, cost per task$14.19 ⟳
Coding Agent Index, Claude Code with Sonnet 5.5 at max, agent time per task87 min ⟳
Coding Agent Index, Claude Code with Opus 5.5 at max66.0 ⟳
Coding Agent Index, Claude Code with Opus 5.5 at max, cost per task$13.04 ⟳
Coding Agent Index, Claude Code with Opus 5.5 at max, agent time per task64 min ⟳
Coding Agent Index, Codex with GPT-6.1 Sol at xhigh62.9 ⟳
Coding Agent Index, Codex with GPT-6.1 Sol at xhigh, cost per task$1.04 ⟳
Coding Agent Index, Codex with GPT-6.1 Sol at xhigh, agent time per task16 min ⟳
Coding Agent Index, Claude Code with Fable 5.1 at max62.2 ⟳
Coding Agent Index, Claude Code with Fable 5.1 at max, cost per task$12.39 ⟳
Coding Agent Index, Claude Code with Fable 5.1 at max, agent time per task35 min ⟳
Coding Agent Index, Codex with GPT-6 Astra at max61.6 ⟳
Coding Agent Index, Codex with GPT-6 Astra at max, cost per task$7.47 ⟳
Coding Agent Index, Codex with GPT-6 Astra at max, agent time per task29 min ⟳
Coding Agent Index, Codex with GPT-6.1 Sol at medium61.4 ⟳
Coding Agent Index, Codex with GPT-6.1 Sol at medium, cost per task$0.70 ⟳
Coding Agent Index, Codex with GPT-6.1 Sol at medium, agent time per task11 min ⟳
Coding Agent Index, Codex with GPT-6.1 Sol at high60.1 ⟳
Coding Agent Index, Codex with GPT-6.1 Sol at high, cost per task$0.89 ⟳
Coding Agent Index, Codex with GPT-6.1 Sol at high, agent time per task13 min ⟳
Coding Agent Index, Codex with GPT-6.1 Sol at max60.1 ⟳
Coding Agent Index, Codex with GPT-6.1 Sol at max, cost per task$1.55 ⟳
Coding Agent Index, Codex with GPT-6.1 Sol at max, agent time per task24 min ⟳
Coding Agent Index, Codex with GPT-6.1 Sol at low57.2 ⟳
Coding Agent Index, Codex with GPT-6.1 Sol at low, cost per task$0.50 ⟳
Coding Agent Index, Codex with GPT-6.1 Sol at low, agent time per task9 min ⟳
Coding Agent Index, Codex with GPT-6 Sol at max56.7 ⟳
Coding Agent Index, Codex with GPT-6 Sol at max, cost per task$2.99 ⟳
Coding Agent Index, Codex with GPT-6 Sol at max, agent time per task22 min ⟳
Coding Agent Index cost per task, Claude Code with Opus 5.5 at max against Codex with GPT-6.1 Sol at xhigh (our arithmetic)12.5×
Coding Agent Index cost per task, Codex with GPT-6.1 Sol at xhigh against Claude Code with Opus 5.5 at max (our arithmetic)0.08
Coding Agent Index cost per task, GPT-6.1 Sol at xhigh against GPT-6 Astra at max, both in Codex (our arithmetic)0.14

OpenAI

OpenAI Help Center, Codex credit rate card · read 29 Sep 2026, 22:10 CEST
Codex credits per million input tokens, GPT-6.1 Sol50
Codex credits per million cached input tokens, GPT-6.1 Sol2.5
Codex credits per million output tokens, GPT-6.1 Sol250
Codex credits per million input tokens, GPT-6 Astra250
Codex credits per million cached input tokens, GPT-6 Astra25
Codex credits per million output tokens, GPT-6 Astra1,250
Codex credits per million input tokens, GPT-6 Sol50
Codex credits per million cached input tokens, GPT-6 Sol5
Codex credits per million output tokens, GPT-6 Sol250
Codex credits per input token, GPT-6 Astra against GPT-6.1 Sol (our arithmetic)5×
Codex credits per cached input token, GPT-6 Astra against GPT-6.1 Sol (our arithmetic)10×

OpenAI

OpenAI, Codex plans and limits · read 29 Sep 2026, 22:10 CEST
OpenAI's estimate of Codex messages per 5 hours on Plus, GPT-6.1 Sol15–160
OpenAI's estimate of Codex messages per 5 hours on Plus, GPT-6 Astra5–45

OpenAI

coding deception, misrepresentation rate, GPT-6.1 Sol1.50%
coding deception, misrepresentation rate, GPT-6 Astra0.51%
coding deception, misrepresentation rate, GPT-6 Sol1.30%
a blocked tool with a warning: runs that tried to get around the block, GPT-6.1 Sol23.5%
a blocked tool with a warning: runs that tried to get around the block, GPT-6 Astra17.4%
simulated internal Codex tasks, flags at severity 3 or higher, GPT-6.1 Sol28
simulated internal Codex tasks, flags at severity 3 or higher, GPT-6 Astra27
simulated internal Codex tasks, flags at severity 3 or higher, GPT-6 Sol42
internal Codex tasks in OpenAI's deployment simulation49,650
fewer severity 3+ flags for GPT-6.1 Sol than GPT-6 Sol, per OpenAI33%
coding deception rate, GPT-6.1 Sol against GPT-6 Astra (our arithmetic)2.9×

r/codex posters

an r/codex post title about GPT-6 SolGPT 6 - Sol is an Idiot
points on 'GPT 6 - Sol is an Idiot', as Reddit showed them335
an r/codex post title about GPT-6 SolGPT-6 Sol is... not good
points on 'GPT-6 Sol is... not good', as Reddit showed them145
an r/codex post title about GPT-6 SolGPT-6 Sol: 33 minutes to do what GPT-5.6 Sol did in 38 seconds — what is going on?
points on 'GPT-6 Sol: 33 minutes to do what GPT-5.6 Sol did in 38 seconds — what is going on?', as Reddit showed them23

Model Fatigue (our runs)

Our Great Wave runs (Codex and Claude Code, one run per model at each tool's default; GPT-6.1 Sol also at low, high, xhigh and max) · run 28 and 29 Sep 2026
our Great Wave run, GPT-6.1 Sol at Codex's default: cost at API list prices$0.15
our Great Wave run, GPT-6.1 Sol at Codex's default: time5 min
our Great Wave run, GPT-6.1 Sol at Codex's default: Codex credits at the rate card's prices (our arithmetic)3.8
our Great Wave run, GPT-6.1 Sol at low: cost at API list prices$0.13
our Great Wave run, GPT-6.1 Sol at low: time5 min
our Great Wave run, GPT-6.1 Sol at low: Codex credits at the rate card's prices (our arithmetic)3.2
our Great Wave run, GPT-6.1 Sol at high: cost at API list prices$0.69
our Great Wave run, GPT-6.1 Sol at high: time22 min
our Great Wave run, GPT-6.1 Sol at high: Codex credits at the rate card's prices (our arithmetic)17.3
our Great Wave run, GPT-6.1 Sol at xhigh: cost at API list prices$0.58
our Great Wave run, GPT-6.1 Sol at xhigh: time20 min
our Great Wave run, GPT-6.1 Sol at xhigh: Codex credits at the rate card's prices (our arithmetic)14.4
our Great Wave run, GPT-6.1 Sol at max: cost at API list prices$0.75
our Great Wave run, GPT-6.1 Sol at max: time29 min
our Great Wave run, GPT-6.1 Sol at max: Codex credits at the rate card's prices (our arithmetic)18.8
our Great Wave run, GPT-6 Sol at the default: cost at API list prices$0.25
our Great Wave run, GPT-6 Sol at the default: time5 min
our Great Wave run, GPT-6 Sol at the default: Codex credits at the rate card's prices (our arithmetic)6.3
our Great Wave run, GPT-6 Astra at the default: cost at API list prices$1.60
our Great Wave run, GPT-6 Astra at the default: time10 min
our Great Wave run, GPT-6 Astra at the default: Codex credits at the rate card's prices (our arithmetic)40.0
our Great Wave run, Opus 5.5 at the default (Claude Code): cost at API list prices (Claude Code's own list-price figure)$1.03
our Great Wave run, Opus 5.5 at the default (Claude Code): time5 min
our GPT-6.1 Sol Great Wave run (default): cached input tokens168,576
our GPT-6.1 Sol Great Wave run (default): fresh input tokens25,427
our GPT-6.1 Sol Great Wave run (default): output tokens8,483
our Great Wave runs, GPT-6.1 Sol's cost against Opus 5.5's (our arithmetic)0.15

The Metropolitan Museum of Art

when the print was first publishedc. 1830–32

Sources

These are the pages the video and this page draw on. We keep a copy of each page as we read it, so a figure can be checked against what the page said at the time.

Credits

The narration in the video is an AI voice, made with ElevenLabs.

OpenAI's Terminal-Bench Science chart appears in the video redrawn from the values embedded in its launch post. This page gives the same values as a table.

Music in the video: "Airport Lounge" by Kevin MacLeod (incompetech.com), licensed under Creative Commons: By Attribution 4.0.