model fatıgue
Measured first look

Is OpenAI's Decisions API 10× faster? We measured 4.4×, not 10×

Published 7 Oct 2026Our test: bank-support messages

Pressing play loads the video from YouTube (Google), in privacy-enhanced mode. Privacy

This is the video written out, with every figure in full, each linked to its source in the table below.

Another decision model▶ 0:00

On 1 October, Cloudflare released two decision models, Clef and Clef-flash. On 6 October, OpenAI opened its own, the Decisions API, as a public beta. That makes three new models in one week that pick an answer from a list. On 2 October we gave Clef and Clef-flash our bank-message test. A few hours after the Decisions API opened, we gave it the same test.

What OpenAI said▶ 0:26

OpenAI first showed the Decisions API at its developer conference on 29 September. In its launch clip, 10,000 customer requests go through two of its APIs. The regular one, called the Responses API, takes 1.6 s a request. The Decisions API takes 150 ms. The clip calls that about 10× faster, and the Decisions guide that went up on 6 October says it answers about 10× faster than the Responses API.

Neither says how much text was in each request, where the requests were sent from, or how many were answered correctly.

What it is▶ 1:00

The Decisions API doesn't write any text, and neither does Jev, a decision model from a company called TypeSafe. Jev is what we measure new decision models against. You give the Decisions API a message or an image and a question with a fixed list of answers, and it picks one. It runs on GPT-6 Luna, the same model you can call through OpenAI's regular API, and for now Luna is the only model it offers.

The guide answered two questions we had before it opened. The price is $0.10 per million input tokens (tokens are the word pieces a model reads), and nothing for the answer. And it tells you how sure it is: every answer comes with a probability for each option on the list, plus a separate confidence.

The guide doesn't say how many options a question can have, so we tried. A question with 255 options worked, which is plenty for our test. At 256, the API returned an error saying the maximum is 255, and it refused a question with only one option. The API also reads images. We didn't test that.

Our test▶ 1:52

This is the test we gave Clef, Clef-flash, Jev and the two Kev models in our Clef vs Jev video. It has 3,080 messages that customers sent to a bank, from the public Banking77 collection. Each message belongs to one of 77 reasons, and every model picks from 78 options, the reasons plus none of these, which is never the right answer on this test. With each message, every model sees the same five similar messages that were already sorted, with their answers, and the same list.

We wrote down the setup and committed it before the first test message went out. On launch night we ran Jev (23:14 to 23:27 UTC), the Decisions API (23:27 to 23:40 UTC) and GPT-6 Luna through the Responses API (23:41 to 00:38 UTC), one after another, so all three were measured side by side. Luna ran with its reasoning switched off, the lowest setting the Responses API accepts for it. Every timed request went out one at a time, from a Mac mini in Berlin.

Clef, Clef-flash, Kev-27B and Kev-9B ran on the same messages, the same way, on 2 October, for our Clef vs Jev video. The two Kev models ran on GPUs we rented.

Speed: 4.4×, not 10×▶ 2:26

From Berlin, the median request to the Decisions API took 239 ms. GPT-6 Luna through the Responses API, on the same messages that night, took 1.05 s. So the Decisions API was 4.4× faster, not 10×. Jev took about the same as the Decisions API, 242 ms, and Cloudflare's Clef-flash, which we tested on 2 October, took 196 ms.

Time per request, from Berlin

The bar runs to the median; the ticks mark the 95th and 99th percentiles. One request at a time.

median95th and 99th percentile
0 s0.5 s1 s1.5 s2 s2.5 sseconds per requestDecisions APIDecisions API: median 239 ms, 95th percentile 370 ms, 99th 594 ms · Model Fatigue, run 7 Oct 2026, 01:13–02:39 CEST (the night of 6 to 7 October)239 msJevJev: median 242 ms, 95th percentile 307 ms, 99th 386 ms · Model Fatigue, run 7 Oct 2026, 01:13–02:39 CEST (the night of 6 to 7 October)242 msGPT-6 Luna (Responses API)GPT-6 Luna (Responses API): median 1,047 ms, 95th percentile 1,604 ms, 99th 2,373 ms · Model Fatigue, run 7 Oct 2026, 01:13–02:39 CEST (the night of 6 to 7 October)1,047 msClef-flashClef-flash: median 196 ms, 95th percentile 1,014 ms, 99th 2,172 ms · Model Fatigue, run 2 Oct 2026, 16:02–22:00 CEST196 msClefClef: median 562 ms, 95th percentile 1,147 ms, 99th 2,001 ms · Model Fatigue, run 2 Oct 2026, 16:02–22:00 CEST562 msKev-27BKev-27B: median 529 ms, 95th percentile 551 ms, 99th 598 ms · Model Fatigue, run 2 Oct 2026, 16:02–22:00 CEST529 msKev-9BKev-9B: median 437 ms, 95th percentile 548 ms, 99th 665 ms · Model Fatigue, run 2 Oct 2026, 16:02–22:00 CEST437 ms
0 s0.5 s1 s1.5 s2 s2.5 sseconds per requestDecisions APIDecisions API: median 239 ms, 95th percentile 370 ms, 99th 594 ms · Model Fatigue, run 7 Oct 2026, 01:13–02:39 CEST (the night of 6 to 7 October)239 msJevJev: median 242 ms, 95th percentile 307 ms, 99th 386 ms · Model Fatigue, run 7 Oct 2026, 01:13–02:39 CEST (the night of 6 to 7 October)242 msGPT-6 Luna (Responses API)GPT-6 Luna (Responses API): median 1,047 ms, 95th percentile 1,604 ms, 99th 2,373 ms · Model Fatigue, run 7 Oct 2026, 01:13–02:39 CEST (the night of 6 to 7 October)1,047 msClef-flashClef-flash: median 196 ms, 95th percentile 1,014 ms, 99th 2,172 ms · Model Fatigue, run 2 Oct 2026, 16:02–22:00 CEST196 msClefClef: median 562 ms, 95th percentile 1,147 ms, 99th 2,001 ms · Model Fatigue, run 2 Oct 2026, 16:02–22:00 CEST562 msKev-27BKev-27B: median 529 ms, 95th percentile 551 ms, 99th 598 ms · Model Fatigue, run 2 Oct 2026, 16:02–22:00 CEST529 msKev-9BKev-9B: median 437 ms, 95th percentile 548 ms, 99th 665 ms · Model Fatigue, run 2 Oct 2026, 16:02–22:00 CEST437 ms
The Decisions API, Jev and Luna the night of 6 October; Clef, Clef-flash and the Kev models on 2 October. The Kev models ran on GPUs we rented on Modal, so their times describe our setup. OpenAI's launch clip gave 150 ms and 1.6 s, without saying how they were measured.
The numbers in this chart

Two things separate our 4.4× from OpenAI's 10×. Our call to the Responses API, GPT-6 Luna with its reasoning switched off, was quicker than the one in OpenAI's clip: 1.05 s instead of 1.6 s. And from Berlin, none of our Decisions requests came back in 150 ms. The fastest took 165 ms. Just reaching OpenAI's API and getting an answer back, with no model involved, took 166 ms at the median before the run and 167 ms after it. That trip is a much bigger share of a quarter-second answer than of a one-second one.

Every Decisions response says how long OpenAI's servers spent on it. The median was 65 ms, and 144 ms at the 95th percentile. Our time per call minus OpenAI's own came to 170 ms at the median, about the length of the trip. So OpenAI's 150 ms is believable for someone calling from close to its servers. OpenAI didn't say how long its requests were, where it sent them from, or how many it sent at once, so its numbers and ours were measured in different ways. Luna's responses through the Responses API don't report a server time, so we can't split its 1.05 s the same way.

Where the Decisions API's time goes

One request at a time from a Mac mini in Berlin, the night of 6 October

time
Our time per request, median239 ms
OpenAI's reported server time per request, median65 ms
OpenAI's reported server time, 95th percentile144 ms
OpenAI's reported server time, 99th percentile367 ms
Our time minus OpenAI's server time, median170 ms
The trip to OpenAI's API and back with no model, median, before the run166 ms
The same trip, median, after the run167 ms
The fastest of our Decisions requests165 ms
Our Decisions requests that came back in under 150 ms0
Server time is the openai-processing-ms header OpenAI returns with each response: OpenAI's number, read from our responses. Responses API calls don't carry it in our records, so Luna's time can't be split this way.

The slowest calls matter if a person is waiting on the answer. One Decisions request in twenty took longer than 370 ms, and one in a hundred longer than 594 ms. For Jev those were 307 ms and 386 ms, and for Luna 1.60 s and 2.37 s. All of these are from one machine in Berlin, sending one request at a time. They say nothing about other places, or about sending many requests at once.

Accuracy▶ 3:38

Speed doesn't help if the answers are wrong. The Decisions API sorted 93.5% of the messages correctly. Jev, the same night, got 93.9%. On 38 messages Jev was right and the Decisions API was wrong, and on 26 it was the other way round. The Decisions API minus Jev comes to −0.39 points, and the 95% interval of that difference runs from −0.91 to +0.13 points, which includes zero. Across this many messages, a gap that small could easily be chance.

GPT-6 Luna through the Responses API, the same model asked a different way, got 94.2%. We had planned to compare the two routes on speed, not accuracy, so this comparison was made after the run, which makes it post hoc. On 41 messages Luna was right and the Decisions API wrong, and on 21 it was the other way round, nearly two to one. Luna's lead is +0.65 points, with a 95% interval from +0.16 to +1.17, which doesn't include zero. So on this test, the fast way of asking Luna was a little less accurate than the slow one. Because we didn't plan this comparison, we'd treat it as something to check again, on this test, not as a fact about the two routes in general.

Cloudflare's Clef-flash and Clef, from 2 October, are still the most accurate models we've measured on this test, at 95.2% and 95.1%.

Share right, with five examples

A dot per model with its 95% interval; the window starts at 90%, not zero

measured95% interval
90%91%92%93%94%95%96%share rightClef-flashClef-flash: 95.2%, interval 94.3 to 95.9 · Model Fatigue, run 2 Oct 2026, 16:02–22:00 CEST95.2%ClefClef: 95.1%, interval 94.2 to 95.8 · Model Fatigue, run 2 Oct 2026, 16:02–22:00 CEST95.1%GPT-6 Luna (Responses API)GPT-6 Luna (Responses API): 94.2%, interval 93.3 to 94.9 · Model Fatigue, run 7 Oct 2026, 01:13–02:39 CEST (the night of 6 to 7 October)94.2%JevJev: 93.9%, interval 93.0 to 94.7 · Model Fatigue, run 7 Oct 2026, 01:13–02:39 CEST (the night of 6 to 7 October)93.9%Kev-27BKev-27B: 93.8%, interval 92.9 to 94.6 · Model Fatigue, run 2 Oct 2026, 16:02–22:00 CEST93.8%Decisions APIDecisions API: 93.5%, interval 92.6 to 94.3 · Model Fatigue, run 7 Oct 2026, 01:13–02:39 CEST (the night of 6 to 7 October)93.5%Kev-9BKev-9B: 91.3%, interval 90.3 to 92.3 · Model Fatigue, run 2 Oct 2026, 16:02–22:00 CEST91.3%
90%91%92%93%94%95%96%share rightClef-flashClef-flash: 95.2%, interval 94.3 to 95.9 · Model Fatigue, run 2 Oct 2026, 16:02–22:00 CEST95.2%ClefClef: 95.1%, interval 94.2 to 95.8 · Model Fatigue, run 2 Oct 2026, 16:02–22:00 CEST95.1%GPT-6 Luna (Responses API)GPT-6 Luna (Responses API): 94.2%, interval 93.3 to 94.9 · Model Fatigue, run 7 Oct 2026, 01:13–02:39 CEST (the night of 6 to 7 October)94.2%JevJev: 93.9%, interval 93.0 to 94.7 · Model Fatigue, run 7 Oct 2026, 01:13–02:39 CEST (the night of 6 to 7 October)93.9%Kev-27BKev-27B: 93.8%, interval 92.9 to 94.6 · Model Fatigue, run 2 Oct 2026, 16:02–22:00 CEST93.8%Decisions APIDecisions API: 93.5%, interval 92.6 to 94.3 · Model Fatigue, run 7 Oct 2026, 01:13–02:39 CEST (the night of 6 to 7 October)93.5%Kev-9BKev-9B: 91.3%, interval 90.3 to 92.3 · Model Fatigue, run 2 Oct 2026, 16:02–22:00 CEST91.3%
The Decisions API, Jev and Luna on 7 October; the others on 2 October. Each interval is for one model's share right; the gaps between two models on the same messages have their own intervals, in the table below.
The numbers in this chart
share right95% interval
Clef-flash95.2%94.3 to 95.9
Clef95.1%94.2 to 95.8
GPT-6 Luna (Responses API)94.2%93.3 to 94.9
Jev93.9%93.0 to 94.7
Kev-27B93.8%92.9 to 94.6
Decisions API93.5%92.6 to 94.3
Kev-9B91.3%90.3 to 92.3

The same messages, two models at a time

Share right, the first model minus the second; the same night from Berlin

difference95% intervalonly the first rightonly the second right
Decisions API minus Jev−0.39 points−0.91 to +0.132638
GPT-6 Luna (Responses API) minus Jev+0.26 points−0.23 to +0.753123
GPT-6 Luna (Responses API) minus Decisions API, post hoc+0.65 points+0.16 to +1.174121
Post hoc: the Luna-against-Decisions comparison wasn't in the plan we committed before the run; we made it afterwards. Clef and Clef-flash ran on 2 October, so their pairs with Jev are on our Clef vs Jev page, against Jev's pass of that day.

Side by side▶ 4:41

Here are three of the test messages, picked by a rule we committed before the run: one the Decisions API got right and Jev got wrong, one the other way round, and one the Decisions API got wrong while it was at least as sure as its cut-off for handling a message without a person (the next section explains the cut-off). Each was drawn at random from the messages short enough to read on screen.

"help me with my transfer". The answer on file is a failed transfer. The Decisions API picks that, though it's only 36% sure, and Jev says none of these, which counts as wrong.

"There is an unauthorized fee." Jev and Clef, the larger of Cloudflare's two models, say a card payment fee was charged, which is right. The Decisions API says an extra charge on a statement, and it's 84% sure. That one's wrong.

"How do I change currencies to euros?" The Decisions API, Jev and Clef all say exchanging money in the app. The answer on file is which currencies the bank supports, so all three count as wrong, and the Decisions API was 99% sure of its answer.

These three are not typical. Of all the test messages, the three models all got 2,841 right, and all got 125 wrong. Clef's answers are from its run on 2 October.

Three messages, side by side

Each model's answer and its probability for it, with the five sorted examples; Clef's from 2 October

messagemodelanswerprobability
help me with my transfer (on file: failed_transfer)Decisions APIfailed_transfer36%right
Jevnone_of_these60%wrong
Cleffailed_transfer58%right
There is an unauthorized fee. (on file: card_payment_fee_charged)Decisions APIextra_charge_on_statement84%wrong
Jevcard_payment_fee_charged48%right
Clefcard_payment_fee_charged34%right
How do I change currencies to euros? (on file: fiat_currency_support)Decisions APIexchange_via_app99%wrong
Jevexchange_via_app94%wrong
Clefexchange_via_app74%wrong
Messages as written in the test set. Picked by build/side_by_side.py, whose rule was committed before the Decisions API's first test answer existed: one the Decisions API got right and Jev wrong, one the reverse, one the Decisions API got wrong at or above its cut-off, each drawn at random from messages of at most seventy characters.

Who gets a person▶ 5:30

That probability for every answer matters if you route messages: let the model handle the ones it's sure of, and send the rest to a person. On a separate set of 573 messages, not the test ones, we set each model's cut-off to aim for about 2% wrong among the messages it handles. For the Decisions API and for Jev, that cut-off came out at 0.99.

On the test, the Decisions API then sent 160 of every thousand messages to a person, and Jev sent 159. The Decisions API was wrong on 2.13% of the messages it handled, and Jev on 2.05%. Both are just over the target.

With the Decisions API, the euros message is the kind that gets through: at 99%, nobody checks it. In all, 55 of its wrong answers were at or above the cut-off. Jev was 94% sure of the same wrong answer, below its cut-off, so Jev would have sent it to a person.

There's one catch, and Jev has it too. The probabilities come in whole percents, and most answers say ninety-nine or a hundred percent: 2,586 of the Decisions API's 3,080 answers did, which is 84%. When it said ninety percent or more, its stated confidence averaged 0.994, and it was right 96.1% of the time. Even the hundred-percent answers were wrong more than once in a hundred: 1.58% of the time for the Decisions API on the test and 1.59% for Jev. On the separate set the figures were 1.06% and 1.86%, so when we aimed for 1% wrong instead, neither model had a setting strict enough to get there.

Who gets a person

Each model's cut-off on its stated confidence, set on the separate set of messages to aim at 2% wrong, then applied to the test

modelruncut-offhandled alonewrong among handledto a person, per thousandAUROC
Decisions API7 October0.9984.0%2.13%1600.85
Jev7 October0.9984.1%2.05%1590.85
GPT-6 Luna (Responses API), written confidence7 October0.9090.9%2.75%910.86
Clef2 October0.8093.2%2.40%680.90
Clef-flash2 October0.7381.7%2.15%1830.82
Kev-27B2 October0.7991.0%2.28%900.92
Kev-9B2 October0.8182.7%2.47%1730.91
Wrong among handled, with its 95% interval: Decisions API 1.6% to 2.8%, Jev 1.6% to 2.7%. AUROC: how well a model's confidence separates its right answers from its wrong ones, where one half is a coin flip and one is perfect. The hundred-percent answers, wrong among them: Decisions API 1.58% on the test and 1.06% on the separate set; Jev 1.59% and 1.86%.

GPT-6 Luna through the Responses API returns no probabilities. We asked it to write down a confidence with its answer, which is a different kind of number, so its row is there for reference. Clef, Clef-flash and the two Kev models are from 2 October.

Price▶ 6:41

A thousand of these decisions cost $0.114 on OpenAI's price list. Jev costs $0.078, Luna through the Responses API $0.143, Clef-flash $0.180, and Clef $0.480. Jev charges $0.042 per million tokens, so OpenAI's price per token is about 2.4× Jev's. But OpenAI counts fewer tokens for the same request: 1,136 on average, where Jev counts 1,846. So per decision, the Decisions API comes to about 1.5× Jev's price.

Dollars per thousand decisions

Price list times the tokens each response reports

$0$0.125$0.25$0.375$0.5dollars per thousand decisionsJevJev: $0.078 · Model Fatigue, run 7 Oct 2026, 01:13–02:39 CEST (the night of 6 to 7 October)$0.078Decisions APIDecisions API: $0.114 · Model Fatigue, run 7 Oct 2026, 01:13–02:39 CEST (the night of 6 to 7 October)$0.114GPT-6 Luna (Responses API)GPT-6 Luna (Responses API): $0.143 · Model Fatigue, run 7 Oct 2026, 01:13–02:39 CEST (the night of 6 to 7 October)$0.143Clef-flashClef-flash: $0.180 · Model Fatigue, run 2 Oct 2026, 16:02–22:00 CEST$0.180ClefClef: $0.480 · Model Fatigue, run 2 Oct 2026, 16:02–22:00 CEST$0.480
$0$0.125$0.25$0.375$0.5dollars per thousand decisionsJevJev: $0.078 · Model Fatigue, run 7 Oct 2026, 01:13–02:39 CEST (the night of 6 to 7 October)$0.078Decisions APIDecisions API: $0.114 · Model Fatigue, run 7 Oct 2026, 01:13–02:39 CEST (the night of 6 to 7 October)$0.114GPT-6 Luna (Responses API)GPT-6 Luna (Responses API): $0.143 · Model Fatigue, run 7 Oct 2026, 01:13–02:39 CEST (the night of 6 to 7 October)$0.143Clef-flashClef-flash: $0.180 · Model Fatigue, run 2 Oct 2026, 16:02–22:00 CEST$0.180ClefClef: $0.480 · Model Fatigue, run 2 Oct 2026, 16:02–22:00 CEST$0.480
No response carries a bill, so these are price lists times tokens, not what we were charged.
The numbers in this chart
dollars per thousand decisions
Jev$0.078
Decisions API$0.114
GPT-6 Luna (Responses API)$0.143
Clef-flash$0.180
Clef$0.480

None of these is a bill. Each is the service's price list times the tokens each response reports, because no response says what it cost. For the Decisions API, Jev, Clef and Clef-flash we count input tokens at each price list. For Luna through the Responses API we count input, cache writes and output at OpenAI's rates. At the price list, the whole Decisions test came to $0.35.

What a thousand decisions cost

Each price list times the tokens each response reports

modelinput tokens per request, averageper thousand decisions
Jev1,846$0.078
Decisions API1,136$0.114
GPT-6 Luna (Responses API)1,066$0.143
Clef-flash1,999$0.180
Clef1,999$0.480
Every request carries the same message, five examples and the same list of answers; the token counts differ because each service counts tokens its own way. Prices: the Decisions API $0.10 per million input tokens and nothing else (OpenAI's guide); Jev $0.042 per million input tokens (TypeSafe's price list). Luna's figure counts input, cache writes and output at OpenAI's rates.

Without examples▶ 7:13

When Cloudflare launched Clef, its own table had Clef 14.5 points ahead of Jev on this collection with no examples (94.20 against 79.74, in macro-F1). So we asked six models the bare question ourselves, the message alone with no sorted examples and the 77 reasons as options: Jev, the two Clefs, the Decisions API, and the two open models we've tested before, Kev-27B and Kev-9B.

Clef got 94.2% right (macro-F1 94.20, as in Cloudflare's table), and Jev 80.7%. The Decisions API got 77.7%, the lowest of the six. On our question, with the five examples, it scores 15.8 points higher, the biggest difference of the six, though the two questions also differ in layout and in offering none of these. So without examples, on this test, it did worse than Jev, by 3.0 points.

Without examples, and with five

Share right on the same messages: the message alone, then with five sorted examples

message alonewith five examples
75%80%85%90%95%100%share rightdifference, pointsClefClef: message alone 94.2% (run 2 Oct 2026, 16:02–22:00 CEST), with five examples 95.1% (run 2 Oct 2026, 16:02–22:00 CEST) · Model Fatigue+0.8Clef-flashClef-flash: message alone 90.9% (run 2 Oct 2026, 16:02–22:00 CEST), with five examples 95.2% (run 2 Oct 2026, 16:02–22:00 CEST) · Model Fatigue+4.3Kev-27BKev-27B: message alone 86.8% (run 2 Oct 2026, 16:02–22:00 CEST), with five examples 93.8% (run 2 Oct 2026, 16:02–22:00 CEST) · Model Fatigue+7.0Kev-9BKev-9B: message alone 83.3% (run 2 Oct 2026, 16:02–22:00 CEST), with five examples 91.3% (run 2 Oct 2026, 16:02–22:00 CEST) · Model Fatigue+8.0JevJev: message alone 80.7% (run 2 Oct 2026, 16:02–22:00 CEST), with five examples 93.9% (run 7 Oct 2026, 01:13–02:39 CEST (the night of 6 to 7 October)) · Model Fatigue+13.2Decisions APIDecisions API: message alone 77.7% (run 7 Oct 2026, 02:43–02:45 CEST), with five examples 93.5% (run 7 Oct 2026, 01:13–02:39 CEST (the night of 6 to 7 October)) · Model Fatigue+15.8
75%80%85%90%95%100%share rightdifference, pointsClefClef: message alone 94.2% (run 2 Oct 2026, 16:02–22:00 CEST), with five examples 95.1% (run 2 Oct 2026, 16:02–22:00 CEST) · Model Fatigue+0.8Clef-flashClef-flash: message alone 90.9% (run 2 Oct 2026, 16:02–22:00 CEST), with five examples 95.2% (run 2 Oct 2026, 16:02–22:00 CEST) · Model Fatigue+4.3Kev-27BKev-27B: message alone 86.8% (run 2 Oct 2026, 16:02–22:00 CEST), with five examples 93.8% (run 2 Oct 2026, 16:02–22:00 CEST) · Model Fatigue+7.0Kev-9BKev-9B: message alone 83.3% (run 2 Oct 2026, 16:02–22:00 CEST), with five examples 91.3% (run 2 Oct 2026, 16:02–22:00 CEST) · Model Fatigue+8.0JevJev: message alone 80.7% (run 2 Oct 2026, 16:02–22:00 CEST), with five examples 93.9% (run 7 Oct 2026, 01:13–02:39 CEST (the night of 6 to 7 October)) · Model Fatigue+13.2Decisions APIDecisions API: message alone 77.7% (run 7 Oct 2026, 02:43–02:45 CEST), with five examples 93.5% (run 7 Oct 2026, 01:13–02:39 CEST (the night of 6 to 7 October)) · Model Fatigue+15.8
Our runs. The Decisions API both ways on 7 October; Jev's message-alone pass on 2 October and its pass with examples on 7 October; the others both ways on 2 October.
The numbers in this chart
message alonewith five examplesdifference, points
Clef94.2%95.1%+0.8 points
Clef-flash90.9%95.2%+4.3 points
Kev-27B86.8%93.8%+7.0 points
Kev-9B83.3%91.3%+8.0 points
Jev80.7%93.9%+13.2 points
Decisions API77.7%93.5%+15.8 points

The Decisions API's pass without examples ran on 7 October, right after the timed run, and wasn't sent one request at a time, so we quote no times from it. Jev's and the other four models' passes without examples ran on 2 October.

Every model, with and without examples

Share right on the same messages: the message alone (the Decision Index's question), then with five sorted examples

modelmessage alonemacro-F1, message alonewith five examplesdifference
Clef94.2%94.2095.1%+0.8 points
Clef-flash90.9%90.8595.2%+4.3 points
Kev-27B86.8%86.4193.8%+7.0 points
Kev-9B83.3%83.0391.3%+8.0 points
Jev80.7%80.0093.9%+13.2 points
Decisions API77.7%76.9593.5%+15.8 points
The message alone: the 77 reasons as options, no examples, no none of these, in the Decision Index's layout; so the difference is not the examples' alone. Run dates: the Decisions API both ways on 7 October; Jev's message-alone pass on 2 October and its pass with examples on 7 October; the others both ways on 2 October. Cloudflare's own table: Clef 94.20, Jev 79.74 (macro-F1).

What this doesn't show▶ 7:54

This is one task, bank messages in English, timed from one place on the night it opened. The test is text only, so it says nothing about images.

We checked whether the Decisions API had memorized the public collection of bank messages our test comes from, and found no sign that it had. Asked 200 messages from the collection's training part with no examples, it got 69.5% of the originals right and 69.0% of the same messages reworded, a gap of +0.5 points with a 95% interval from −3.5 to +5.0. When we sent 200 of the test messages a second time, it gave the same answer to 200 of them.

And OpenAI calls this a public beta. Its terms say a beta service "may be changed at any time without notice", so it may not behave the same next week, or once it's out of beta.

Which to use▶ 8:25

If Jev already works for you, this is no reason to switch. On this test the Decisions API was about as accurate, about as fast from Berlin on a typical request, routed about the same, and cost about 1.5× as much. It's worth a look if you already build on OpenAI, or if you need it to read images, which we didn't test. If accuracy matters most, Cloudflare's Clef-flash was a little ahead on this test, at about 2.3× Jev's price. And the Decisions API is fast, just not 10× faster than OpenAI's regular API from here: 4.4×.

The Decisions API and Jev, the same night

Our test from Berlin

Decisions APIJev
Share right93.5%93.9%
Median time per request, from Berlin239 ms242 ms
Per thousand decisions, price list$0.114$0.078
Sent to a person, per thousand160159
A probability for every answeryesyes

AI Actions (made by the company behind this channel)▶ 9:05

The company behind this channel makes AI Actions, an iPhone app. One of its actions for iOS Shortcuts is called Choose Category, and it runs on Jev. It sorts a piece of text, like a reminder about an invoice, into categories you set. You choose how sure it has to be, and when it isn't sure enough about any category, it says it's unsure instead of guessing. It's linked in the video's description.

Next time▶ 9:30

When the Decisions API is out of beta, we'll run it again. If you sort messages with one of these models, tell us in the video's comments what we should test them on next.

Every number

These are all 223 figures behind the video and this page, grouped by whose they are, with the page each came from and when we read it. Figures marked ⟳ can move. When a re-read finds a change, the new value shows next to the one from the video.

Model Fatigue

Our test, the night the Decisions API opened: Jev, the Decisions API and GPT-6 Luna through the Responses API, 3,080 Banking77 test messages, five sorted examples each, 78 options, one request at a time from a Mac mini in Berlin · run 7 Oct 2026, 01:13–02:39 CEST (the night of 6 to 7 October)
Banking77 test messages every model answered3,080
the Decisions API, share of the test messages sorted correctly93.5%
the Decisions API, share right, 95% interval, low92.6
the Decisions API, share right, 95% interval, high94.3
the Decisions API, times it answered none of these4
the Decisions API, median time per request from Berlin239 ms
the Decisions API, 95th percentile time per request from Berlin370 ms
the Decisions API, 99th percentile time per request from Berlin594 ms
the Decisions API, dollars per 1,000 decisions (price list × the tokens each response reports)$0.114
the Decisions API, share of the test messages it handled without a person84.0%
the Decisions API, wrong among the messages it handled without a person2.13%
the Decisions API, messages per 1,000 sent to a person160
the Decisions API, calibration error of its stated confidence (ECE; lower is closer)0.041
the Decisions API, how well its confidence separates its right answers from its wrong ones (AUROC)0.85
Jev, share of the test messages sorted correctly93.9%
Jev, share right, 95% interval, low93.0
Jev, share right, 95% interval, high94.7
Jev, times it answered none of these13
Jev, median time per request from Berlin242 ms
Jev, 95th percentile time per request from Berlin307 ms
Jev, 99th percentile time per request from Berlin386 ms
Jev, dollars per 1,000 decisions (price list × the tokens each response reports)$0.078
Jev, share of the test messages it handled without a person84.1%
Jev, wrong among the messages it handled without a person2.05%
Jev, messages per 1,000 sent to a person159
Jev, calibration error of its stated confidence (ECE; lower is closer)0.035
Jev, how well its confidence separates its right answers from its wrong ones (AUROC)0.85
GPT-6 Luna through the Responses API, share of the test messages sorted correctly94.2%
GPT-6 Luna through the Responses API, share right, 95% interval, low93.3
GPT-6 Luna through the Responses API, share right, 95% interval, high94.9
GPT-6 Luna through the Responses API, times it answered none of these5
GPT-6 Luna through the Responses API, median time per request from Berlin1,047 ms
GPT-6 Luna through the Responses API, 95th percentile time per request from Berlin1,604 ms
GPT-6 Luna through the Responses API, 99th percentile time per request from Berlin2,373 ms
GPT-6 Luna through the Responses API, median time per request from Berlin, in seconds1.05 s
GPT-6 Luna through the Responses API, 95th percentile time per request from Berlin, in seconds1.60 s
GPT-6 Luna through the Responses API, 99th percentile time per request from Berlin, in seconds2.37 s
GPT-6 Luna through the Responses API, dollars per 1,000 decisions (price list × the tokens each response reports)$0.143
GPT-6 Luna through the Responses API, share of the test messages it handled without a person90.9%
GPT-6 Luna through the Responses API, wrong among the messages it handled without a person2.75%
GPT-6 Luna through the Responses API, messages per 1,000 sent to a person91
GPT-6 Luna through the Responses API, calibration error of its stated confidence (ECE; lower is closer)0.024
GPT-6 Luna through the Responses API, how well its confidence separates its right answers from its wrong ones (AUROC)0.86
the Decisions API, wrong among handled, 95% interval, low1.6%
the Decisions API, wrong among handled, 95% interval, high2.8%
Jev, wrong among handled, 95% interval, low1.6%
Jev, wrong among handled, 95% interval, high2.7%
The Decisions API's wrong answers at 0.99 or above, which the 2% cut-off would let through without a person55
Messages sent to the Decisions API a second time that got the same answer200
Messages sent to the Decisions API a second time200
Messages Jev got right and the Decisions API got wrong38
Messages the Decisions API got right and Jev got wrong26
Messages Jev got right and Luna got wrong23
Messages Luna got right and Jev got wrong31
Messages Luna got right and the Decisions API got wrong (post hoc)41
Messages the Decisions API got right and Luna got wrong (post hoc)21
Luna's median time over the Decisions API's, same messages, same night (our arithmetic)4.4×
The Decisions API's median time minus the trip, before the run (our arithmetic)73 ms
The Decisions API's cost per decision over Jev's (our arithmetic)1.5×
The Decisions API test pass, all 3,080 messages, at the price list$0.35
Message A: the Decisions API's probability for the answer it picked36%
Message A: the Decisions API's probability for its second answer19%
Message A: Jev's probability for the answer it picked60%
Message A: Jev's probability for its second answer22%
Message B: the Decisions API's probability for the answer it picked84%
Message B: the Decisions API's probability for its second answer6%
Message B: Jev's probability for the answer it picked48%
Message B: Jev's probability for its second answer36%
Message C: the Decisions API's probability for the answer it picked99%
Message C: the Decisions API's probability for its second answer1%
Message C: Jev's probability for the answer it picked94%
Message C: Jev's probability for its second answer6%
The Decisions API minus Jev, share right, same messages, same night−0.39 points
The Decisions API minus Jev, 95% interval, low−0.91
The Decisions API minus Jev, 95% interval, high+0.13
GPT-6 Luna through the Responses API minus Jev, share right, same messages, same night+0.26 points
Luna minus Jev, 95% interval, low−0.23
Luna minus Jev, 95% interval, high+0.75
Luna through the Responses API minus the Decisions API, same messages (post hoc: compared after the run, not in the frozen plan)+0.65 points
Luna minus the Decisions API, 95% interval, low (post hoc)+0.16
Luna minus the Decisions API, 95% interval, high (post hoc)+1.17
Luna-only right answers per Decisions-only right answer (post hoc, our arithmetic)2.0

Model Fatigue

Our frozen protocol for the bank-message test (FREEZE.md) · frozen 30 Sep 2026
Reasons a message can be sorted into77
Already-sorted example messages shown with each message5
Jev's price per million input tokens (TypeSafe's rate card)$0.042

Model Fatigue

Our results write-up (RESULTS.md, "OpenAI's Decisions API (added 2026-10-07)") · written 7 Oct 2026
Options in each question: the 77 reasons plus none of these78
The stricter target no setting could meet for the Decisions API or Jev1%
Options in the Decision Index's question (the reasons, without none of these)77

Model Fatigue

Our pass of the Decisions API over the separate set of 573 messages its cut-off was set on (not test messages) · run 7 Oct 2026, 01:09 CEST
Messages in the separate set every model's cut-off was set on (not test messages)573
the Decisions API, cut-off on its stated confidence, set on the separate set for a 2% error target0.99
The Decisions API, separate set: wrong among its answers stated at 1.001.06%

Model Fatigue

Our passes of Jev and GPT-6 Luna over the separate set of 573 messages their cut-offs were set on (not test messages) · run 30 Sep 2026, 21:28 CEST
Jev, cut-off on its stated confidence, set on the separate set for a 2% error target0.99
GPT-6 Luna through the Responses API, cut-off on its stated confidence, set on the separate set for a 2% error target0.90
Jev, separate set: wrong among its answers stated at 1.001.86%

Model Fatigue

Our test of Clef, Clef-flash, Kev-27B and Kev-9B on the same messages (the Clef vs Jev video's run) · run 2 Oct 2026, 16:02–22:00 CEST
Clef, share of the test messages sorted correctly95.1%
Clef, share right, 95% interval, low94.2
Clef, share right, 95% interval, high95.8
Clef, times it answered none of these0
Clef, median time per request from Berlin562 ms
Clef, 95th percentile time per request from Berlin1,147 ms
Clef, 99th percentile time per request from Berlin2,001 ms
Clef, dollars per 1,000 decisions (price list × the tokens each response reports)$0.480
Clef, share of the test messages it handled without a person93.2%
Clef, wrong among the messages it handled without a person2.40%
Clef, messages per 1,000 sent to a person68
Clef, calibration error of its stated confidence (ECE; lower is closer)0.030
Clef, how well its confidence separates its right answers from its wrong ones (AUROC)0.90
Clef-flash, share of the test messages sorted correctly95.2%
Clef-flash, share right, 95% interval, low94.3
Clef-flash, share right, 95% interval, high95.9
Clef-flash, times it answered none of these0
Clef-flash, median time per request from Berlin196 ms
Clef-flash, 95th percentile time per request from Berlin1,014 ms
Clef-flash, 99th percentile time per request from Berlin2,172 ms
Clef-flash, dollars per 1,000 decisions (price list × the tokens each response reports)$0.180
Clef-flash, share of the test messages it handled without a person81.7%
Clef-flash, wrong among the messages it handled without a person2.15%
Clef-flash, messages per 1,000 sent to a person183
Clef-flash, calibration error of its stated confidence (ECE; lower is closer)0.143
Clef-flash, how well its confidence separates its right answers from its wrong ones (AUROC)0.82
Kev-27B, share of the test messages sorted correctly93.8%
Kev-27B, share right, 95% interval, low92.9
Kev-27B, share right, 95% interval, high94.6
Kev-27B, times it answered none of these21
Kev-27B, median time per request from Berlin529 ms
Kev-27B, 95th percentile time per request from Berlin551 ms
Kev-27B, 99th percentile time per request from Berlin598 ms
Kev-27B, share of the test messages it handled without a person91.0%
Kev-27B, wrong among the messages it handled without a person2.28%
Kev-27B, messages per 1,000 sent to a person90
Kev-27B, calibration error of its stated confidence (ECE; lower is closer)0.014
Kev-27B, how well its confidence separates its right answers from its wrong ones (AUROC)0.92
Kev-9B, share of the test messages sorted correctly91.3%
Kev-9B, share right, 95% interval, low90.3
Kev-9B, share right, 95% interval, high92.3
Kev-9B, times it answered none of these110
Kev-9B, median time per request from Berlin437 ms
Kev-9B, 95th percentile time per request from Berlin548 ms
Kev-9B, 99th percentile time per request from Berlin665 ms
Kev-9B, share of the test messages it handled without a person82.7%
Kev-9B, wrong among the messages it handled without a person2.47%
Kev-9B, messages per 1,000 sent to a person173
Kev-9B, calibration error of its stated confidence (ECE; lower is closer)0.038
Kev-9B, how well its confidence separates its right answers from its wrong ones (AUROC)0.91
Clef, input tokens per request as its service counted them, average1,999
Clef-flash, input tokens per request as its service counted them, average1,999
Message A: Clef's probability for the answer it picked58%
Message A: Clef's probability for its second answer10%
Message B: Clef's probability for the answer it picked34%
Message B: Clef's probability for its second answer24%
Message C: Clef's probability for the answer it picked74%
Message C: Clef's probability for its second answer17%
Jev, share right with the message alone (the Decision Index's question)80.7%
Jev, macro-F1 with the message alone80.00
Clef, share right with the message alone (the Decision Index's question)94.2%
Clef, macro-F1 with the message alone94.20
Clef, share right on our question (five examples) minus on the message alone, points+0.8 points
Clef-flash, share right with the message alone (the Decision Index's question)90.9%
Clef-flash, macro-F1 with the message alone90.85
Clef-flash, share right on our question (five examples) minus on the message alone, points+4.3 points
Kev-27B, share right with the message alone (the Decision Index's question)86.8%
Kev-27B, macro-F1 with the message alone86.41
Kev-27B, share right on our question (five examples) minus on the message alone, points+7.0 points
Kev-9B, share right with the message alone (the Decision Index's question)83.3%
Kev-9B, macro-F1 with the message alone83.03
Kev-9B, share right on our question (five examples) minus on the message alone, points+8.0 points

Model Fatigue

Our passes of Clef, Clef-flash, Kev-27B and Kev-9B over the separate set of 573 messages their cut-offs were set on (not test messages) · run 2 Oct 2026, 15:38–20:53 CEST
Clef, cut-off on its stated confidence, set on the separate set for a 2% error target0.80
Clef-flash, cut-off on its stated confidence, set on the separate set for a 2% error target0.73
Kev-27B, cut-off on its stated confidence, set on the separate set for a 2% error target0.79
Kev-9B, cut-off on its stated confidence, set on the separate set for a 2% error target0.81

OpenAI

OpenAI's time per request for the Decisions API (launch clip)150 ms
OpenAI's time per request for the Responses API (launch clip)1.6 s
OpenAI's speed-up, Decisions API over the Responses API (launch clip)10×
Customer requests in OpenAI's launch clip10,000

OpenAI

OpenAI, the Decisions API guide · read 7 Oct 2026, 01:05 CEST
OpenAI's speed-up in the Decisions guide10×
The Decisions API's price per million input tokens (nothing else billed)$0.10

Model Fatigue

Our arithmetic on the raw call records (derived.json, this page's own file) · run 7 Oct 2026, 01:13–02:39 CEST (the night of 6 to 7 October)
Decisions calls that came back to Berlin in under 150 ms0
The fastest Decisions call from Berlin165 ms
The trip from Berlin to OpenAI's API and back with no model work, median, before the run166 ms
The same trip, median, after the run167 ms
When the timed run went out (first network-floor check to the last)23:13 to 00:39 UTC
When Jev's test pass ran23:14 to 23:27 UTC
When the Decisions API's test pass ran23:27 to 23:40 UTC
When Luna's test pass ran23:41 to 00:38 UTC
the Decisions API, input tokens per request as its service counted them, average1,136
Jev, input tokens per request as its service counted them, average1,846
GPT-6 Luna through the Responses API, input tokens per request as its service counted them, average1,066
Decisions answers stated at 0.99 or 1.002,586
Decisions answers stated at 1.002,146
Decisions answers stated at 0.90 or more2,902
The Decisions API's average stated confidence on those answers0.994
Share of those answers that were right96.1%
The Decisions API, test: wrong among its answers stated at 1.001.58%
Jev, test: wrong among its answers stated at 1.001.59%
Time OpenAI's servers report spending on each Decisions call (openai-processing-ms), median65 ms
OpenAI's reported server time per Decisions call, 95th percentile144 ms
OpenAI's reported server time per Decisions call, 99th percentile367 ms
Our time per Decisions call minus OpenAI's reported server time, median (our arithmetic)170 ms
Jev's token count per request over the Decisions API's (our arithmetic)1.6×
Share of Decisions answers stated at 0.99 or 1.00 (our arithmetic)84%
Share of Decisions answers stated at 0.90 or more (our arithmetic)94%

Model Fatigue

Our protocol addendum for the Decisions API, committed before the first test call (commit 23057db1) · frozen 7 Oct 2026, 01:13 CEST
The error target every cut-off was set for: wrong answers among those a model handles alone2%

Model Fatigue, arithmetic on our runs

Our arithmetic across the two nights' passes: the Decisions API, Jev and Luna on 7 October, Clef, Clef-flash, the Kev models and every model's message-alone pass but the Decisions API's on 2 October · run 2 Oct 2026 and 7 Oct 2026
Test messages the Decisions API, Jev and Clef all got right2,841
Test messages all three got wrong125
the Decisions API, share right on our question (five examples) minus on the message alone, points+15.8 points
Jev, share right on our question (five examples) minus on the message alone, points+13.2 points
Clef-flash's cost per decision over Jev's (our arithmetic)2.3×
Jev minus the Decisions API with the message alone (our arithmetic)3.0 points

Model Fatigue

Our calls to the Decisions API with 1 to 1,000 options, to find its limit · sent 7 Oct 2026, 01:08 and 07:14 CEST
The most options the Decisions API accepted in one question255
Options in the question the API refused as too many256
The fewest options the API accepted in one question2

Model Fatigue

Our passes of the Decisions API without examples, and the memorisation check (about six requests at a time, from another machine; no times quoted from them) · run 7 Oct 2026, 02:43–02:45 CEST
the Decisions API, share right with the message alone (the Decision Index's question)77.7%
the Decisions API, macro-F1 with the message alone76.95
Messages from the public collection's training part, asked as written and reworded200
The Decisions API, share right on the original messages (no examples)69.5%
The Decisions API, share right on the same messages reworded69.0%
Originals minus rewordings (our arithmetic)+0.5 points
Originals minus rewordings, 95% interval, low−3.5
Originals minus rewordings, 95% interval, high+5.0

Cloudflare

Clef on BANKING77 in Cloudflare's table (macro-F1)94.20
Jev on BANKING77 in Cloudflare's table (macro-F1)79.74
Clef minus Jev on BANKING77 in Cloudflare's table (our arithmetic)14.5 points

Model Fatigue, arithmetic on the two price lists

Our arithmetic on two price lists: OpenAI's Decisions guide (read 7 Oct 2026, 01:05 CEST) and Jev's rate card in our frozen protocol (30 Sep 2026) · read 7 Oct 2026, 01:05 CEST, and 30 Sep 2026
The Decisions API's price per token over Jev's (our arithmetic)2.4×

Sources

These are the pages the video and this page draw on. We keep a copy of each page as we read it, so a figure can be checked against what the page said at the time.

Credits

The narration in the video is an AI voice, made with ElevenLabs.

Music in the video: "Airport Lounge" by Kevin MacLeod (incompetech.com), licensed under Creative Commons: By Attribution 4.0.

Further reading