This is the video written out, with every figure in full, each linked to its source in the table below.
Another decision model▶ 0:00
On 1 October, Cloudflare released two decision models, Clef and Clef-flash. On 6 October, OpenAI opened its own, the Decisions API, as a public beta. That makes three new models in one week that pick an answer from a list. On 2 October we gave Clef and Clef-flash our bank-message test. A few hours after the Decisions API opened, we gave it the same test.
What OpenAI said▶ 0:26
OpenAI first showed the Decisions API at its developer conference on 29 September. In its launch clip, 10,000 customer requests go through two of its APIs. The regular one, called the Responses API, takes 1.6 s a request. The Decisions API takes 150 ms. The clip calls that about 10× faster, and the Decisions guide that went up on 6 October says it answers about 10× faster than the Responses API.
Neither says how much text was in each request, where the requests were sent from, or how many were answered correctly.
What it is▶ 1:00
The Decisions API doesn't write any text, and neither does Jev, a decision model from a company called TypeSafe. Jev is what we measure new decision models against. You give the Decisions API a message or an image and a question with a fixed list of answers, and it picks one. It runs on GPT-6 Luna, the same model you can call through OpenAI's regular API, and for now Luna is the only model it offers.
The guide answered two questions we had before it opened. The price is $0.10 per million input tokens (tokens are the word pieces a model reads), and nothing for the answer. And it tells you how sure it is: every answer comes with a probability for each option on the list, plus a separate confidence.
The guide doesn't say how many options a question can have, so we tried. A question with 255 options worked, which is plenty for our test. At 256, the API returned an error saying the maximum is 255, and it refused a question with only one option. The API also reads images. We didn't test that.
Our test▶ 1:52
This is the test we gave Clef, Clef-flash, Jev and the two Kev models in our Clef vs Jev video. It has 3,080 messages that customers sent to a bank, from the public Banking77 collection. Each message belongs to one of 77 reasons, and every model picks from 78 options, the reasons plus none of these, which is never the right answer on this test. With each message, every model sees the same five similar messages that were already sorted, with their answers, and the same list.
We wrote down the setup and committed it before the first test message went out. On launch night we ran Jev (23:14 to 23:27 UTC), the Decisions API (23:27 to 23:40 UTC) and GPT-6 Luna through the Responses API (23:41 to 00:38 UTC), one after another, so all three were measured side by side. Luna ran with its reasoning switched off, the lowest setting the Responses API accepts for it. Every timed request went out one at a time, from a Mac mini in Berlin.
Clef, Clef-flash, Kev-27B and Kev-9B ran on the same messages, the same way, on 2 October, for our Clef vs Jev video. The two Kev models ran on GPUs we rented.
Speed: 4.4×, not 10×▶ 2:26
From Berlin, the median request to the Decisions API took 239 ms. GPT-6 Luna through the Responses API, on the same messages that night, took 1.05 s. So the Decisions API was 4.4× faster, not 10×. Jev took about the same as the Decisions API, 242 ms, and Cloudflare's Clef-flash, which we tested on 2 October, took 196 ms.
Time per request, from Berlin
The bar runs to the median; the ticks mark the 95th and 99th percentiles. One request at a time.
Two things separate our 4.4× from OpenAI's 10×. Our call to the Responses API, GPT-6 Luna with its reasoning switched off, was quicker than the one in OpenAI's clip: 1.05 s instead of 1.6 s. And from Berlin, none of our Decisions requests came back in 150 ms. The fastest took 165 ms. Just reaching OpenAI's API and getting an answer back, with no model involved, took 166 ms at the median before the run and 167 ms after it. That trip is a much bigger share of a quarter-second answer than of a one-second one.
Every Decisions response says how long OpenAI's servers spent on it. The median was 65 ms, and 144 ms at the 95th percentile. Our time per call minus OpenAI's own came to 170 ms at the median, about the length of the trip. So OpenAI's 150 ms is believable for someone calling from close to its servers. OpenAI didn't say how long its requests were, where it sent them from, or how many it sent at once, so its numbers and ours were measured in different ways. Luna's responses through the Responses API don't report a server time, so we can't split its 1.05 s the same way.
Where the Decisions API's time goes
One request at a time from a Mac mini in Berlin, the night of 6 October
| time | |
|---|---|
| Our time per request, median | 239 ms |
| OpenAI's reported server time per request, median | 65 ms |
| OpenAI's reported server time, 95th percentile | 144 ms |
| OpenAI's reported server time, 99th percentile | 367 ms |
| Our time minus OpenAI's server time, median | 170 ms |
| The trip to OpenAI's API and back with no model, median, before the run | 166 ms |
| The same trip, median, after the run | 167 ms |
| The fastest of our Decisions requests | 165 ms |
| Our Decisions requests that came back in under 150 ms | 0 |
The slowest calls matter if a person is waiting on the answer. One Decisions request in twenty took longer than 370 ms, and one in a hundred longer than 594 ms. For Jev those were 307 ms and 386 ms, and for Luna 1.60 s and 2.37 s. All of these are from one machine in Berlin, sending one request at a time. They say nothing about other places, or about sending many requests at once.
Accuracy▶ 3:38
Speed doesn't help if the answers are wrong. The Decisions API sorted 93.5% of the messages correctly. Jev, the same night, got 93.9%. On 38 messages Jev was right and the Decisions API was wrong, and on 26 it was the other way round. The Decisions API minus Jev comes to −0.39 points, and the 95% interval of that difference runs from −0.91 to +0.13 points, which includes zero. Across this many messages, a gap that small could easily be chance.
GPT-6 Luna through the Responses API, the same model asked a different way, got 94.2%. We had planned to compare the two routes on speed, not accuracy, so this comparison was made after the run, which makes it post hoc. On 41 messages Luna was right and the Decisions API wrong, and on 21 it was the other way round, nearly two to one. Luna's lead is +0.65 points, with a 95% interval from +0.16 to +1.17, which doesn't include zero. So on this test, the fast way of asking Luna was a little less accurate than the slow one. Because we didn't plan this comparison, we'd treat it as something to check again, on this test, not as a fact about the two routes in general.
Cloudflare's Clef-flash and Clef, from 2 October, are still the most accurate models we've measured on this test, at 95.2% and 95.1%.
Share right, with five examples
A dot per model with its 95% interval; the window starts at 90%, not zero
The same messages, two models at a time
Share right, the first model minus the second; the same night from Berlin
| difference | 95% interval | only the first right | only the second right | |
|---|---|---|---|---|
| Decisions API minus Jev | −0.39 points | −0.91 to +0.13 | 26 | 38 |
| GPT-6 Luna (Responses API) minus Jev | +0.26 points | −0.23 to +0.75 | 31 | 23 |
| GPT-6 Luna (Responses API) minus Decisions API, post hoc | +0.65 points | +0.16 to +1.17 | 41 | 21 |
Side by side▶ 4:41
Here are three of the test messages, picked by a rule we committed before the run: one the Decisions API got right and Jev got wrong, one the other way round, and one the Decisions API got wrong while it was at least as sure as its cut-off for handling a message without a person (the next section explains the cut-off). Each was drawn at random from the messages short enough to read on screen.
"help me with my transfer". The answer on file is a failed transfer. The Decisions API picks that, though it's only 36% sure, and Jev says none of these, which counts as wrong.
"There is an unauthorized fee." Jev and Clef, the larger of Cloudflare's two models, say a card payment fee was charged, which is right. The Decisions API says an extra charge on a statement, and it's 84% sure. That one's wrong.
"How do I change currencies to euros?" The Decisions API, Jev and Clef all say exchanging money in the app. The answer on file is which currencies the bank supports, so all three count as wrong, and the Decisions API was 99% sure of its answer.
These three are not typical. Of all the test messages, the three models all got 2,841 right, and all got 125 wrong. Clef's answers are from its run on 2 October.
Three messages, side by side
Each model's answer and its probability for it, with the five sorted examples; Clef's from 2 October
| message | model | answer | probability | |
|---|---|---|---|---|
| help me with my transfer (on file: failed_transfer) | Decisions API | failed_transfer | 36% | right |
| Jev | none_of_these | 60% | wrong | |
| Clef | failed_transfer | 58% | right | |
| There is an unauthorized fee. (on file: card_payment_fee_charged) | Decisions API | extra_charge_on_statement | 84% | wrong |
| Jev | card_payment_fee_charged | 48% | right | |
| Clef | card_payment_fee_charged | 34% | right | |
| How do I change currencies to euros? (on file: fiat_currency_support) | Decisions API | exchange_via_app | 99% | wrong |
| Jev | exchange_via_app | 94% | wrong | |
| Clef | exchange_via_app | 74% | wrong |
Who gets a person▶ 5:30
That probability for every answer matters if you route messages: let the model handle the ones it's sure of, and send the rest to a person. On a separate set of 573 messages, not the test ones, we set each model's cut-off to aim for about 2% wrong among the messages it handles. For the Decisions API and for Jev, that cut-off came out at 0.99.
On the test, the Decisions API then sent 160 of every thousand messages to a person, and Jev sent 159. The Decisions API was wrong on 2.13% of the messages it handled, and Jev on 2.05%. Both are just over the target.
With the Decisions API, the euros message is the kind that gets through: at 99%, nobody checks it. In all, 55 of its wrong answers were at or above the cut-off. Jev was 94% sure of the same wrong answer, below its cut-off, so Jev would have sent it to a person.
There's one catch, and Jev has it too. The probabilities come in whole percents, and most answers say ninety-nine or a hundred percent: 2,586 of the Decisions API's 3,080 answers did, which is 84%. When it said ninety percent or more, its stated confidence averaged 0.994, and it was right 96.1% of the time. Even the hundred-percent answers were wrong more than once in a hundred: 1.58% of the time for the Decisions API on the test and 1.59% for Jev. On the separate set the figures were 1.06% and 1.86%, so when we aimed for 1% wrong instead, neither model had a setting strict enough to get there.
Who gets a person
Each model's cut-off on its stated confidence, set on the separate set of messages to aim at 2% wrong, then applied to the test
| model | run | cut-off | handled alone | wrong among handled | to a person, per thousand | AUROC |
|---|---|---|---|---|---|---|
| Decisions API | 7 October | 0.99 | 84.0% | 2.13% | 160 | 0.85 |
| Jev | 7 October | 0.99 | 84.1% | 2.05% | 159 | 0.85 |
| GPT-6 Luna (Responses API), written confidence | 7 October | 0.90 | 90.9% | 2.75% | 91 | 0.86 |
| Clef | 2 October | 0.80 | 93.2% | 2.40% | 68 | 0.90 |
| Clef-flash | 2 October | 0.73 | 81.7% | 2.15% | 183 | 0.82 |
| Kev-27B | 2 October | 0.79 | 91.0% | 2.28% | 90 | 0.92 |
| Kev-9B | 2 October | 0.81 | 82.7% | 2.47% | 173 | 0.91 |
GPT-6 Luna through the Responses API returns no probabilities. We asked it to write down a confidence with its answer, which is a different kind of number, so its row is there for reference. Clef, Clef-flash and the two Kev models are from 2 October.
Price▶ 6:41
A thousand of these decisions cost $0.114 on OpenAI's price list. Jev costs $0.078, Luna through the Responses API $0.143, Clef-flash $0.180, and Clef $0.480. Jev charges $0.042 per million tokens, so OpenAI's price per token is about 2.4× Jev's. But OpenAI counts fewer tokens for the same request: 1,136 on average, where Jev counts 1,846. So per decision, the Decisions API comes to about 1.5× Jev's price.
Dollars per thousand decisions
Price list times the tokens each response reports
None of these is a bill. Each is the service's price list times the tokens each response reports, because no response says what it cost. For the Decisions API, Jev, Clef and Clef-flash we count input tokens at each price list. For Luna through the Responses API we count input, cache writes and output at OpenAI's rates. At the price list, the whole Decisions test came to $0.35.
What a thousand decisions cost
Each price list times the tokens each response reports
| model | input tokens per request, average | per thousand decisions |
|---|---|---|
| Jev | 1,846 | $0.078 |
| Decisions API | 1,136 | $0.114 |
| GPT-6 Luna (Responses API) | 1,066 | $0.143 |
| Clef-flash | 1,999 | $0.180 |
| Clef | 1,999 | $0.480 |
Without examples▶ 7:13
When Cloudflare launched Clef, its own table had Clef 14.5 points ahead of Jev on this collection with no examples (94.20 against 79.74, in macro-F1). So we asked six models the bare question ourselves, the message alone with no sorted examples and the 77 reasons as options: Jev, the two Clefs, the Decisions API, and the two open models we've tested before, Kev-27B and Kev-9B.
Clef got 94.2% right (macro-F1 94.20, as in Cloudflare's table), and Jev 80.7%. The Decisions API got 77.7%, the lowest of the six. On our question, with the five examples, it scores 15.8 points higher, the biggest difference of the six, though the two questions also differ in layout and in offering none of these. So without examples, on this test, it did worse than Jev, by 3.0 points.
Without examples, and with five
Share right on the same messages: the message alone, then with five sorted examples
The numbers in this chart
| message alone | with five examples | difference, points | |
|---|---|---|---|
| Clef | 94.2% | 95.1% | +0.8 points |
| Clef-flash | 90.9% | 95.2% | +4.3 points |
| Kev-27B | 86.8% | 93.8% | +7.0 points |
| Kev-9B | 83.3% | 91.3% | +8.0 points |
| Jev | 80.7% | 93.9% | +13.2 points |
| Decisions API | 77.7% | 93.5% | +15.8 points |
The Decisions API's pass without examples ran on 7 October, right after the timed run, and wasn't sent one request at a time, so we quote no times from it. Jev's and the other four models' passes without examples ran on 2 October.
Every model, with and without examples
Share right on the same messages: the message alone (the Decision Index's question), then with five sorted examples
| model | message alone | macro-F1, message alone | with five examples | difference |
|---|---|---|---|---|
| Clef | 94.2% | 94.20 | 95.1% | +0.8 points |
| Clef-flash | 90.9% | 90.85 | 95.2% | +4.3 points |
| Kev-27B | 86.8% | 86.41 | 93.8% | +7.0 points |
| Kev-9B | 83.3% | 83.03 | 91.3% | +8.0 points |
| Jev | 80.7% | 80.00 | 93.9% | +13.2 points |
| Decisions API | 77.7% | 76.95 | 93.5% | +15.8 points |
What this doesn't show▶ 7:54
This is one task, bank messages in English, timed from one place on the night it opened. The test is text only, so it says nothing about images.
We checked whether the Decisions API had memorized the public collection of bank messages our test comes from, and found no sign that it had. Asked 200 messages from the collection's training part with no examples, it got 69.5% of the originals right and 69.0% of the same messages reworded, a gap of +0.5 points with a 95% interval from −3.5 to +5.0. When we sent 200 of the test messages a second time, it gave the same answer to 200 of them.
And OpenAI calls this a public beta. Its terms say a beta service "may be changed at any time without notice", so it may not behave the same next week, or once it's out of beta.
Which to use▶ 8:25
If Jev already works for you, this is no reason to switch. On this test the Decisions API was about as accurate, about as fast from Berlin on a typical request, routed about the same, and cost about 1.5× as much. It's worth a look if you already build on OpenAI, or if you need it to read images, which we didn't test. If accuracy matters most, Cloudflare's Clef-flash was a little ahead on this test, at about 2.3× Jev's price. And the Decisions API is fast, just not 10× faster than OpenAI's regular API from here: 4.4×.
The Decisions API and Jev, the same night
Our test from Berlin
AI Actions (made by the company behind this channel)▶ 9:05
The company behind this channel makes AI Actions, an iPhone app. One of its actions for iOS Shortcuts is called Choose Category, and it runs on Jev. It sorts a piece of text, like a reminder about an invoice, into categories you set. You choose how sure it has to be, and when it isn't sure enough about any category, it says it's unsure instead of guessing. It's linked in the video's description.
Next time▶ 9:30
When the Decisions API is out of beta, we'll run it again. If you sort messages with one of these models, tell us in the video's comments what we should test them on next.