model fatıgue
Measured

Does Cloudflare's Clef really beat Jev?

Published 3 Oct 2026Our test: bank-support messages

Pressing play loads the video from YouTube (Google), in privacy-enhanced mode. Privacy

This is the video written out, with every figure in full, each linked to its source in the table below.

Cloudflare's claim▶ 0:00

Cloudflare's launch post on 1 October says "the market is getting increasingly saturated with decision models". Then it releases two more, Clef and Clef-flash. Their weights are open, and they take the same requests as Jev, TypeSafe's decision model.

A decision model takes a message and a fixed list of answers, like the teams that could handle a request, and returns a probability for each answer instead of writing text. Cloudflare's post compares its two models with Jev and three others in a table of ten tests, and Clef is ahead of Jev on eight of them. The fuller table on Clef's Hugging Face card has 41 tests, and Clef is ahead on 26. Jev is ahead on the other 15, among them several general knowledge and reasoning tests (MMLU-Pro, BBH, GPQA Diamond).

Cloudflare's main table

The ten tests in Cloudflare's launch post, Clef, Clef-flash and Jev (the post also lists three other models)

testClefClef-flashJev
BFCL (case exact)98.4798.7695.75
ToolRet (nDCG@10)69.1966.4365.28
API-Bank (accuracy)91.9393.1188.19
Home appliances (case exact)82.9597.7352.27
When2Call (accuracy)72.3765.5880.97
BANKING77 (macro-F1)94.2090.9379.74
CLINC150+OOS (macro-F1)97.4366.7789.27
BRIGHT (nDCG@10)45.9139.2647.52
Amazon ESCI (macro-F1)57.4857.3955.21
PhishNChips (accuracy)79.6075.0562.55
Cloudflare's numbers, as published. Its Clef scores are its own run of the Decision Index suite; Jev's are the Decision Index's own run.

A second table in the post covers four business workflows from TypeSafe's own evaluation suite, and Clef is ahead of Jev on three of them. We didn't run those.

Cloudflare's workflow table

Four business workflows from TypeSafe's own evaluation suite, as Cloudflare ran them

workflowClefClef-flashJev
Invoice processing64.757.161.8
Customer service76.377.076.0
Security incidents62.961.761.7
Agent trace observability68.569.871.6
Cloudflare's numbers, as published in its launch post. We didn't run these.

The test we checked is in both the post's main table and the card's. BANKING77 is a public set of messages that customers sent to a bank, each filed under one of 77 reasons, like a card that hasn't arrived or a declined transfer. Its test split has 3,080 messages. On it, Cloudflare's table gives Clef 94.20 and Jev 79.74, a lead of 14.5 points. The score is macro-F1, the average over the reasons of how well the model handles each one, counting both the messages it misses and the ones it files there wrongly, and on these messages it comes out within a point of the share each model gets right. If you sort support tickets, a gap that size would make you switch. So is Clef really that far ahead, and is it worth what it costs?

The basics▶ 1:32

Cloudflare released Clef and Clef-flash on 1 October. It built them on open Qwen models, Qwen3.8-27B for Clef and Qwen3.5-9B for Clef-flash, kept those weights frozen, and trained small adapters inside them and a scoring head on top, with what its post calls "our own internal synthetic datasets". The weights are on Hugging Face under the Apache 2.0 licence, so you can run them yourself, and Cloudflare hosts both on Workers AI.

The three models

As their makers publish them

JevClefClef-flash
made byTypeSafeCloudflareCloudflare
built onnot publishedQwen3.8-27BQwen3.5-9B
weightsnot publishedApache 2.0, on Hugging FaceApache 2.0, on Hugging Face
hosted byTypeSafeCloudflare Workers AICloudflare Workers AI
price per million input tokens$0.042$0.240$0.090
output tokensfreenot chargednot charged
imagesnoyesyes
version we testedjev-1.13.0clefclef-flash
Prices read 2 October and again 3 October, unchanged. Workers AI lists no output price for either model, and our runs returned no output tokens.

You pay for what you send in, and the answers come back free on all three. Per token, Clef costs 5.7× what Jev does and Clef-flash 2.1×. Both Cloudflare models can read images, which Jev can't. We didn't test that.

Their test, run ourselves▶ 2:01

Cloudflare's table comes from the Decision Index, a community benchmark that says it is "unofficial and community-maintained; not affiliated with TypeSafe AI". BANKING77 is one of its tests. Cloudflare ran Clef and Clef-flash on edition 0.2.1 itself and labels those scores self-reported. Its number for Jev is the benchmark's own run. So the Clef scores and the Jev score in that table come from different runs.

We sent all 3,080 messages to Clef, Clef-flash and Jev, each request exactly as the Decision Index sends it: no context, the instruction "Classify the banking intent of this user request:" with the message, and the 77 reasons as the options. We sent the same requests to Kev-27B and Kev-9B, which come up later.

Their test: the Decision Index's request

The message alone and the reasons as options, no examples. All five models, the same messages, sent the same day.

modelmacro-F195% intervalshare rightCloudflare's table
Clef94.2093.29 to 94.9494.2%94.20
Clef-flash90.8589.78 to 91.7490.9%90.93
Kev-27B86.4185.14 to 87.4186.8%not in it
Kev-9B83.0381.57 to 84.1083.3%84.83, an older Kev-9B
Jev80.0078.61 to 81.0880.7%79.74
Our run of 2 October. Cloudflare's column is its launch table: Clef and Clef-flash from Cloudflare's own run, Jev from the Decision Index's.

Clef scored 94.20, the same as Cloudflare's table. Clef-flash scored 90.85 against the table's 90.93, and Jev 80.00 against 79.74. All three sit inside our 95% interval. Kev-9B's 84.83 in that table is an older checkpoint of it, so it doesn't compare with ours. So the table holds: on this test Clef is 14.2 points ahead of Jev in macro-F1, 13.5 points in the share right.

How we ran it: we fixed the procedure before sending a single test message, covering which messages, which request for each model, how answers are scored, how each model's cut-off for sending a message to a person is set, and how costs are counted. Clef, Clef-flash and Kev each got a short addendum on 2 October, written before their test passes, that changes only the model name and the address the request goes to. Kev's servers had answered a timing probe of the first forty test messages just before its addendum was written; nothing was set from it, and its answers weren't scored.

Their request is the Decision Index's own, copied byte for byte from its code. We sent it six at a time from a laptop in Berlin, so we don't use its times for speed. The one place they come in is Kev's cost projection below, which says so.

Our request carries the message, the five most similar messages from BANKING77's training split with their reasons, and 78 options: the 77 reasons plus "none of these", which a real queue needs and which is always wrong on this test. One fixed embedding model finds the examples in a pool of 9,405 labelled training messages, with test duplicates and our held-out messages removed, so no test message ever appears as its own example. Every model gets the same examples. We sent it to Jev, Clef and Clef-flash on the afternoon of 2 October, one request at a time from a Mac mini on a home connection in Berlin, Jev first. Kev-27B and Kev-9B ran the same evening on H100s we rented on Modal in Europe.

The cut-offs for sending a message to a person were set on 573 training messages held out from the example pool, before the test, aiming at 2% wrong answers among the messages a model routes on its own. Costs are each provider's list price times the input tokens it counted on every call, and none of the three hosted models charges for output. For Cloudflare, the neurons it reported on every call add up to the same dollars. Kev's cost is Modal's bill for our GPU time.

Five examples▶ 2:52

The Decision Index's request gives the model the customer's message and the list of reasons, and nothing else. If you sort support tickets for real, you also have old tickets that someone has already sorted. You can look up the five most like the new one and put them in the request with their right answers, which is called few-shot prompting. That is our request.

Take one test message: "Can I be given a new passcode?", filed under forgotten passcode. On their request, Jev picks change PIN at 0.63, and Clef picks forgotten passcode at 0.90. The five most similar sorted messages ("Need a new passcode.", "Where can I get a new passcode?" and three more) were all filed under forgotten passcode. With them in the request, Jev picks forgotten passcode at 1.00. Jev rounds its probabilities to two decimals.

Three messages, three models

Each model's answer and its probability on their request, then with five examples. ✓ marks the reason the dataset files it under.

messagemodeltheir requestwith five examples
“Can I be given a new passcode?” (filed under passcode forgotten)Jevchange pin (0.63)passcode forgotten (1.00) ✓
Clef-flashpasscode forgotten (0.77) ✓passcode forgotten (0.89) ✓
Clefpasscode forgotten (0.90) ✓passcode forgotten (0.96) ✓
“How do I contact customer support about a transfer?” (filed under declined transfer)Jevpending transfer (0.48)none of these (0.81)
Clef-flashtransfer not received by recipient (0.10)declined transfer (0.48) ✓
Clefbalance not updated after bank transfer (0.28)declined transfer (0.27) ✓
“Any fee for topping up?” (filed under top up by card charge)Jevtop up by card charge (0.64) ✓top up by card charge (0.44) ✓
Clef-flashtop up by bank transfer charge (0.40)top up by card charge (0.59) ✓
Cleftop up by card charge (0.67) ✓top up by bank transfer charge (0.60)
Our run of 2 October, probabilities as each model returned them. The messages were picked by a rule written down before it ran (first a message Jev gets wrong without examples and right with them, the commonest kind of change, then one each way), not by hand.

Across the test, 437 of Jev's answers went from wrong to right and 30 the other way. To check that the examples and not our different wording carry that, an earlier run of our request with the examples taken out had Jev at 78.3%, a little under what it got on the Decision Index's request.

Without examples, and with five

Share right on the same messages: their request (the message alone), then ours with five examples

their requestwith five examples
75%80%85%90%95%100%share rightdifference, pointsClefClef: their request 94.2% (run 2 Oct 2026, 17:03–21:41 CEST), with five examples 95.1% (run 2 Oct 2026, 15:49–16:54 CEST) · Model Fatigue+0.8Clef-flashClef-flash: their request 90.9% (run 2 Oct 2026, 17:03–21:41 CEST), with five examples 95.2% (run 2 Oct 2026, 15:49–16:54 CEST) · Model Fatigue+4.3Kev-27BKev-27B: their request 86.8% (run 2 Oct 2026, 17:03–21:41 CEST), with five examples 93.8% (run 2 Oct 2026, 21:05–22:00 CEST) · Model Fatigue+7.0Kev-9BKev-9B: their request 83.3% (run 2 Oct 2026, 17:03–21:41 CEST), with five examples 91.3% (run 2 Oct 2026, 21:05–22:00 CEST) · Model Fatigue+8.0JevJev: their request 80.7% (run 2 Oct 2026, 17:03–21:41 CEST), with five examples 93.9% (run 2 Oct 2026, 15:49–16:54 CEST) · Model Fatigue+13.2
75%80%85%90%95%100%share rightdifference, pointsClefClef: their request 94.2% (run 2 Oct 2026, 17:03–21:41 CEST), with five examples 95.1% (run 2 Oct 2026, 15:49–16:54 CEST) · Model Fatigue+0.8Clef-flashClef-flash: their request 90.9% (run 2 Oct 2026, 17:03–21:41 CEST), with five examples 95.2% (run 2 Oct 2026, 15:49–16:54 CEST) · Model Fatigue+4.3Kev-27BKev-27B: their request 86.8% (run 2 Oct 2026, 17:03–21:41 CEST), with five examples 93.8% (run 2 Oct 2026, 21:05–22:00 CEST) · Model Fatigue+7.0Kev-9BKev-9B: their request 83.3% (run 2 Oct 2026, 17:03–21:41 CEST), with five examples 91.3% (run 2 Oct 2026, 21:05–22:00 CEST) · Model Fatigue+8.0JevJev: their request 80.7% (run 2 Oct 2026, 17:03–21:41 CEST), with five examples 93.9% (run 2 Oct 2026, 15:49–16:54 CEST) · Model Fatigue+13.2
Our runs of 2 October. The two requests also differ in wording and in offering none of these; Jev without examples on our wording scored 78.3% in an earlier run.
The numbers in this chart
their requestwith five examplesdifference, points
Clef94.2%95.1%+0.8 points
Clef-flash90.9%95.2%+4.3 points
Kev-27B86.8%93.8%+7.0 points
Kev-9B83.3%91.3%+8.0 points
Jev80.7%93.9%+13.2 points

With five examples Jev got 93.9% right, Clef 95.1% and Clef-flash 95.2%. Clef barely needed the examples. Its lead in the share right, 13.5 points without them, is down to 1.1 points. That point isn't chance. On the same messages Clef got 54 right that Jev got wrong, and Jev 19 that Clef got wrong, and the 95% interval on the difference runs from +0.62 to +1.69 points.

Our test: with five examples

The same messages, each with its five most similar sorted training messages, and none of these as an extra option

modelshare right95% intervalminus Jev, points95% intervalnone of these
Clef-flash95.2%94.3 to 95.9+1.23+0.68 to +1.820
Clef95.1%94.2 to 95.8+1.14+0.62 to +1.690
Jev93.9%93.0 to 94.714
Kev-27B93.8%92.9 to 94.6−0.10−0.55 to +0.3621
Kev-9B91.3%90.3 to 92.3−2.60−3.28 to −1.92110
Our runs of 2 October. Differences are on the same messages, against Jev's pass that day. None of these is always wrong on this test.

Here is one each way. "How do I contact customer support about a transfer?" is filed under declined transfer. Clef gets it at 0.27 and Jev picks none of these. "Any fee for topping up?" is filed under the fee for topping up by card. Jev gets it at 0.44, and Clef picks the fee for topping up by bank transfer. At the cut-offs in "Who gets a person", all four of those answers would have gone to a person. Clef and Jev split 73 messages this way, one right and the other wrong. On 47 of them both answers were below their cut-offs, so if a person checks the unsure ones, they'd have seen most of these messages whichever model you used.

The bill▶ 5:38

In our run with five examples, sorting a million messages would cost $77.53 on Jev, $179.91 on Clef-flash and $479.75 on Clef. Most of the gap is the price per token. The rest is that Cloudflare's models count 8% more input tokens than Jev for the same request, 1,999 against 1,846.

Dollars per thousand messages

Our test with five examples

$0$0.2$0.4$0.6$0.8dollars per thousand messagesJevJev: $0.0775 · Model Fatigue, run 2 Oct 2026, 15:49–16:54 CEST$0.0775Clef-flashClef-flash: $0.1799 · Model Fatigue, run 2 Oct 2026, 15:49–16:54 CEST$0.1799ClefClef: $0.4798 · Model Fatigue, run 2 Oct 2026, 15:49–16:54 CEST$0.4798Kev-9BKev-9B: $0.6211 · Model Fatigue, run 2 Oct 2026, 21:05–22:00 CEST$0.6211Kev-27BKev-27B: $0.7295 · Model Fatigue, run 2 Oct 2026, 21:05–22:00 CEST$0.7295
$0$0.2$0.4$0.6$0.8dollars per thousand messagesJevJev: $0.0775 · Model Fatigue, run 2 Oct 2026, 15:49–16:54 CEST$0.0775Clef-flashClef-flash: $0.1799 · Model Fatigue, run 2 Oct 2026, 15:49–16:54 CEST$0.1799ClefClef: $0.4798 · Model Fatigue, run 2 Oct 2026, 15:49–16:54 CEST$0.4798Kev-9BKev-9B: $0.6211 · Model Fatigue, run 2 Oct 2026, 21:05–22:00 CEST$0.6211Kev-27BKev-27B: $0.7295 · Model Fatigue, run 2 Oct 2026, 21:05–22:00 CEST$0.7295
Filled: list price times tokens counted. Outlined: our rented GPU time, one request at a time, before start-up and idle time.
The numbers in this chart
dollars per thousand messages
Jev$0.0775
Clef-flash$0.1799
Clef$0.4798
Kev-9B$0.6211
Kev-27B$0.7295

What it costs

Our test with five examples

modelper thousand messagesper millioninput tokens per requestagainst Jev
Jev$0.0775$77.531,846
Clef-flash$0.1799$179.911,9992.3×
Clef$0.4798$479.751,9996.2×
Kev-27B$0.7295GPU time9759.4×
Kev-9B$0.6211GPU time975
List price times the input tokens each provider counted; per million is our arithmetic. The Kev rows are Modal's bill for our rented H100s, one request at a time, before start-up and idle time; Kev counts tokens with its own tokenizer.

So with examples, Clef beats Jev by about one message in a hundred for 6.2× the price, and Clef-flash gets the same point for 2.3× the price. Without examples the requests are shorter: per thousand messages, their request cost $0.448 on Clef, $0.168 on Clef-flash and $0.072 on Jev.

Speed▶ 6:17

Cloudflare's table also says its models answer faster than Jev. It gives median times of 209.3 ms for Clef, 38.8 ms for Clef-flash and 524.1 ms for Jev, which makes Clef 2.5× as fast and Clef-flash 13.5× by our arithmetic. Those times aren't measured the same way. The Decision Index says its Jev time is a "network round-trip from our lab" and not comparable with times measured on a model running on its own hardware. Cloudflare's leaderboard marks Clef's scores and latency as self-reported, with the latency "not measured on the board's hardware". Its post counts on its network for the rest: "our Clef models are hosted on Workers AI", "leading to low network latency".

So we timed every call in our run with five examples, from Berlin. Our requests entered Cloudflare's network at its Berlin edge. From somewhere else the order could change.

Time per request, from Berlin

The bar runs to the median; the ticks mark p95 and p99. One request at a time.

median95th and 99th percentile
0 s0.5 s1 s1.5 s2 s2.5 sseconds per requestClef-flashClef-flash: median 196 ms, 95th percentile 1,014 ms, 99th 2,172 ms · Model Fatigue, run 2 Oct 2026, 15:49–16:54 CEST196 msJevJev: median 235 ms, 95th percentile 335 ms, 99th 442 ms · Model Fatigue, run 2 Oct 2026, 15:49–16:54 CEST235 msKev-9BKev-9B: median 437 ms, 95th percentile 548 ms, 99th 665 ms · Model Fatigue, run 2 Oct 2026, 21:05–22:00 CEST437 msKev-27BKev-27B: median 529 ms, 95th percentile 551 ms, 99th 598 ms · Model Fatigue, run 2 Oct 2026, 21:05–22:00 CEST529 msClefClef: median 562 ms, 95th percentile 1,147 ms, 99th 2,001 ms · Model Fatigue, run 2 Oct 2026, 15:49–16:54 CEST562 ms
0 s0.5 s1 s1.5 s2 s2.5 sseconds per requestClef-flashClef-flash: median 196 ms, 95th percentile 1,014 ms, 99th 2,172 ms · Model Fatigue, run 2 Oct 2026, 15:49–16:54 CEST196 msJevJev: median 235 ms, 95th percentile 335 ms, 99th 442 ms · Model Fatigue, run 2 Oct 2026, 15:49–16:54 CEST235 msKev-9BKev-9B: median 437 ms, 95th percentile 548 ms, 99th 665 ms · Model Fatigue, run 2 Oct 2026, 21:05–22:00 CEST437 msKev-27BKev-27B: median 529 ms, 95th percentile 551 ms, 99th 598 ms · Model Fatigue, run 2 Oct 2026, 21:05–22:00 CEST529 msClefClef: median 562 ms, 95th percentile 1,147 ms, 99th 2,001 ms · Model Fatigue, run 2 Oct 2026, 15:49–16:54 CEST562 ms
Our run of 2 October from a Mac mini in Berlin. Kev ran on our own H100s in Europe.
The numbers in this chart

Time per request from Berlin

Our run with five examples, one request at a time

Ours: wall time per call from a Mac mini in Berlin, Cloudflare's models entering at its Berlin edge, Kev on our own H100s in Europe. Cloudflare's: self-reported for Clef, the Decision Index's round trip from its lab for Jev. The two kinds of time aren't comparable.

At the median, Clef-flash answered in 196 ms, Jev in 235 ms and Clef in 562 ms. Clef-flash had the widest spread, with the slowest calls of any model at the 99th percentile: 139 of its 159 calls over a second came in the first quarter of its pass. We ran it once and didn't look into why. For scale, a bare request with no model behind it took 171 ms to TypeSafe's host from the same machine at the start of the run.

Who gets a person▶ 7:23

In a real queue you'd let the model route the messages it's sure about and send the rest to a person. For each model we set a cut-off on the probability of its answer, using held-out training messages and aiming at 2% wrong among the messages it routes on its own. Then we ran the test with examples.

Routing at each model's cut-off

Each model routes the messages above its cut-off and sends the rest to a person

modelcut-offrouted on its ownwrong among routed (95% interval)to a person, per thousandwrong and routed, per thousandAUROCcalibration error
Clef0.8093.2%2.40% (1.90 to 3.03)6822.40.900.030
Kev-27B0.7991.0%2.28% (1.79 to 2.90)9020.80.920.014
Jev0.9984.4%2.04% (1.56 to 2.66)15617.20.840.035
Kev-9B0.8182.7%2.47% (1.94 to 3.15)17320.50.910.038
Clef-flash0.7381.7%2.15% (1.65 to 2.79)18317.50.820.143
Our runs of 2 October. Cut-offs set on held-out training messages before the test, aiming at 2% wrong among the routed. AUROC: how well a model's probabilities rank its right answers above its wrong ones (one is perfect, a half is chance). Calibration error: how far its probabilities sit from how often it is right (lower is closer).

Per thousand messages, Clef sent 68 to a person, Jev 156 and Clef-flash 183. Among the messages each routed on its own, all three came out between two and two and a half percent wrong: Clef 2.40%, Jev 2.04% and Clef-flash 2.15%.

Per thousand messages, Clef costs $0.40 more than Jev and sends 88.3 fewer messages to a person, so each check it saves costs $0.0046 in API fees. At $15 an hour, that buys 1.1 s of someone's time. Unless a person can check a message in about a second, Clef is the cheaper one once you add the checking to the API bill.

Clef routes more on its own, so it also lets more wrong answers through: 22.4 per thousand against Jev's 17.2. The extra messages it routes are harder. Of the 318 that Clef routed and Jev sent to a person, Clef got 31 wrong.

How far can you take each model's probability at face value? Clef's come close to how often it turns out right (calibration error 0.030), and it is the best of the three at ranking its right answers above its wrong ones (AUROC 0.90, against 0.82 for Clef-flash). Jev's are close on average but too high whenever it is less than sure: its 481 answers below its cut-off of 0.99 averaged 0.84 and were right 72% of the time. Clef-flash's are too low. On the 736 answers it gave between 0.7 and 0.8 it was right 95% of the time. Check those two on your own data before you take their numbers as they come.

Probability given against the share right

Each model's answers grouped by the probability it gave. On the dashed line a model is right exactly as often as it says.

where stated confidence equals the share righta tenth of the range, sized by how many answers fell in it
Kev-27BECE 0.0140.20.20.60.61.01.0Kev-27B: 9 answers stated about 0.25, 0.22 of them right · Model Fatigue, run 2 Oct 2026, 21:05–22:00 CESTKev-27B: 25 answers stated about 0.37, 0.36 of them right · Model Fatigue, run 2 Oct 2026, 21:05–22:00 CESTKev-27B: 51 answers stated about 0.45, 0.43 of them right · Model Fatigue, run 2 Oct 2026, 21:05–22:00 CESTKev-27B: 55 answers stated about 0.55, 0.45 of them right · Model Fatigue, run 2 Oct 2026, 21:05–22:00 CESTKev-27B: 62 answers stated about 0.65, 0.56 of them right · Model Fatigue, run 2 Oct 2026, 21:05–22:00 CESTKev-27B: 75 answers stated about 0.75, 0.77 of them right · Model Fatigue, run 2 Oct 2026, 21:05–22:00 CESTKev-27B: 204 answers stated about 0.86, 0.86 of them right · Model Fatigue, run 2 Oct 2026, 21:05–22:00 CESTKev-27B: 2,598 answers stated about 0.98, 0.99 of them right · Model Fatigue, run 2 Oct 2026, 21:05–22:00 CESTClefECE 0.0300.20.20.60.61.01.0Clef: 5 answers stated about 0.28, 0.40 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:54 CESTClef: 13 answers stated about 0.36, 0.69 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:54 CESTClef: 23 answers stated about 0.45, 0.57 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:54 CESTClef: 38 answers stated about 0.55, 0.50 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:54 CESTClef: 55 answers stated about 0.65, 0.56 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:54 CESTClef: 69 answers stated about 0.76, 0.70 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:54 CESTClef: 187 answers stated about 0.86, 0.83 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:54 CESTClef: 2,690 answers stated about 0.96, 0.99 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:54 CESTJevECE 0.0350.20.20.60.61.01.0Jev: 1 answers stated about 0.28, 0.00 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:54 CESTJev: 3 answers stated about 0.35, 0.67 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:54 CESTJev: 17 answers stated about 0.45, 0.35 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:54 CESTJev: 37 answers stated about 0.55, 0.49 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:54 CESTJev: 39 answers stated about 0.64, 0.44 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:54 CESTJev: 46 answers stated about 0.75, 0.63 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:54 CESTJev: 81 answers stated about 0.85, 0.72 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:54 CESTJev: 2,856 answers stated about 1.00, 0.97 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:54 CESTKev-9BECE 0.0380.20.20.60.61.01.0Kev-9B: 2 answers stated about 0.29, 0.50 of them right · Model Fatigue, run 2 Oct 2026, 21:05–22:00 CESTKev-9B: 21 answers stated about 0.36, 0.24 of them right · Model Fatigue, run 2 Oct 2026, 21:05–22:00 CESTKev-9B: 88 answers stated about 0.46, 0.38 of them right · Model Fatigue, run 2 Oct 2026, 21:05–22:00 CESTKev-9B: 102 answers stated about 0.55, 0.47 of them right · Model Fatigue, run 2 Oct 2026, 21:05–22:00 CESTKev-9B: 115 answers stated about 0.65, 0.64 of them right · Model Fatigue, run 2 Oct 2026, 21:05–22:00 CESTKev-9B: 180 answers stated about 0.75, 0.81 of them right · Model Fatigue, run 2 Oct 2026, 21:05–22:00 CESTKev-9B: 440 answers stated about 0.86, 0.92 of them right · Model Fatigue, run 2 Oct 2026, 21:05–22:00 CESTKev-9B: 2,132 answers stated about 0.96, 0.99 of them right · Model Fatigue, run 2 Oct 2026, 21:05–22:00 CESTClef-flashECE 0.1430.20.20.60.61.01.0Clef-flash: 7 answers stated about 0.26, 0.14 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:54 CESTClef-flash: 12 answers stated about 0.35, 0.33 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:54 CESTClef-flash: 59 answers stated about 0.46, 0.68 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:54 CESTClef-flash: 87 answers stated about 0.56, 0.78 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:54 CESTClef-flash: 241 answers stated about 0.66, 0.87 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:54 CESTClef-flash: 736 answers stated about 0.76, 0.95 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:54 CESTClef-flash: 1,339 answers stated about 0.85, 0.98 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:54 CESTClef-flash: 599 answers stated about 0.92, 1.00 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:54 CEST
Kev-27BECE 0.0140.20.20.60.61.01.0Kev-27B: 9 answers stated about 0.25, 0.22 of them right · Model Fatigue, run 2 Oct 2026, 21:05–22:00 CESTKev-27B: 25 answers stated about 0.37, 0.36 of them right · Model Fatigue, run 2 Oct 2026, 21:05–22:00 CESTKev-27B: 51 answers stated about 0.45, 0.43 of them right · Model Fatigue, run 2 Oct 2026, 21:05–22:00 CESTKev-27B: 55 answers stated about 0.55, 0.45 of them right · Model Fatigue, run 2 Oct 2026, 21:05–22:00 CESTKev-27B: 62 answers stated about 0.65, 0.56 of them right · Model Fatigue, run 2 Oct 2026, 21:05–22:00 CESTKev-27B: 75 answers stated about 0.75, 0.77 of them right · Model Fatigue, run 2 Oct 2026, 21:05–22:00 CESTKev-27B: 204 answers stated about 0.86, 0.86 of them right · Model Fatigue, run 2 Oct 2026, 21:05–22:00 CESTKev-27B: 2,598 answers stated about 0.98, 0.99 of them right · Model Fatigue, run 2 Oct 2026, 21:05–22:00 CESTClefECE 0.0300.20.20.60.61.01.0Clef: 5 answers stated about 0.28, 0.40 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:54 CESTClef: 13 answers stated about 0.36, 0.69 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:54 CESTClef: 23 answers stated about 0.45, 0.57 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:54 CESTClef: 38 answers stated about 0.55, 0.50 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:54 CESTClef: 55 answers stated about 0.65, 0.56 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:54 CESTClef: 69 answers stated about 0.76, 0.70 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:54 CESTClef: 187 answers stated about 0.86, 0.83 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:54 CESTClef: 2,690 answers stated about 0.96, 0.99 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:54 CESTJevECE 0.0350.20.20.60.61.01.0Jev: 1 answers stated about 0.28, 0.00 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:54 CESTJev: 3 answers stated about 0.35, 0.67 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:54 CESTJev: 17 answers stated about 0.45, 0.35 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:54 CESTJev: 37 answers stated about 0.55, 0.49 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:54 CESTJev: 39 answers stated about 0.64, 0.44 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:54 CESTJev: 46 answers stated about 0.75, 0.63 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:54 CESTJev: 81 answers stated about 0.85, 0.72 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:54 CESTJev: 2,856 answers stated about 1.00, 0.97 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:54 CESTKev-9BECE 0.0380.20.20.60.61.01.0Kev-9B: 2 answers stated about 0.29, 0.50 of them right · Model Fatigue, run 2 Oct 2026, 21:05–22:00 CESTKev-9B: 21 answers stated about 0.36, 0.24 of them right · Model Fatigue, run 2 Oct 2026, 21:05–22:00 CESTKev-9B: 88 answers stated about 0.46, 0.38 of them right · Model Fatigue, run 2 Oct 2026, 21:05–22:00 CESTKev-9B: 102 answers stated about 0.55, 0.47 of them right · Model Fatigue, run 2 Oct 2026, 21:05–22:00 CESTKev-9B: 115 answers stated about 0.65, 0.64 of them right · Model Fatigue, run 2 Oct 2026, 21:05–22:00 CESTKev-9B: 180 answers stated about 0.75, 0.81 of them right · Model Fatigue, run 2 Oct 2026, 21:05–22:00 CESTKev-9B: 440 answers stated about 0.86, 0.92 of them right · Model Fatigue, run 2 Oct 2026, 21:05–22:00 CESTKev-9B: 2,132 answers stated about 0.96, 0.99 of them right · Model Fatigue, run 2 Oct 2026, 21:05–22:00 CESTClef-flashECE 0.1430.20.20.60.61.01.0Clef-flash: 7 answers stated about 0.26, 0.14 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:54 CESTClef-flash: 12 answers stated about 0.35, 0.33 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:54 CESTClef-flash: 59 answers stated about 0.46, 0.68 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:54 CESTClef-flash: 87 answers stated about 0.56, 0.78 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:54 CESTClef-flash: 241 answers stated about 0.66, 0.87 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:54 CESTClef-flash: 736 answers stated about 0.76, 0.95 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:54 CESTClef-flash: 1,339 answers stated about 0.85, 0.98 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:54 CESTClef-flash: 599 answers stated about 0.92, 1.00 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:54 CEST
Our runs of 2 October, with five examples. Dot size: how many answers fell in that tenth. Groups below a fifth are left off the chart; the table lists them.
The numbers in this chart
statedanswersmean statedshare right
Kev-27B0.1 to 0.210.191.00
Kev-27B0.2 to 0.390.250.22
Kev-27B0.3 to 0.4250.370.36
Kev-27B0.4 to 0.5510.450.43
Kev-27B0.5 to 0.6550.550.45
Kev-27B0.6 to 0.7620.650.56
Kev-27B0.7 to 0.8750.750.77
Kev-27B0.8 to 0.92040.860.86
Kev-27B0.9 to 1.02,5980.980.99
Clef0.2 to 0.350.280.40
Clef0.3 to 0.4130.360.69
Clef0.4 to 0.5230.450.57
Clef0.5 to 0.6380.550.50
Clef0.6 to 0.7550.650.56
Clef0.7 to 0.8690.760.70
Clef0.8 to 0.91870.860.83
Clef0.9 to 1.02,6900.960.99
Jev0.2 to 0.310.280.00
Jev0.3 to 0.430.350.67
Jev0.4 to 0.5170.450.35
Jev0.5 to 0.6370.550.49
Jev0.6 to 0.7390.640.44
Jev0.7 to 0.8460.750.63
Jev0.8 to 0.9810.850.72
Jev0.9 to 1.02,8561.000.97
Kev-9B0.2 to 0.320.290.50
Kev-9B0.3 to 0.4210.360.24
Kev-9B0.4 to 0.5880.460.38
Kev-9B0.5 to 0.61020.550.47
Kev-9B0.6 to 0.71150.650.64
Kev-9B0.7 to 0.81800.750.81
Kev-9B0.8 to 0.94400.860.92
Kev-9B0.9 to 1.02,1320.960.99
Clef-flash0.2 to 0.370.260.14
Clef-flash0.3 to 0.4120.350.33
Clef-flash0.4 to 0.5590.460.68
Clef-flash0.5 to 0.6870.560.78
Clef-flash0.6 to 0.72410.660.87
Clef-flash0.7 to 0.87360.760.95
Clef-flash0.8 to 0.91,3390.850.98
Clef-flash0.9 to 1.05990.921.00

Kev 1.0▶ 9:03

Kev, from the developer Jared Palmer, is one more open-weight model that takes Jev's requests. Kev 1.0 came out on 1 October, the same day as Clef, and gathers models he had already published; its release notes say "Nothing in it is newly trained." Kev-27B is built on Qwen3.8-27B, the same base as Clef, and Kev-9B on Qwen3.5-9B-Base. Its model cards list BANKING77 among their training data, and its code trains on the training split and leaves the test split out. We ran both on H100s we rented on Modal in Europe, the same two ways as the others.

With five examples, Kev-27B got 93.8% right, level with Jev (−0.10 points, interval −0.55 to +0.36) and behind Clef (−1.23 points, interval −1.75 to −0.71). Kev-9B got 91.3%, mostly because it answered none of these 110 times on messages that have a real reason; with that option set aside its top real reason was right 93.6% of the time. On the Decision Index's request Kev-27B scored 86.41, between Jev and Clef. Two teams started from the same base model, and Cloudflare's version does far better when it sees only the message.

Kev-27B's probabilities matched how often it was right more closely than any other model's here (calibration error 0.014), on a dataset whose training split it learned from. At the same aim it sent 90 messages per thousand to a person.

We paid for Kev by the hour of GPU, not per token, so what a message costs depends on how many you send the card at once. Sent one at a time, as we did, it came to $0.7295 per thousand messages, 9.4× Jev's price, before start-up and idle time, which were 35% of Kev-27B's bill. Modal's bill works out at $4.92 per container-hour. From how fast Kev got through the Decision Index pass, sent six at a time from the laptop, six at a time would come to roughly $0.157 per thousand. That pass used shorter requests than our test, 853 input tokens a call against 975 by Kev's count, and we didn't run our test that way, so read it as a rough guide. From Berlin its median answer took 529 ms, of which the model itself took 110 ms.

What Kev cost on our GPUs

Rented H100s on Modal, billed by the second

Kev-27BKev-9B
Modal's bill per container-hour$4.92$4.91
per thousand messages, one at a time, as run$0.7295$0.6211
per thousand, six at a time (rough projection)$0.157$0.112
start-up and idle, share of the bill35%46%
median time from Berlin529 ms437 ms
of which the model itself110 ms35 ms
Our arithmetic on Modal's own bill, per pass. The six-at-a-time projection comes from how fast the Decision Index pass went, sent from the laptop on shorter requests than our test, so it is a rough guide; we didn't run our test that way. The model's own time is what Kev's server reports for each request.

Kev's own release notes put Kev-27B slightly behind Jev on datasets no Kev trained on: 52.3 against 54.0 on the release's index for datasets held out from training (breadth).

What this doesn't settle▶ 10:40

Cloudflare says it trained Clef on top of Qwen with its own synthetic data. Neither its post nor the model cards say whether BANKING77, or anything made from it, was in any of the training. Kev's cards list it. TypeSafe's models page says Jev isn't trained on customer requests and doesn't list what it was trained on.

We ran a check on Clef and Clef-flash that compares 200 training messages they might have seen with reworded copies of them. Both did better on the originals, but so does a vote of the five nearest labelled examples, which can't remember anything (+15.0 points), and both gaps were smaller than that. So the check can't tell memory from the cost of rewording. That matters most for the result without examples, where Clef's lead is big.

Have they seen these messages before?

Training messages against reworded copies of the same messages, without examples

modeltraining messagesreworded copiesdifference
Clef94.5%87.0%+7.5 points
Clef-flash95.5%83.5%+12.0 points
vote of the five nearest labelled examples+15.0 points
Our check of 2 October on 200 messages; the vote, which can't remember anything, is from M37's jev-2 run. A model that remembered the messages would score much higher on the originals than the vote does; neither did, so the check is inconclusive.

We tested one task, in English. In the Hugging Face table Jev is ahead of Clef on 15 of the 41 tests, so other tasks may come out differently. We didn't test images. On a second pass over the same messages, Clef and Clef-flash gave the same answer to 200 of 200. And this is one day of calls, from Berlin.

Which to use▶ 11:33

If you have no sorted messages to show the model, use Clef. It was 14.2 points ahead of Jev on the Decision Index's request, and Clef-flash 10.8 points, for 38% of Clef's price.

If you do have sorted messages, Clef, Clef-flash and Jev come within about a point of each other, and the other differences decide it.

For the smallest bill, use Jev, at about a sixth of Clef's price per message.

If a person checks the answers the model isn't sure about, Clef sends them fewer than half as many messages as Jev. Unless a check costs you less than $0.0046, that makes Clef the cheaper one overall, with a few more wrong answers getting through.

If no person checks the answers, Clef-flash got as many right as Clef in our run, for 38% of Clef's price and 2.3× Jev's.

For speed from Berlin, Clef-flash and Jev were close at the median, and Clef was the slowest.

If you want to run one on your own hardware, Clef, Clef-flash and Kev are all open, and Clef and Clef-flash were the most accurate of them here. Kev-27B matched Jev with examples, but it learned from this dataset's training split, and its own release notes put it slightly behind Jev on data no Kev trained on. What it costs depends on how you host it.

Every number

These are all 438 figures behind the video and this page, grouped by whose they are, with the page each came from and when we read it. Figures marked ⟳ can move. When a re-read finds a change, the new value shows next to the one from the video.

Model Fatigue

Our run with five examples: Jev, Clef and Clef-flash on the same day · run 2 Oct 2026, 15:49–16:54 CEST
BANKING77 test messages every model answered3,080
Jev, the version that answeredjev-1.13.0
Jev, share right with five examples93.9%
Jev, share right, 95% interval, low93.0
Jev, share right, 95% interval, high94.7
Jev, messages sorted correctly with five examples2,893
Jev, times it answered none-of-these (always wrong on this test)14
Jev, median time per request from Berlin235 ms
Jev, 95th percentile time per request from Berlin335 ms
Jev, 99th percentile time per request from Berlin442 ms
Jev, dollars per 1,000 messages (price list × input tokens counted)$0.0775
Jev, calibration error of its probabilities (ECE; lower is closer)0.035
Jev, how well its probabilities separate its right answers from its wrong ones (AUROC)0.84
Jev, cut-off on its top probability, set on the held-out training messages0.99
Jev, share of messages routed without a person84.4%
Jev, wrong among the messages it routed itself2.04%
Jev, wrong among routed, 95% interval, low1.56
Jev, wrong among routed, 95% interval, high2.66
Jev, messages sent to a person481
Jev, wrong answers routed without a person53
Clef, share right with five examples95.1%
Clef, share right, 95% interval, low94.2
Clef, share right, 95% interval, high95.8
Clef, messages sorted correctly with five examples2,928
Clef, times it answered none-of-these (always wrong on this test)0
Clef, median time per request from Berlin562 ms
Clef, 95th percentile time per request from Berlin1,147 ms
Clef, 99th percentile time per request from Berlin2,001 ms
Clef, dollars per 1,000 messages (price list × input tokens counted)$0.4798
Clef, calibration error of its probabilities (ECE; lower is closer)0.030
Clef, how well its probabilities separate its right answers from its wrong ones (AUROC)0.90
Clef, cut-off on its top probability, set on the held-out training messages0.80
Clef, share of messages routed without a person93.2%
Clef, wrong among the messages it routed itself2.40%
Clef, wrong among routed, 95% interval, low1.90
Clef, wrong among routed, 95% interval, high3.03
Clef, messages sent to a person209
Clef, wrong answers routed without a person69
Clef-flash, share right with five examples95.2%
Clef-flash, share right, 95% interval, low94.3
Clef-flash, share right, 95% interval, high95.9
Clef-flash, messages sorted correctly with five examples2,931
Clef-flash, times it answered none-of-these (always wrong on this test)0
Clef-flash, median time per request from Berlin196 ms
Clef-flash, 95th percentile time per request from Berlin1,014 ms
Clef-flash, 99th percentile time per request from Berlin2,172 ms
Clef-flash, dollars per 1,000 messages (price list × input tokens counted)$0.1799
Clef-flash, calibration error of its probabilities (ECE; lower is closer)0.143
Clef-flash, how well its probabilities separate its right answers from its wrong ones (AUROC)0.82
Clef-flash, cut-off on its top probability, set on the held-out training messages0.73
Clef-flash, share of messages routed without a person81.7%
Clef-flash, wrong among the messages it routed itself2.15%
Clef-flash, wrong among routed, 95% interval, low1.65
Clef-flash, wrong among routed, 95% interval, high2.79
Clef-flash, messages sent to a person564
Clef-flash, wrong answers routed without a person54
Jev, answers given 0.2 to 0.3: how many1
Jev, answers given 0.2 to 0.3: average probability given0.28
Jev, answers given 0.2 to 0.3: share right0.00
Jev, answers given 0.3 to 0.4: how many3
Jev, answers given 0.3 to 0.4: average probability given0.35
Jev, answers given 0.3 to 0.4: share right0.67
Jev, answers given 0.4 to 0.5: how many17
Jev, answers given 0.4 to 0.5: average probability given0.45
Jev, answers given 0.4 to 0.5: share right0.35
Jev, answers given 0.5 to 0.6: how many37
Jev, answers given 0.5 to 0.6: average probability given0.55
Jev, answers given 0.5 to 0.6: share right0.49
Jev, answers given 0.6 to 0.7: how many39
Jev, answers given 0.6 to 0.7: average probability given0.64
Jev, answers given 0.6 to 0.7: share right0.44
Jev, answers given 0.7 to 0.8: how many46
Jev, answers given 0.7 to 0.8: average probability given0.75
Jev, answers given 0.7 to 0.8: share right0.63
Jev, answers given 0.8 to 0.9: how many81
Jev, answers given 0.8 to 0.9: average probability given0.85
Jev, answers given 0.8 to 0.9: share right0.72
Jev, answers given 0.9 to 1.0: how many2,856
Jev, answers given 0.9 to 1.0: average probability given1.00
Jev, answers given 0.9 to 1.0: share right0.97
Clef, answers given 0.2 to 0.3: how many5
Clef, answers given 0.2 to 0.3: average probability given0.28
Clef, answers given 0.2 to 0.3: share right0.40
Clef, answers given 0.3 to 0.4: how many13
Clef, answers given 0.3 to 0.4: average probability given0.36
Clef, answers given 0.3 to 0.4: share right0.69
Clef, answers given 0.4 to 0.5: how many23
Clef, answers given 0.4 to 0.5: average probability given0.45
Clef, answers given 0.4 to 0.5: share right0.57
Clef, answers given 0.5 to 0.6: how many38
Clef, answers given 0.5 to 0.6: average probability given0.55
Clef, answers given 0.5 to 0.6: share right0.50
Clef, answers given 0.6 to 0.7: how many55
Clef, answers given 0.6 to 0.7: average probability given0.65
Clef, answers given 0.6 to 0.7: share right0.56
Clef, answers given 0.7 to 0.8: how many69
Clef, answers given 0.7 to 0.8: average probability given0.76
Clef, answers given 0.7 to 0.8: share right0.70
Clef, answers given 0.8 to 0.9: how many187
Clef, answers given 0.8 to 0.9: average probability given0.86
Clef, answers given 0.8 to 0.9: share right0.83
Clef, answers given 0.9 to 1.0: how many2,690
Clef, answers given 0.9 to 1.0: average probability given0.96
Clef, answers given 0.9 to 1.0: share right0.99
Clef-flash, answers given 0.2 to 0.3: how many7
Clef-flash, answers given 0.2 to 0.3: average probability given0.26
Clef-flash, answers given 0.2 to 0.3: share right0.14
Clef-flash, answers given 0.3 to 0.4: how many12
Clef-flash, answers given 0.3 to 0.4: average probability given0.35
Clef-flash, answers given 0.3 to 0.4: share right0.33
Clef-flash, answers given 0.4 to 0.5: how many59
Clef-flash, answers given 0.4 to 0.5: average probability given0.46
Clef-flash, answers given 0.4 to 0.5: share right0.68
Clef-flash, answers given 0.5 to 0.6: how many87
Clef-flash, answers given 0.5 to 0.6: average probability given0.56
Clef-flash, answers given 0.5 to 0.6: share right0.78
Clef-flash, answers given 0.6 to 0.7: how many241
Clef-flash, answers given 0.6 to 0.7: average probability given0.66
Clef-flash, answers given 0.6 to 0.7: share right0.87
Clef-flash, answers given 0.7 to 0.8: how many736
Clef-flash, answers given 0.7 to 0.8: average probability given0.76
Clef-flash, answers given 0.7 to 0.8: share right0.95
Clef-flash, answers given 0.8 to 0.9: how many1,339
Clef-flash, answers given 0.8 to 0.9: average probability given0.85
Clef-flash, answers given 0.8 to 0.9: share right0.98
Clef-flash, answers given 0.9 to 1.0: how many599
Clef-flash, answers given 0.9 to 1.0: average probability given0.92
Clef-flash, answers given 0.9 to 1.0: share right1.00
Messages Clef got right and Jev got wrong54
Messages Jev got right and Clef got wrong19
Messages Clef-flash got right and Jev got wrong60
Messages Jev got right and Clef-flash got wrong22
Clef-flash, that group's lowest probability0.7
Clef-flash, that group's highest probability0.8
Clef-flash, answers it gave a probability of 0.7 to 0.8736
Those answers: average stated probability0.76
Those answers: share right95%
A bare request to TypeSafe's host from the Mac mini, warm, median of 20, start of the run171 ms
A bare request to Cloudflare's API host, warm, median of 20, start of the run (not the inference path)204 ms
Clef, same answer on a second pass200
Clef-flash, same answer on a second pass200
Messages in the second pass200
Jev, dollars per million messages (our arithmetic)$77.53
Jev, input tokens per request (our arithmetic)1,846
Jev, sent to a person per 1,000 messages (our arithmetic)156
Jev, wrong answers routed without a person per 1,000 messages (our arithmetic)17.2
Clef, dollars per million messages (our arithmetic)$479.75
Clef, input tokens per request (our arithmetic)1,999
Clef, sent to a person per 1,000 messages (our arithmetic)68
Clef, wrong answers routed without a person per 1,000 messages (our arithmetic)22.4
Clef-flash, dollars per million messages (our arithmetic)$179.91
Clef-flash, input tokens per request (our arithmetic)1,999
Clef-flash, sent to a person per 1,000 messages (our arithmetic)183
Clef-flash, wrong answers routed without a person per 1,000 messages (our arithmetic)17.5
Clef minus Jev, share right with five examples, same messages+1.14 points
Clef minus Jev, 95% interval, low+0.62
Clef minus Jev, 95% interval, high+1.69
Clef-flash minus Jev, share right with five examples, same messages+1.23 points
Clef-flash minus Jev, 95% interval, low+0.68
Clef-flash minus Jev, 95% interval, high+1.82
Clef's price per message against Jev's, with five examples (our arithmetic)6.2×
Clef-flash's price per message against Jev's, with five examples (our arithmetic)2.3×
Clef-flash's price per message as a share of Clef's (our arithmetic)38%
How many more input tokens Cloudflare's models count than Jev for the same request (our arithmetic)8%
Fewer messages per 1,000 sent to a person by Clef than by Jev (our arithmetic)88
The same, unrounded (our arithmetic)88.3
Clef's extra API cost per 1,000 messages over Jev's (our arithmetic)$0.40
Break-even: Clef's extra cost per person's check it saves (our arithmetic)$0.0046
More wrong answers per 1,000 routed without a person by Clef than by Jev (our arithmetic)5.2
Clef minus Jev, share right with five examples (our arithmetic)+1.1 points

Model Fatigue

Our dev-slice and example-pool counts (dev-slice/stats.json) · frozen 30 Sep 2026
Reasons (intents) a message can be sorted into77
Training messages held out to set each model's cut-off before the test573
Labelled training messages the five examples are drawn from9,405

Model Fatigue

Our protocol (FREEZE.md), fixed before any test message was sent · frozen 30 Sep 2026
Options in our request: the reasons plus none-of-these78
Labelled examples in each of our requests5
Wrong answers allowed among those a model routes on its own, the aim each cut-off was set for2%

Cloudflare

Edition of the Decision Index Cloudflare ran0.2.1
Tests in the Decision Index table on Clef's model card (our count)41
Of those, tests where Clef is ahead of Jev (our count)26
Of those, tests where Jev is ahead of Clef (our count)15

Cloudflare

Cloudflare Workers AI pricing · read 2 Oct 2026, 20:52 CEST
Clef, price per million input tokens on Workers AI$0.240
Clef-flash, price per million input tokens on Workers AI$0.090
Clef's price per token against Jev's (our arithmetic)5.7×
Clef-flash's price per token against Jev's (our arithmetic)2.1×

TypeSafe

Jev, price per million input tokens$0.042

Model Fatigue

Our replication of the Decision Index's BANKING77 request, all five models · run 2 Oct 2026, 17:03–21:41 CEST
Jev, macro-F1 on the Decision Index's request80.00
Jev, macro-F1 on their request, 95% interval, low78.61
Jev, macro-F1 on their request, 95% interval, high81.08
Jev, share right on the Decision Index's request80.7%
Clef, macro-F1 on the Decision Index's request94.20
Clef, macro-F1 on their request, 95% interval, low93.29
Clef, macro-F1 on their request, 95% interval, high94.94
Clef, share right on the Decision Index's request94.2%
Clef-flash, macro-F1 on the Decision Index's request90.85
Clef-flash, macro-F1 on their request, 95% interval, low89.78
Clef-flash, macro-F1 on their request, 95% interval, high91.74
Clef-flash, share right on the Decision Index's request90.9%
Kev-27B, macro-F1 on the Decision Index's request86.41
Kev-27B, macro-F1 on their request, 95% interval, low85.14
Kev-27B, macro-F1 on their request, 95% interval, high87.41
Kev-27B, share right on the Decision Index's request86.8%
Kev-9B, macro-F1 on the Decision Index's request83.03
Kev-9B, macro-F1 on their request, 95% interval, low81.57
Kev-9B, macro-F1 on their request, 95% interval, high84.10
Kev-9B, share right on the Decision Index's request83.3%
Jev, share right with five examples minus on their request (our arithmetic)+13.2 points
Clef, share right with five examples minus on their request (our arithmetic)+0.8 points
Clef-flash, share right with five examples minus on their request (our arithmetic)+4.3 points
Kev-27B, share right with five examples minus on their request (our arithmetic)+7.0 points
Kev-9B, share right with five examples minus on their request (our arithmetic)+8.0 points
Jev, dollars per 1,000 messages on the Decision Index's request (price list × tokens; our arithmetic)$0.072
Clef, dollars per 1,000 messages on the Decision Index's request (price list × tokens; our arithmetic)$0.448
Clef-flash, dollars per 1,000 messages on the Decision Index's request (price list × tokens; our arithmetic)$0.168
Clef minus Jev, macro-F1 on their request (our arithmetic)+14.2 points
Clef-flash minus Jev, macro-F1 on their request (our arithmetic)+10.8 points
Clef minus Jev, share right on their request (our arithmetic)+13.5 points

Model Fatigue

Our run of Kev-27B and Kev-9B on H100s we rented on Modal (Europe) · run 2 Oct 2026, 21:05–22:00 CEST
Kev-27B, share right with five examples93.8%
Kev-27B, share right, 95% interval, low92.9
Kev-27B, share right, 95% interval, high94.6
Kev-27B, messages sorted correctly with five examples2,890
Kev-27B, times it answered none-of-these (always wrong on this test)21
Kev-27B, median time per request from Berlin529 ms
Kev-27B, 95th percentile time per request from Berlin551 ms
Kev-27B, 99th percentile time per request from Berlin598 ms
Kev-27B, dollars per 1,000 messages (our GPU time as Modal billed it, one call at a time)$0.7295
Kev-27B, calibration error of its probabilities (ECE; lower is closer)0.014
Kev-27B, how well its probabilities separate its right answers from its wrong ones (AUROC)0.92
Kev-27B, cut-off on its top probability, set on the held-out training messages0.79
Kev-27B, share of messages routed without a person91.0%
Kev-27B, wrong among the messages it routed itself2.28%
Kev-27B, wrong among routed, 95% interval, low1.79
Kev-27B, wrong among routed, 95% interval, high2.90
Kev-27B, messages sent to a person276
Kev-27B, wrong answers routed without a person64
Kev-9B, share right with five examples91.3%
Kev-9B, share right, 95% interval, low90.3
Kev-9B, share right, 95% interval, high92.3
Kev-9B, messages sorted correctly with five examples2,813
Kev-9B, times it answered none-of-these (always wrong on this test)110
Kev-9B, median time per request from Berlin437 ms
Kev-9B, 95th percentile time per request from Berlin548 ms
Kev-9B, 99th percentile time per request from Berlin665 ms
Kev-9B, dollars per 1,000 messages (our GPU time as Modal billed it, one call at a time)$0.6211
Kev-9B, calibration error of its probabilities (ECE; lower is closer)0.038
Kev-9B, how well its probabilities separate its right answers from its wrong ones (AUROC)0.91
Kev-9B, cut-off on its top probability, set on the held-out training messages0.81
Kev-9B, share of messages routed without a person82.7%
Kev-9B, wrong among the messages it routed itself2.47%
Kev-9B, wrong among routed, 95% interval, low1.94
Kev-9B, wrong among routed, 95% interval, high3.15
Kev-9B, messages sent to a person533
Kev-9B, wrong answers routed without a person63
Kev-27B, answers given 0.1 to 0.2: how many1
Kev-27B, answers given 0.1 to 0.2: average probability given0.19
Kev-27B, answers given 0.1 to 0.2: share right1.00
Kev-27B, answers given 0.2 to 0.3: how many9
Kev-27B, answers given 0.2 to 0.3: average probability given0.25
Kev-27B, answers given 0.2 to 0.3: share right0.22
Kev-27B, answers given 0.3 to 0.4: how many25
Kev-27B, answers given 0.3 to 0.4: average probability given0.37
Kev-27B, answers given 0.3 to 0.4: share right0.36
Kev-27B, answers given 0.4 to 0.5: how many51
Kev-27B, answers given 0.4 to 0.5: average probability given0.45
Kev-27B, answers given 0.4 to 0.5: share right0.43
Kev-27B, answers given 0.5 to 0.6: how many55
Kev-27B, answers given 0.5 to 0.6: average probability given0.55
Kev-27B, answers given 0.5 to 0.6: share right0.45
Kev-27B, answers given 0.6 to 0.7: how many62
Kev-27B, answers given 0.6 to 0.7: average probability given0.65
Kev-27B, answers given 0.6 to 0.7: share right0.56
Kev-27B, answers given 0.7 to 0.8: how many75
Kev-27B, answers given 0.7 to 0.8: average probability given0.75
Kev-27B, answers given 0.7 to 0.8: share right0.77
Kev-27B, answers given 0.8 to 0.9: how many204
Kev-27B, answers given 0.8 to 0.9: average probability given0.86
Kev-27B, answers given 0.8 to 0.9: share right0.86
Kev-27B, answers given 0.9 to 1.0: how many2,598
Kev-27B, answers given 0.9 to 1.0: average probability given0.98
Kev-27B, answers given 0.9 to 1.0: share right0.99
Kev-9B, answers given 0.2 to 0.3: how many2
Kev-9B, answers given 0.2 to 0.3: average probability given0.29
Kev-9B, answers given 0.2 to 0.3: share right0.50
Kev-9B, answers given 0.3 to 0.4: how many21
Kev-9B, answers given 0.3 to 0.4: average probability given0.36
Kev-9B, answers given 0.3 to 0.4: share right0.24
Kev-9B, answers given 0.4 to 0.5: how many88
Kev-9B, answers given 0.4 to 0.5: average probability given0.46
Kev-9B, answers given 0.4 to 0.5: share right0.38
Kev-9B, answers given 0.5 to 0.6: how many102
Kev-9B, answers given 0.5 to 0.6: average probability given0.55
Kev-9B, answers given 0.5 to 0.6: share right0.47
Kev-9B, answers given 0.6 to 0.7: how many115
Kev-9B, answers given 0.6 to 0.7: average probability given0.65
Kev-9B, answers given 0.6 to 0.7: share right0.64
Kev-9B, answers given 0.7 to 0.8: how many180
Kev-9B, answers given 0.7 to 0.8: average probability given0.75
Kev-9B, answers given 0.7 to 0.8: share right0.81
Kev-9B, answers given 0.8 to 0.9: how many440
Kev-9B, answers given 0.8 to 0.9: average probability given0.86
Kev-9B, answers given 0.8 to 0.9: share right0.92
Kev-9B, answers given 0.9 to 1.0: how many2,132
Kev-9B, answers given 0.9 to 1.0: average probability given0.96
Kev-9B, answers given 0.9 to 1.0: share right0.99
Messages Kev-27B got right and Jev got wrong24
Messages Jev got right and Kev-27B got wrong27
Messages Kev-9B got right and Jev got wrong17
Messages Jev got right and Kev-9B got wrong97
Kev-27B, share right with none-of-these set aside (its top real reason)94.1%
Kev-9B, share right with none-of-these set aside (its top real reason)93.6%
Kev-27B, dollars per million messages (our arithmetic)$729.51
Kev-27B, input tokens per request (our arithmetic)975
Kev-27B, sent to a person per 1,000 messages (our arithmetic)90
Kev-27B, wrong answers routed without a person per 1,000 messages (our arithmetic)20.8
Kev-9B, dollars per million messages (our arithmetic)$621.12
Kev-9B, input tokens per request (our arithmetic)975
Kev-9B, sent to a person per 1,000 messages (our arithmetic)173
Kev-9B, wrong answers routed without a person per 1,000 messages (our arithmetic)20.5
Kev-27B minus Jev, share right with five examples, same messages−0.10 points
Kev-27B minus Jev, 95% interval, low−0.55
Kev-27B minus Jev, 95% interval, high+0.36
Kev-9B minus Jev, share right with five examples, same messages−2.60 points
Kev-9B minus Jev, 95% interval, low−3.28
Kev-9B minus Jev, 95% interval, high−1.92
Kev-27B minus Clef, share right with five examples, same messages−1.23 points
Kev-27B minus Clef, 95% interval, low−1.75
Kev-27B minus Clef, 95% interval, high−0.71
Kev-27B's cost per message as run against Jev's (our arithmetic)9.4×

Model Fatigue

Our arithmetic on the raw call records and Cloudflare's table (derived.json, this page's own file) · run 2 Oct 2026, 15:49–22:00 CEST, across the passes it draws on
A person's time, our assumption for the break-even$15
Messages Clef routed on its own that Jev sent to a person318
Of those, how many Clef got wrong31
Messages where one of Clef and Jev was right and the other wrong73
Of those, messages where both answers were below their cut-offs47
Of those, messages where the model that was right would have sent it to a person53
Jev's answers below 0.99481
Jev's answers below 0.99: average stated probability0.84
Jev's answers below 0.99: share right72%
Jev's answers wrong on their request and right with five examples437
Jev's answers right on their request and wrong with five examples30
Message A, Jev, probability of its answer on their request0.63
Message A, Jev, probability of its answer on with five examples1.00
Message A, Clef, probability of its answer on their request0.90
Message A, Clef, probability of its answer on with five examples0.96
Message A, Clef-flash, probability of its answer on their request0.77
Message A, Clef-flash, probability of its answer on with five examples0.89
Message B, Jev, probability of its answer on their request0.48
Message B, Jev, probability of its answer on with five examples0.81
Message B, Clef, probability of its answer on their request0.28
Message B, Clef, probability of its answer on with five examples0.27
Message B, Clef-flash, probability of its answer on their request0.10
Message B, Clef-flash, probability of its answer on with five examples0.48
Message C, Jev, probability of its answer on their request0.64
Message C, Jev, probability of its answer on with five examples0.44
Message C, Clef, probability of its answer on their request0.67
Message C, Clef, probability of its answer on with five examples0.60
Message C, Clef-flash, probability of its answer on their request0.40
Message C, Clef-flash, probability of its answer on with five examples0.59
Clef-flash, calls that took over a second159
Of those, calls in the first quarter of its pass139
Kev-27B, median time the model took by its own server's count110 ms
Kev-9B, median time the model took by its own server's count35 ms
Kev-27B, input tokens per call on the Decision Index's request (its own count)853
Kev-27B, input tokens per call on our test (its own count)975
Break-even as seconds of a person's time at that rate (our arithmetic)1.1 s

Model Fatigue (M37's jev-2 run)

Jev on our request with the examples taken out (jev-2, 21 September)78.3%
Jev in that control, times it answered none-of-these131

Model Fatigue

Modal's bill for our Kev deployment, read back per pass (kev_gpu_cost.json) · billed 2 Oct 2026, the whole deployment (Modal's hourly rows 20:00 to 22:59 CEST)
Kev-27B, Modal's bill per container-hour (our arithmetic on the bill)$4.92
Kev-27B, dollars per 1,000 at six requests at a time, projected from measured throughput$0.157
Kev-27B, everything Modal billed for this model$4.76
Kev-27B, start-up and idle time on the bill$1.69
Kev-9B, Modal's bill per container-hour (our arithmetic on the bill)$4.91
Kev-9B, dollars per 1,000 at six requests at a time, projected from measured throughput$0.112
Kev-9B, everything Modal billed for this model$4.67
Kev-9B, start-up and idle time on the bill$2.14
Kev-27B, start-up and idle as a share of its bill (our arithmetic)35%
Kev-9B, start-up and idle as a share of its bill (our arithmetic)46%

Jared Palmer

Kev-27B on breadth-v1, datasets no Kev trained on (release notes)52.3
Jev on breadth-v1, as Kev's release notes report it54.0

Cloudflare

Cloudflare's table, BFCL · case exact: Clef98.47
Cloudflare's table, BFCL · case exact: Clef-flash98.76
Cloudflare's table, BFCL · case exact: Jev95.75
Cloudflare's table, ToolRet · nDCG@10: Clef69.19
Cloudflare's table, ToolRet · nDCG@10: Clef-flash66.43
Cloudflare's table, ToolRet · nDCG@10: Jev65.28
Cloudflare's table, API-Bank · accuracy: Clef91.93
Cloudflare's table, API-Bank · accuracy: Clef-flash93.11
Cloudflare's table, API-Bank · accuracy: Jev88.19
Cloudflare's table, Home appliances · case exact: Clef82.95
Cloudflare's table, Home appliances · case exact: Clef-flash97.73
Cloudflare's table, Home appliances · case exact: Jev52.27
Cloudflare's table, When2Call · accuracy: Clef72.37
Cloudflare's table, When2Call · accuracy: Clef-flash65.58
Cloudflare's table, When2Call · accuracy: Jev80.97
Cloudflare's table, BANKING77 · macro-F1: Clef94.20
Cloudflare's table, BANKING77 · macro-F1: Clef-flash90.93
Cloudflare's table, BANKING77 · macro-F1: Jev79.74
Cloudflare's table, CLINC150+OOS · macro-F1: Clef97.43
Cloudflare's table, CLINC150+OOS · macro-F1: Clef-flash66.77
Cloudflare's table, CLINC150+OOS · macro-F1: Jev89.27
Cloudflare's table, BRIGHT · nDCG@10: Clef45.91
Cloudflare's table, BRIGHT · nDCG@10: Clef-flash39.26
Cloudflare's table, BRIGHT · nDCG@10: Jev47.52
Cloudflare's table, Amazon ESCI · macro-F1: Clef57.48
Cloudflare's table, Amazon ESCI · macro-F1: Clef-flash57.39
Cloudflare's table, Amazon ESCI · macro-F1: Jev55.21
Cloudflare's table, PhishNChips · accuracy: Clef79.60
Cloudflare's table, PhishNChips · accuracy: Clef-flash75.05
Cloudflare's table, PhishNChips · accuracy: Jev62.55
Cloudflare's workflow table, Invoice processing: Clef64.7
Cloudflare's workflow table, Invoice processing: Clef-flash57.1
Cloudflare's workflow table, Invoice processing: Jev61.8
Cloudflare's workflow table, Customer service: Clef76.3
Cloudflare's workflow table, Customer service: Clef-flash77.0
Cloudflare's workflow table, Customer service: Jev76.0
Cloudflare's workflow table, Security incidents: Clef62.9
Cloudflare's workflow table, Security incidents: Clef-flash61.7
Cloudflare's workflow table, Security incidents: Jev61.7
Cloudflare's workflow table, Agent trace observability: Clef68.5
Cloudflare's workflow table, Agent trace observability: Clef-flash69.8
Cloudflare's workflow table, Agent trace observability: Jev71.6
Cloudflare's latency table, median ms: Clef209.3
Cloudflare's latency table, median ms: Clef-flash38.8
Cloudflare's latency table, median ms: Jev524.1
Cloudflare's latency table, p95 ms: Clef238.6
Cloudflare's latency table, p95 ms: Clef-flash122.4
Cloudflare's latency table, p95 ms: Jev536.0
Kev 9B (the board's older checkpoint) on BANKING77 in Cloudflare's table84.83
Benchmarks Cloudflare's post says it ran43
Clef minus Jev on BANKING77 in Cloudflare's table (our arithmetic)+14.5 points
Jev's median time over Clef's in Cloudflare's table (our arithmetic)2.5×
Jev's median time over Clef-flash's in Cloudflare's table (our arithmetic)13.5×

Model Fatigue

Our memorisation check: 200 training messages and reworded copies · run 2 Oct 2026, 17:36–17:38 CEST
Clef, share right on training messages94.5%
Clef, share right on reworded copies of them87.0%
Clef-flash, share right on training messages95.5%
Clef-flash, share right on reworded copies of them83.5%
Training messages in the check (and as many reworded copies)200
Clef, originals minus reworded copies+7.5 points
Clef-flash, originals minus reworded copies+12.0 points

Model Fatigue (M37's jev-2 run)

Our results write-up (RESULTS.md), jev-2's memorisation yardstick · run 21–24 Sep 2026
A vote of the five nearest labelled examples, which can't remember anything: originals minus reworded copies+15.0 points

Sources

These are the pages the video and this page draw on. We keep a copy of each page as we read it, so a figure can be checked against what the page said at the time.

Credits

The narration in the video is an AI voice, made with ElevenLabs.

Music in the video: "Airport Lounge" by Kevin MacLeod (incompetech.com), licensed under Creative Commons: By Attribution 4.0.

Further reading