This is the video written out, with every figure in full, each linked to its source in the table below.
Cloudflare's claim▶ 0:00
Cloudflare's launch post on 1 October says "the market is getting increasingly saturated with decision models". Then it releases two more, Clef and Clef-flash. Their weights are open, and they take the same requests as Jev, TypeSafe's decision model.
A decision model takes a message and a fixed list of answers, like the teams that could handle a request, and returns a probability for each answer instead of writing text. Cloudflare's post compares its two models with Jev and three others in a table of ten tests, and Clef is ahead of Jev on eight of them. The fuller table on Clef's Hugging Face card has 41 tests, and Clef is ahead on 26. Jev is ahead on the other 15, among them several general knowledge and reasoning tests (MMLU-Pro, BBH, GPQA Diamond).
Cloudflare's main table
The ten tests in Cloudflare's launch post, Clef, Clef-flash and Jev (the post also lists three other models)
| test | Clef | Clef-flash | Jev |
|---|---|---|---|
| BFCL (case exact) | 98.47 | 98.76 | 95.75 |
| ToolRet (nDCG@10) | 69.19 | 66.43 | 65.28 |
| API-Bank (accuracy) | 91.93 | 93.11 | 88.19 |
| Home appliances (case exact) | 82.95 | 97.73 | 52.27 |
| When2Call (accuracy) | 72.37 | 65.58 | 80.97 |
| BANKING77 (macro-F1) | 94.20 | 90.93 | 79.74 |
| CLINC150+OOS (macro-F1) | 97.43 | 66.77 | 89.27 |
| BRIGHT (nDCG@10) | 45.91 | 39.26 | 47.52 |
| Amazon ESCI (macro-F1) | 57.48 | 57.39 | 55.21 |
| PhishNChips (accuracy) | 79.60 | 75.05 | 62.55 |
A second table in the post covers four business workflows from TypeSafe's own evaluation suite, and Clef is ahead of Jev on three of them. We didn't run those.
Cloudflare's workflow table
Four business workflows from TypeSafe's own evaluation suite, as Cloudflare ran them
| workflow | Clef | Clef-flash | Jev |
|---|---|---|---|
| Invoice processing | 64.7 | 57.1 | 61.8 |
| Customer service | 76.3 | 77.0 | 76.0 |
| Security incidents | 62.9 | 61.7 | 61.7 |
| Agent trace observability | 68.5 | 69.8 | 71.6 |
The test we checked is in both the post's main table and the card's. BANKING77 is a public set of messages that customers sent to a bank, each filed under one of 77 reasons, like a card that hasn't arrived or a declined transfer. Its test split has 3,080 messages. On it, Cloudflare's table gives Clef 94.20 and Jev 79.74, a lead of 14.5 points. The score is macro-F1, the average over the reasons of how well the model handles each one, counting both the messages it misses and the ones it files there wrongly, and on these messages it comes out within a point of the share each model gets right. If you sort support tickets, a gap that size would make you switch. So is Clef really that far ahead, and is it worth what it costs?
The basics▶ 1:32
Cloudflare released Clef and Clef-flash on 1 October. It built them on open Qwen models, Qwen3.8-27B for Clef and Qwen3.5-9B for Clef-flash, kept those weights frozen, and trained small adapters inside them and a scoring head on top, with what its post calls "our own internal synthetic datasets". The weights are on Hugging Face under the Apache 2.0 licence, so you can run them yourself, and Cloudflare hosts both on Workers AI.
The three models
As their makers publish them
| Jev | Clef | Clef-flash | |
|---|---|---|---|
| made by | TypeSafe | Cloudflare | Cloudflare |
| built on | not published | Qwen3.8-27B | Qwen3.5-9B |
| weights | not published | Apache 2.0, on Hugging Face | Apache 2.0, on Hugging Face |
| hosted by | TypeSafe | Cloudflare Workers AI | Cloudflare Workers AI |
| price per million input tokens | $0.042 | $0.240 | $0.090 |
| output tokens | free | not charged | not charged |
| images | no | yes | yes |
| version we tested | jev-1.13.0 | clef | clef-flash |
You pay for what you send in, and the answers come back free on all three. Per token, Clef costs 5.7× what Jev does and Clef-flash 2.1×. Both Cloudflare models can read images, which Jev can't. We didn't test that.
Their test, run ourselves▶ 2:01
Cloudflare's table comes from the Decision Index, a community benchmark that says it is "unofficial and community-maintained; not affiliated with TypeSafe AI". BANKING77 is one of its tests. Cloudflare ran Clef and Clef-flash on edition 0.2.1 itself and labels those scores self-reported. Its number for Jev is the benchmark's own run. So the Clef scores and the Jev score in that table come from different runs.
We sent all 3,080 messages to Clef, Clef-flash and Jev, each request exactly as the Decision Index sends it: no context, the instruction "Classify the banking intent of this user request:" with the message, and the 77 reasons as the options. We sent the same requests to Kev-27B and Kev-9B, which come up later.
Their test: the Decision Index's request
The message alone and the reasons as options, no examples. All five models, the same messages, sent the same day.
| model | macro-F1 | 95% interval | share right | Cloudflare's table |
|---|---|---|---|---|
| Clef | 94.20 | 93.29 to 94.94 | 94.2% | 94.20 |
| Clef-flash | 90.85 | 89.78 to 91.74 | 90.9% | 90.93 |
| Kev-27B | 86.41 | 85.14 to 87.41 | 86.8% | not in it |
| Kev-9B | 83.03 | 81.57 to 84.10 | 83.3% | 84.83, an older Kev-9B |
| Jev | 80.00 | 78.61 to 81.08 | 80.7% | 79.74 |
Clef scored 94.20, the same as Cloudflare's table. Clef-flash scored 90.85 against the table's 90.93, and Jev 80.00 against 79.74. All three sit inside our 95% interval. Kev-9B's 84.83 in that table is an older checkpoint of it, so it doesn't compare with ours. So the table holds: on this test Clef is 14.2 points ahead of Jev in macro-F1, 13.5 points in the share right.
How we ran it: we fixed the procedure before sending a single test message, covering which messages, which request for each model, how answers are scored, how each model's cut-off for sending a message to a person is set, and how costs are counted. Clef, Clef-flash and Kev each got a short addendum on 2 October, written before their test passes, that changes only the model name and the address the request goes to. Kev's servers had answered a timing probe of the first forty test messages just before its addendum was written; nothing was set from it, and its answers weren't scored.
Their request is the Decision Index's own, copied byte for byte from its code. We sent it six at a time from a laptop in Berlin, so we don't use its times for speed. The one place they come in is Kev's cost projection below, which says so.
Our request carries the message, the five most similar messages from BANKING77's training split with their reasons, and 78 options: the 77 reasons plus "none of these", which a real queue needs and which is always wrong on this test. One fixed embedding model finds the examples in a pool of 9,405 labelled training messages, with test duplicates and our held-out messages removed, so no test message ever appears as its own example. Every model gets the same examples. We sent it to Jev, Clef and Clef-flash on the afternoon of 2 October, one request at a time from a Mac mini on a home connection in Berlin, Jev first. Kev-27B and Kev-9B ran the same evening on H100s we rented on Modal in Europe.
The cut-offs for sending a message to a person were set on 573 training messages held out from the example pool, before the test, aiming at 2% wrong answers among the messages a model routes on its own. Costs are each provider's list price times the input tokens it counted on every call, and none of the three hosted models charges for output. For Cloudflare, the neurons it reported on every call add up to the same dollars. Kev's cost is Modal's bill for our GPU time.
Five examples▶ 2:52
The Decision Index's request gives the model the customer's message and the list of reasons, and nothing else. If you sort support tickets for real, you also have old tickets that someone has already sorted. You can look up the five most like the new one and put them in the request with their right answers, which is called few-shot prompting. That is our request.
Take one test message: "Can I be given a new passcode?", filed under forgotten passcode. On their request, Jev picks change PIN at 0.63, and Clef picks forgotten passcode at 0.90. The five most similar sorted messages ("Need a new passcode.", "Where can I get a new passcode?" and three more) were all filed under forgotten passcode. With them in the request, Jev picks forgotten passcode at 1.00. Jev rounds its probabilities to two decimals.
Three messages, three models
Each model's answer and its probability on their request, then with five examples. ✓ marks the reason the dataset files it under.
| message | model | their request | with five examples |
|---|---|---|---|
| “Can I be given a new passcode?” (filed under passcode forgotten) | Jev | change pin (0.63) | passcode forgotten (1.00) ✓ |
| Clef-flash | passcode forgotten (0.77) ✓ | passcode forgotten (0.89) ✓ | |
| Clef | passcode forgotten (0.90) ✓ | passcode forgotten (0.96) ✓ | |
| “How do I contact customer support about a transfer?” (filed under declined transfer) | Jev | pending transfer (0.48) | none of these (0.81) |
| Clef-flash | transfer not received by recipient (0.10) | declined transfer (0.48) ✓ | |
| Clef | balance not updated after bank transfer (0.28) | declined transfer (0.27) ✓ | |
| “Any fee for topping up?” (filed under top up by card charge) | Jev | top up by card charge (0.64) ✓ | top up by card charge (0.44) ✓ |
| Clef-flash | top up by bank transfer charge (0.40) | top up by card charge (0.59) ✓ | |
| Clef | top up by card charge (0.67) ✓ | top up by bank transfer charge (0.60) |
Across the test, 437 of Jev's answers went from wrong to right and 30 the other way. To check that the examples and not our different wording carry that, an earlier run of our request with the examples taken out had Jev at 78.3%, a little under what it got on the Decision Index's request.
Without examples, and with five
Share right on the same messages: their request (the message alone), then ours with five examples
The numbers in this chart
| their request | with five examples | difference, points | |
|---|---|---|---|
| Clef | 94.2% | 95.1% | +0.8 points |
| Clef-flash | 90.9% | 95.2% | +4.3 points |
| Kev-27B | 86.8% | 93.8% | +7.0 points |
| Kev-9B | 83.3% | 91.3% | +8.0 points |
| Jev | 80.7% | 93.9% | +13.2 points |
With five examples Jev got 93.9% right, Clef 95.1% and Clef-flash 95.2%. Clef barely needed the examples. Its lead in the share right, 13.5 points without them, is down to 1.1 points. That point isn't chance. On the same messages Clef got 54 right that Jev got wrong, and Jev 19 that Clef got wrong, and the 95% interval on the difference runs from +0.62 to +1.69 points.
Our test: with five examples
The same messages, each with its five most similar sorted training messages, and none of these as an extra option
| model | share right | 95% interval | minus Jev, points | 95% interval | none of these |
|---|---|---|---|---|---|
| Clef-flash | 95.2% | 94.3 to 95.9 | +1.23 | +0.68 to +1.82 | 0 |
| Clef | 95.1% | 94.2 to 95.8 | +1.14 | +0.62 to +1.69 | 0 |
| Jev | 93.9% | 93.0 to 94.7 | 14 | ||
| Kev-27B | 93.8% | 92.9 to 94.6 | −0.10 | −0.55 to +0.36 | 21 |
| Kev-9B | 91.3% | 90.3 to 92.3 | −2.60 | −3.28 to −1.92 | 110 |
Here is one each way. "How do I contact customer support about a transfer?" is filed under declined transfer. Clef gets it at 0.27 and Jev picks none of these. "Any fee for topping up?" is filed under the fee for topping up by card. Jev gets it at 0.44, and Clef picks the fee for topping up by bank transfer. At the cut-offs in "Who gets a person", all four of those answers would have gone to a person. Clef and Jev split 73 messages this way, one right and the other wrong. On 47 of them both answers were below their cut-offs, so if a person checks the unsure ones, they'd have seen most of these messages whichever model you used.
The bill▶ 5:38
In our run with five examples, sorting a million messages would cost $77.53 on Jev, $179.91 on Clef-flash and $479.75 on Clef. Most of the gap is the price per token. The rest is that Cloudflare's models count 8% more input tokens than Jev for the same request, 1,999 against 1,846.
Dollars per thousand messages
Our test with five examples
What it costs
Our test with five examples
| model | per thousand messages | per million | input tokens per request | against Jev |
|---|---|---|---|---|
| Jev | $0.0775 | $77.53 | 1,846 | |
| Clef-flash | $0.1799 | $179.91 | 1,999 | 2.3× |
| Clef | $0.4798 | $479.75 | 1,999 | 6.2× |
| Kev-27B | $0.7295 | GPU time | 975 | 9.4× |
| Kev-9B | $0.6211 | GPU time | 975 |
So with examples, Clef beats Jev by about one message in a hundred for 6.2× the price, and Clef-flash gets the same point for 2.3× the price. Without examples the requests are shorter: per thousand messages, their request cost $0.448 on Clef, $0.168 on Clef-flash and $0.072 on Jev.
Speed▶ 6:17
Cloudflare's table also says its models answer faster than Jev. It gives median times of 209.3 ms for Clef, 38.8 ms for Clef-flash and 524.1 ms for Jev, which makes Clef 2.5× as fast and Clef-flash 13.5× by our arithmetic. Those times aren't measured the same way. The Decision Index says its Jev time is a "network round-trip from our lab" and not comparable with times measured on a model running on its own hardware. Cloudflare's leaderboard marks Clef's scores and latency as self-reported, with the latency "not measured on the board's hardware". Its post counts on its network for the rest: "our Clef models are hosted on Workers AI", "leading to low network latency".
So we timed every call in our run with five examples, from Berlin. Our requests entered Cloudflare's network at its Berlin edge. From somewhere else the order could change.
Time per request, from Berlin
The bar runs to the median; the ticks mark p95 and p99. One request at a time.
Time per request from Berlin
Our run with five examples, one request at a time
| model | median | p95 | p99 | Cloudflare's table, median |
|---|---|---|---|---|
| Clef-flash | 196 ms | 1,014 ms | 2,172 ms | 38.8 ms |
| Jev | 235 ms | 335 ms | 442 ms | 524.1 ms |
| Kev-9B | 437 ms | 548 ms | 665 ms | |
| Kev-27B | 529 ms | 551 ms | 598 ms | |
| Clef | 562 ms | 1,147 ms | 2,001 ms | 209.3 ms |
At the median, Clef-flash answered in 196 ms, Jev in 235 ms and Clef in 562 ms. Clef-flash had the widest spread, with the slowest calls of any model at the 99th percentile: 139 of its 159 calls over a second came in the first quarter of its pass. We ran it once and didn't look into why. For scale, a bare request with no model behind it took 171 ms to TypeSafe's host from the same machine at the start of the run.
Who gets a person▶ 7:23
In a real queue you'd let the model route the messages it's sure about and send the rest to a person. For each model we set a cut-off on the probability of its answer, using held-out training messages and aiming at 2% wrong among the messages it routes on its own. Then we ran the test with examples.
Routing at each model's cut-off
Each model routes the messages above its cut-off and sends the rest to a person
| model | cut-off | routed on its own | wrong among routed (95% interval) | to a person, per thousand | wrong and routed, per thousand | AUROC | calibration error |
|---|---|---|---|---|---|---|---|
| Clef | 0.80 | 93.2% | 2.40% (1.90 to 3.03) | 68 | 22.4 | 0.90 | 0.030 |
| Kev-27B | 0.79 | 91.0% | 2.28% (1.79 to 2.90) | 90 | 20.8 | 0.92 | 0.014 |
| Jev | 0.99 | 84.4% | 2.04% (1.56 to 2.66) | 156 | 17.2 | 0.84 | 0.035 |
| Kev-9B | 0.81 | 82.7% | 2.47% (1.94 to 3.15) | 173 | 20.5 | 0.91 | 0.038 |
| Clef-flash | 0.73 | 81.7% | 2.15% (1.65 to 2.79) | 183 | 17.5 | 0.82 | 0.143 |
Per thousand messages, Clef sent 68 to a person, Jev 156 and Clef-flash 183. Among the messages each routed on its own, all three came out between two and two and a half percent wrong: Clef 2.40%, Jev 2.04% and Clef-flash 2.15%.
Per thousand messages, Clef costs $0.40 more than Jev and sends 88.3 fewer messages to a person, so each check it saves costs $0.0046 in API fees. At $15 an hour, that buys 1.1 s of someone's time. Unless a person can check a message in about a second, Clef is the cheaper one once you add the checking to the API bill.
Clef routes more on its own, so it also lets more wrong answers through: 22.4 per thousand against Jev's 17.2. The extra messages it routes are harder. Of the 318 that Clef routed and Jev sent to a person, Clef got 31 wrong.
How far can you take each model's probability at face value? Clef's come close to how often it turns out right (calibration error 0.030), and it is the best of the three at ranking its right answers above its wrong ones (AUROC 0.90, against 0.82 for Clef-flash). Jev's are close on average but too high whenever it is less than sure: its 481 answers below its cut-off of 0.99 averaged 0.84 and were right 72% of the time. Clef-flash's are too low. On the 736 answers it gave between 0.7 and 0.8 it was right 95% of the time. Check those two on your own data before you take their numbers as they come.
Probability given against the share right
Each model's answers grouped by the probability it gave. On the dashed line a model is right exactly as often as it says.
The numbers in this chart
| stated | answers | mean stated | share right | |
|---|---|---|---|---|
| Kev-27B | 0.1 to 0.2 | 1 | 0.19 | 1.00 |
| Kev-27B | 0.2 to 0.3 | 9 | 0.25 | 0.22 |
| Kev-27B | 0.3 to 0.4 | 25 | 0.37 | 0.36 |
| Kev-27B | 0.4 to 0.5 | 51 | 0.45 | 0.43 |
| Kev-27B | 0.5 to 0.6 | 55 | 0.55 | 0.45 |
| Kev-27B | 0.6 to 0.7 | 62 | 0.65 | 0.56 |
| Kev-27B | 0.7 to 0.8 | 75 | 0.75 | 0.77 |
| Kev-27B | 0.8 to 0.9 | 204 | 0.86 | 0.86 |
| Kev-27B | 0.9 to 1.0 | 2,598 | 0.98 | 0.99 |
| Clef | 0.2 to 0.3 | 5 | 0.28 | 0.40 |
| Clef | 0.3 to 0.4 | 13 | 0.36 | 0.69 |
| Clef | 0.4 to 0.5 | 23 | 0.45 | 0.57 |
| Clef | 0.5 to 0.6 | 38 | 0.55 | 0.50 |
| Clef | 0.6 to 0.7 | 55 | 0.65 | 0.56 |
| Clef | 0.7 to 0.8 | 69 | 0.76 | 0.70 |
| Clef | 0.8 to 0.9 | 187 | 0.86 | 0.83 |
| Clef | 0.9 to 1.0 | 2,690 | 0.96 | 0.99 |
| Jev | 0.2 to 0.3 | 1 | 0.28 | 0.00 |
| Jev | 0.3 to 0.4 | 3 | 0.35 | 0.67 |
| Jev | 0.4 to 0.5 | 17 | 0.45 | 0.35 |
| Jev | 0.5 to 0.6 | 37 | 0.55 | 0.49 |
| Jev | 0.6 to 0.7 | 39 | 0.64 | 0.44 |
| Jev | 0.7 to 0.8 | 46 | 0.75 | 0.63 |
| Jev | 0.8 to 0.9 | 81 | 0.85 | 0.72 |
| Jev | 0.9 to 1.0 | 2,856 | 1.00 | 0.97 |
| Kev-9B | 0.2 to 0.3 | 2 | 0.29 | 0.50 |
| Kev-9B | 0.3 to 0.4 | 21 | 0.36 | 0.24 |
| Kev-9B | 0.4 to 0.5 | 88 | 0.46 | 0.38 |
| Kev-9B | 0.5 to 0.6 | 102 | 0.55 | 0.47 |
| Kev-9B | 0.6 to 0.7 | 115 | 0.65 | 0.64 |
| Kev-9B | 0.7 to 0.8 | 180 | 0.75 | 0.81 |
| Kev-9B | 0.8 to 0.9 | 440 | 0.86 | 0.92 |
| Kev-9B | 0.9 to 1.0 | 2,132 | 0.96 | 0.99 |
| Clef-flash | 0.2 to 0.3 | 7 | 0.26 | 0.14 |
| Clef-flash | 0.3 to 0.4 | 12 | 0.35 | 0.33 |
| Clef-flash | 0.4 to 0.5 | 59 | 0.46 | 0.68 |
| Clef-flash | 0.5 to 0.6 | 87 | 0.56 | 0.78 |
| Clef-flash | 0.6 to 0.7 | 241 | 0.66 | 0.87 |
| Clef-flash | 0.7 to 0.8 | 736 | 0.76 | 0.95 |
| Clef-flash | 0.8 to 0.9 | 1,339 | 0.85 | 0.98 |
| Clef-flash | 0.9 to 1.0 | 599 | 0.92 | 1.00 |
Kev 1.0▶ 9:03
Kev, from the developer Jared Palmer, is one more open-weight model that takes Jev's requests. Kev 1.0 came out on 1 October, the same day as Clef, and gathers models he had already published; its release notes say "Nothing in it is newly trained." Kev-27B is built on Qwen3.8-27B, the same base as Clef, and Kev-9B on Qwen3.5-9B-Base. Its model cards list BANKING77 among their training data, and its code trains on the training split and leaves the test split out. We ran both on H100s we rented on Modal in Europe, the same two ways as the others.
With five examples, Kev-27B got 93.8% right, level with Jev (−0.10 points, interval −0.55 to +0.36) and behind Clef (−1.23 points, interval −1.75 to −0.71). Kev-9B got 91.3%, mostly because it answered none of these 110 times on messages that have a real reason; with that option set aside its top real reason was right 93.6% of the time. On the Decision Index's request Kev-27B scored 86.41, between Jev and Clef. Two teams started from the same base model, and Cloudflare's version does far better when it sees only the message.
Kev-27B's probabilities matched how often it was right more closely than any other model's here (calibration error 0.014), on a dataset whose training split it learned from. At the same aim it sent 90 messages per thousand to a person.
We paid for Kev by the hour of GPU, not per token, so what a message costs depends on how many you send the card at once. Sent one at a time, as we did, it came to $0.7295 per thousand messages, 9.4× Jev's price, before start-up and idle time, which were 35% of Kev-27B's bill. Modal's bill works out at $4.92 per container-hour. From how fast Kev got through the Decision Index pass, sent six at a time from the laptop, six at a time would come to roughly $0.157 per thousand. That pass used shorter requests than our test, 853 input tokens a call against 975 by Kev's count, and we didn't run our test that way, so read it as a rough guide. From Berlin its median answer took 529 ms, of which the model itself took 110 ms.
What Kev cost on our GPUs
Rented H100s on Modal, billed by the second
| Kev-27B | Kev-9B | |
|---|---|---|
| Modal's bill per container-hour | $4.92 | $4.91 |
| per thousand messages, one at a time, as run | $0.7295 | $0.6211 |
| per thousand, six at a time (rough projection) | $0.157 | $0.112 |
| start-up and idle, share of the bill | 35% | 46% |
| median time from Berlin | 529 ms | 437 ms |
| of which the model itself | 110 ms | 35 ms |
Kev's own release notes put Kev-27B slightly behind Jev on datasets no Kev trained on: 52.3 against 54.0 on the release's index for datasets held out from training (breadth).
What this doesn't settle▶ 10:40
Cloudflare says it trained Clef on top of Qwen with its own synthetic data. Neither its post nor the model cards say whether BANKING77, or anything made from it, was in any of the training. Kev's cards list it. TypeSafe's models page says Jev isn't trained on customer requests and doesn't list what it was trained on.
We ran a check on Clef and Clef-flash that compares 200 training messages they might have seen with reworded copies of them. Both did better on the originals, but so does a vote of the five nearest labelled examples, which can't remember anything (+15.0 points), and both gaps were smaller than that. So the check can't tell memory from the cost of rewording. That matters most for the result without examples, where Clef's lead is big.
Have they seen these messages before?
Training messages against reworded copies of the same messages, without examples
| model | training messages | reworded copies | difference |
|---|---|---|---|
| Clef | 94.5% | 87.0% | +7.5 points |
| Clef-flash | 95.5% | 83.5% | +12.0 points |
| vote of the five nearest labelled examples | +15.0 points |
We tested one task, in English. In the Hugging Face table Jev is ahead of Clef on 15 of the 41 tests, so other tasks may come out differently. We didn't test images. On a second pass over the same messages, Clef and Clef-flash gave the same answer to 200 of 200. And this is one day of calls, from Berlin.
Which to use▶ 11:33
If you have no sorted messages to show the model, use Clef. It was 14.2 points ahead of Jev on the Decision Index's request, and Clef-flash 10.8 points, for 38% of Clef's price.
If you do have sorted messages, Clef, Clef-flash and Jev come within about a point of each other, and the other differences decide it.
For the smallest bill, use Jev, at about a sixth of Clef's price per message.
If a person checks the answers the model isn't sure about, Clef sends them fewer than half as many messages as Jev. Unless a check costs you less than $0.0046, that makes Clef the cheaper one overall, with a few more wrong answers getting through.
If no person checks the answers, Clef-flash got as many right as Clef in our run, for 38% of Clef's price and 2.3× Jev's.
For speed from Berlin, Clef-flash and Jev were close at the median, and Clef was the slowest.
If you want to run one on your own hardware, Clef, Clef-flash and Kev are all open, and Clef and Clef-flash were the most accurate of them here. Kev-27B matched Jev with examples, but it learned from this dataset's training split, and its own release notes put it slightly behind Jev on data no Kev trained on. What it costs depends on how you host it.