This is the video written out, with every figure in full, each linked to its source in the table below.
A new number one▶ 0:00
JevBench is a leaderboard of decision models like Jev, made by Benchmark Heaven. On Saturday, 3 October, and Sunday, 4 October, people kept sending it requests, on its GitHub page, to rank new models. One of Sunday's requests had five models in it, from a developer called Chinh Nguyen. On Monday 5 October the leaderboard's maker put out v1.6.0, with all 92 of its ranked systems tested from scratch on new questions, and the biggest of those five, Quyet-1.0-Large, came out first on its headline score, ahead of Jev.
So we tested it. Does the new number one beat Jev on someone else's test too, or only on JevBench's?
The leaderboard▶ 0:30
A decision model gets a situation and a question with a fixed set of answers. Instead of writing text, it returns a probability for each answer. JevBench asks each open model 1,500 of those questions, most of them kept private so that nobody can train on them. Jev only runs through TypeSafe's service, and in v1.6.0 it answered 600.
JevBench's headline score, Capability, is the average of two parts. Intelligence scores getting decisions right, weighted toward the hard questions and corrected for lucky guesses. Calibration scores whether a model's confidence can be trusted. JevBench calls a model Jev-class when it costs at most twice what Jev cost in the previous version and takes at most twice as long to answer.
JevBench v1.6.0, the top three and Clef-Flash
Capability is the average of the two parts; Jev-class place is among the Jev-class systems, official place among all the ranked systems; JevBench's numbers
Among the Jev-class systems, Quyet-1.0-Large was first with 81.7, deck-31B second with 77.6 and Jev third with 76.5. On JevBench's official ranking, which also weighs price and speed, Quyet was first and Jev second. All of Quyet's lead over Jev was in Intelligence, 73.4 against 62.4. On Calibration, JevBench had them about level, 90.0 against 90.6, with Jev slightly ahead. Jev's two parts are as JevBench shows them, adjusted to the open models' question set.
JevBench Capability, v1.6.0
JevBench's headline score, from zero; Jev-class systems
Who they are▶ 1:52
Chinh Nguyen released Quyet-1.0-Large under the Apache 2.0 licence, so its weights are free to download. It is Google's open model Gemma-4-31B-it, fine-tuned on about a million example decisions in English and Vietnamese. It doesn't write its answers. Its code puts each question into a fixed prompt and reads which letter the model would say next, one letter per option, and turns those letters into probabilities with a confidence setting its author tuned.
deck-31B, from another developer, is the same Gemma at the same version with no training at all. Its makers compressed it to eight bits and read its answers the same way, from the next letter, with a prompt and settings borrowed from another open project, Cygnet. So the first and second on Capability share one base model, one trained and one not. On JevBench they are level on Intelligence, 73.4 and 73.0. The gap between them is in Calibration, 90.0 against 82.2.
Our test, with ten options▶ 2:53
Our test is the one from our Clef vs Jev video: 3,080 messages that customers sent to a bank's support team, from the Banking77 test set, each to be sorted into one of seventy-seven reasons or into none of these, 78 options in all. With each message, every model sees the five most similar messages that were already sorted, with their answers.
Quyet's software only accepts ten options per question. So every model, Jev included, answered a ten-option version: the 9 reasons whose sorted examples look most like the new message, plus none of these, listed in a fixed order that says nothing about which is likeliest. The right reason is on the list for 3,073 of the 3,080 messages, and for the other 7, none of these counts as right. A shorter list should be an easier question, so these numbers don't compare with the full-list ones on our Clef vs Jev page. For Jev it made no difference: it got 93.9% right on the full list and 93.9% on the short one, on the same day.
We fixed the procedure before sending a single test message: which messages, which examples, how answers are scored, how each model's cut-off for sending a message to a person is set, and how costs are counted. Jev ran on TypeSafe's service and Clef-Flash on Cloudflare's Workers AI. Quyet, deck-31B and plain Gemma ran on H100 GPUs we rented on Modal in Europe, each through its own code: Quyet's package, deck-31B's server, and plain Gemma through Quyet's own prompt and readout, so that the only difference between it and Quyet that can change an answer is the training. Every request went one at a time from a Mac mini in Berlin, on the evening of 5 October.
The result▶ 4:01
Our test: share right
The same messages for every model, each with its five most similar sorted messages and ten options
| model | share right | 95% interval | minus Jev | 95% interval | none of these |
|---|---|---|---|---|---|
| Clef-Flash | 95.3% | 94.5 to 96.0 | +1.40 points | +0.84 to +1.98 | 1 |
| Quyet-1.0-Large | 94.5% | 93.6 to 95.2 | +0.58 points | +0.13 to +1.07 | 10 |
| deck-31B | 94.0% | 93.1 to 94.7 | +0.06 points | −0.36 to +0.49 | 24 |
| Jev | 93.9% | 93.0 to 94.7 | 19 | ||
| plain Gemma (Quyet's prompt) | 93.8% | 92.9 to 94.6 | −0.13 points | −0.55 to +0.29 | 31 |
Quyet got 94.5% right and Jev 93.9%. On the same messages, Quyet got 36 right that Jev got wrong, and Jev 18 the other way. That is a lead of +0.58 points, with a 95% interval from +0.13 to +1.07: small, but not one that chance would easily produce.
deck-31B, with no training, got 94.0%, level with Jev (+0.06 points, interval −0.36 to +0.49). Plain Gemma through Quyet's prompt got 93.8%, also level with Jev (−0.13 points). So the training is worth +0.71 points here, with an interval from +0.26 to +1.20.
On JevBench, Quyet is well ahead of Jev. On our test it is ahead too, by about half a point, and deck-31B, second on Capability, is level with Jev. The two are different scales, so the gaps don't compare as numbers.
Our test, ten options: share right
A dot per model with its 95% interval; the window starts at 92%, not zero
The model ranked eleventh▶ 4:51
Clef-Flash, the smaller of Cloudflare's two Clef models, got the most right of the five in our run, 95.3%. That is +1.40 points ahead of Jev, with an interval from +0.84 to +1.98, and +0.81 points ahead of Quyet. We only decided to compare Clef-Flash with Quyet after seeing the results, and a gap picked out after the fact is more likely to be luck, so treat that second gap with more caution. On JevBench's Capability, Clef-Flash is number 11 among the Jev-class systems, with 63.8.
A leaderboard and one task can disagree like this for several reasons. JevBench mixes many kinds of task, from sorting messages like ours to working out how much an insurance policy pays for a leak, and the public questions we read have no sorted examples next to them. It gives the hard questions the most weight, and confidence is half its score. Ours is one task, with five sorted examples next to every message, scored on getting it right. If you count plain right answers on JevBench's multiple-choice questions, Clef-Flash gets 67% and Quyet 87%. On one bank's messages, with those examples, Clef-Flash comes out ahead. We didn't test which of those differences is the one that matters.
Who gets a person▶ 6:02
JevBench had Quyet and Jev about level on confidence. On our test, Quyet's confidence was the best of the five at telling its right answers from its wrong ones: its AUROC is 0.90, against Jev's 0.84. That matters for one job, where the model answers the messages it is sure about and a person checks the rest.
We set each model's cut-off on a separate set of 573 messages, aiming at 2% wrong among the messages it answers alone, then applied it to the test.
Who gets a person
Each model handles the messages above its cut-off and sends the rest to a person; cut-offs set on a separate set, aiming at 2% wrong
Quyet handled 91.6% of the messages on its own, and 2.45% of those were wrong, a little over the target. Jev handled 83.8%, with 1.94% wrong, inside it. Clef-Flash handled 93.2%, with 2.79% wrong. Plain Gemma handled only 64.0%: its answers were nearly as good as Quyet's, but its confidence was much worse at sorting them (AUROC 0.79). So Quyet's training, with the confidence setting fitted to it, moved its accuracy by under a point, and the share it could handle on its own from 64.0% to 91.6%.
deck-31B, read with its borrowed settings, handled 1.2%. On the separate set, its 9 most confident answers were right and the next was wrong, and further down its error never came back down to the target: the lowest it reached was 2.7%. So the lowest cut-off that met the target sat near the very top. Its temperature setting isn't the cause. With the temperature removed, the same procedure still leaves deck-31B answering almost nothing on its own.
Price and speed▶ 7:23
Price and speed
Per thousand decisions on this ten-option question; times from a Mac mini in Berlin, one request at a time
| model | dollars per thousand | how it is counted | median time | 95th percentile |
|---|---|---|---|---|
| Jev | $0.026 | TypeSafe's price list × tokens | 229 ms | 282 ms |
| Clef-Flash | $0.042 | Workers AI price list × tokens | 290 ms | 886 ms |
| Quyet-1.0-Large | $0.602 | our H100 while the test ran | 431 ms | 459 ms |
| plain Gemma (Quyet's prompt) | $0.658 | our H100 while the test ran | 461 ms | 496 ms |
| deck-31B | $0.817 | our H100 while the test ran | 539 ms | 845 ms |
On this ten-option question, Jev costs $0.026 per thousand decisions and Clef-Flash $0.042, each at its list price for the tokens counted. Nobody we could find sells Quyet as a service, so it has no price. We ran it on a rented GPU, sending one message at a time, the way a speed test does. Counting only the GPU time while our test messages ran, that came to $0.602 per thousand, 23× what Jev costs. Modal's whole bill for Quyet was $4.35, $2.30 of it for start-up, warm-ups and idle time, and spread over all 3,853 calls we sent it that is $1.13 per thousand. Sending many messages at once would bring both figures down. We didn't measure by how much.
JevBench lists Quyet at $0.045 per thousand. That is JevBench's estimate, based on what plain Gemma costs on OpenRouter, a service that sells it for writing text, times the tokens in an earlier run. On that basis, with the tokens of our run, Quyet would be $0.024. No one sells it at that price.
Dollars per thousand decisions
Our test, ten options
From Berlin, Jev answered in 229 ms at the median, Clef-Flash in 290 ms and Quyet in 431 ms. Quyet itself took only 66 ms on our GPU; the rest was the trip to the GPU and back, through Modal. For Quyet, as for every model it runs itself, JevBench takes the time it measured (116 ms), doubles it and adds fifteen hundredths of a second, which it calls an assumption, and lists 382 ms.
Time per request, from Berlin
The bar runs to the median; the ticks mark p95 and p99. One request at a time.
What this doesn't settle▶ 8:47
This was one task, in English, in a ten-option version, on one day, from one city. The open models ran on our rented GPUs, so their speed and price describe our setup. None of the makers we checked says whether this bank dataset was in its training. Quyet's author at least lists public classification datasets in its mix, without naming them.
The second and third on Capability were closer than they looked. JevBench v1.6.0 measured Jev through its own service on 600 questions and the open models on 1,500. When it checked Jev a second time, on new questions, Jev came out at 77.5, level with deck-31B. JevBench says that is within the margin of its first measurement, and it kept the official number.
After our run, JevBench published v1.6.1, scored on 6 October at 00:39 UTC. It tests hosted APIs such as Jev on the full question set, as it does the open models, and no longer adjusts their scores. Its official ranking is now Sage 1.3.0 from Levanto Labs first, Jev second and Quyet third, so the video's line that Quyet is still first on the official ranking and Jev second describes v1.6.0. On Capability, Quyet is still first with 81.7. Sage, which v1.6.0 had left out of the Jev-class list on cost, is now second with 78.6, Jev's score is 77.1, and Clef-Flash is number 12. Our test didn't include Sage. The video's leaderboard numbers are v1.6.0's, and that version is unchanged and still online.
JevBench v1.6.0 and v1.6.1 side by side
v1.6.0 is the version the video reports, unchanged and still online; v1.6.1 was published after our run. JevBench's numbers
Which to use▶ 9:34
If you sort messages and can show the model a few sorted examples, Clef-Flash got the most right in our run. If a person checks whatever the model isn't sure about, Quyet's confidence was the best of the five at telling right from wrong; only Clef-Flash handed fewer messages to a person, and it made more mistakes doing it. For the smallest bill and the quickest answer from Berlin, Jev. To run a model on your own GPU, Quyet, or plain Gemma if you can live with weaker confidence for nearly the same accuracy. Clef-Flash's weights are open too, but we only measured it on Cloudflare's service.
On this test, the new number one does beat Jev, by about half a point. And Clef-Flash, number 11 on JevBench's Jev-class list and 18 Capability points behind Quyet there, got more right than both in our run.