model fatıgue
Measured

Does JevBench's new #1 really beat Jev?

Published 7 Oct 2026Our test: bank-support messages, ten options

Pressing play loads the video from YouTube (Google), in privacy-enhanced mode. Privacy

This is the video written out, with every figure in full, each linked to its source in the table below.

A new number one▶ 0:00

JevBench is a leaderboard of decision models like Jev, made by Benchmark Heaven. On Saturday, 3 October, and Sunday, 4 October, people kept sending it requests, on its GitHub page, to rank new models. One of Sunday's requests had five models in it, from a developer called Chinh Nguyen. On Monday 5 October the leaderboard's maker put out v1.6.0, with all 92 of its ranked systems tested from scratch on new questions, and the biggest of those five, Quyet-1.0-Large, came out first on its headline score, ahead of Jev.

So we tested it. Does the new number one beat Jev on someone else's test too, or only on JevBench's?

The leaderboard▶ 0:30

A decision model gets a situation and a question with a fixed set of answers. Instead of writing text, it returns a probability for each answer. JevBench asks each open model 1,500 of those questions, most of them kept private so that nobody can train on them. Jev only runs through TypeSafe's service, and in v1.6.0 it answered 600.

JevBench's headline score, Capability, is the average of two parts. Intelligence scores getting decisions right, weighted toward the hard questions and corrected for lucky guesses. Calibration scores whether a model's confidence can be trusted. JevBench calls a model Jev-class when it costs at most twice what Jev cost in the previous version and takes at most twice as long to answer.

JevBench v1.6.0, the top three and Clef-Flash

Capability is the average of the two parts; Jev-class place is among the Jev-class systems, official place among all the ranked systems; JevBench's numbers

modelCapability95% intervalIntelligenceCalibrationJev-class placeofficial placequestions
Quyet-1.0-Large81.778.0 to 83.673.490.0111,500
deck-31B77.673.8 to 80.273.082.2251,500
Jev76.569.2 to 78.862.490.632600
Clef-Flash63.859.4 to 65.542.285.511211,500

Among the Jev-class systems, Quyet-1.0-Large was first with 81.7, deck-31B second with 77.6 and Jev third with 76.5. On JevBench's official ranking, which also weighs price and speed, Quyet was first and Jev second. All of Quyet's lead over Jev was in Intelligence, 73.4 against 62.4. On Calibration, JevBench had them about level, 90.0 against 90.6, with Jev slightly ahead. Jev's two parts are as JevBench shows them, adjusted to the open models' question set.

JevBench Capability, v1.6.0

JevBench's headline score, from zero; Jev-class systems

0255075100CapabilityQuyet-1.0-LargeQuyet-1.0-Large: 81.7 · JevBench (Benchmark Heaven), read 5 Oct 2026, 20:18 CEST81.7deck-31Bdeck-31B: 77.6 · JevBench (Benchmark Heaven), read 5 Oct 2026, 20:18 CEST77.6JevJev: 76.5 · JevBench (Benchmark Heaven), read 5 Oct 2026, 20:18 CEST76.5Clef-FlashClef-Flash: 63.8 · JevBench (Benchmark Heaven), read 5 Oct 2026, 20:18 CEST63.8
0255075100CapabilityQuyet-1.0-LargeQuyet-1.0-Large: 81.7 · JevBench (Benchmark Heaven), read 5 Oct 2026, 20:18 CEST81.7deck-31Bdeck-31B: 77.6 · JevBench (Benchmark Heaven), read 5 Oct 2026, 20:18 CEST77.6JevJev: 76.5 · JevBench (Benchmark Heaven), read 5 Oct 2026, 20:18 CEST76.5Clef-FlashClef-Flash: 63.8 · JevBench (Benchmark Heaven), read 5 Oct 2026, 20:18 CEST63.8
JevBench's numbers, v1.6.0 of 5 October. A different scale from our test's share right.
The numbers in this chart
Capability
Quyet-1.0-Large81.7
deck-31B77.6
Jev76.5
Clef-Flash63.8

Who they are▶ 1:52

Chinh Nguyen released Quyet-1.0-Large under the Apache 2.0 licence, so its weights are free to download. It is Google's open model Gemma-4-31B-it, fine-tuned on about a million example decisions in English and Vietnamese. It doesn't write its answers. Its code puts each question into a fixed prompt and reads which letter the model would say next, one letter per option, and turns those letters into probabilities with a confidence setting its author tuned.

deck-31B, from another developer, is the same Gemma at the same version with no training at all. Its makers compressed it to eight bits and read its answers the same way, from the next letter, with a prompt and settings borrowed from another open project, Cygnet. So the first and second on Capability share one base model, one trained and one not. On JevBench they are level on Intelligence, 73.4 and 73.0. The gap between them is in Calibration, 90.0 against 82.2.

Our test, with ten options▶ 2:53

Our test is the one from our Clef vs Jev video: 3,080 messages that customers sent to a bank's support team, from the Banking77 test set, each to be sorted into one of seventy-seven reasons or into none of these, 78 options in all. With each message, every model sees the five most similar messages that were already sorted, with their answers.

Quyet's software only accepts ten options per question. So every model, Jev included, answered a ten-option version: the 9 reasons whose sorted examples look most like the new message, plus none of these, listed in a fixed order that says nothing about which is likeliest. The right reason is on the list for 3,073 of the 3,080 messages, and for the other 7, none of these counts as right. A shorter list should be an easier question, so these numbers don't compare with the full-list ones on our Clef vs Jev page. For Jev it made no difference: it got 93.9% right on the full list and 93.9% on the short one, on the same day.

We fixed the procedure before sending a single test message: which messages, which examples, how answers are scored, how each model's cut-off for sending a message to a person is set, and how costs are counted. Jev ran on TypeSafe's service and Clef-Flash on Cloudflare's Workers AI. Quyet, deck-31B and plain Gemma ran on H100 GPUs we rented on Modal in Europe, each through its own code: Quyet's package, deck-31B's server, and plain Gemma through Quyet's own prompt and readout, so that the only difference between it and Quyet that can change an answer is the training. Every request went one at a time from a Mac mini in Berlin, on the evening of 5 October.

The result▶ 4:01

Our test: share right

The same messages for every model, each with its five most similar sorted messages and ten options

modelshare right95% intervalminus Jev95% intervalnone of these
Clef-Flash95.3%94.5 to 96.0+1.40 points+0.84 to +1.981
Quyet-1.0-Large94.5%93.6 to 95.2+0.58 points+0.13 to +1.0710
deck-31B94.0%93.1 to 94.7+0.06 points−0.36 to +0.4924
Jev93.9%93.0 to 94.719
plain Gemma (Quyet's prompt)93.8%92.9 to 94.6−0.13 points−0.55 to +0.2931

Quyet got 94.5% right and Jev 93.9%. On the same messages, Quyet got 36 right that Jev got wrong, and Jev 18 the other way. That is a lead of +0.58 points, with a 95% interval from +0.13 to +1.07: small, but not one that chance would easily produce.

deck-31B, with no training, got 94.0%, level with Jev (+0.06 points, interval −0.36 to +0.49). Plain Gemma through Quyet's prompt got 93.8%, also level with Jev (−0.13 points). So the training is worth +0.71 points here, with an interval from +0.26 to +1.20.

On JevBench, Quyet is well ahead of Jev. On our test it is ahead too, by about half a point, and deck-31B, second on Capability, is level with Jev. The two are different scales, so the gaps don't compare as numbers.

Our test, ten options: share right

A dot per model with its 95% interval; the window starts at 92%, not zero

measured95% interval
92%93%94%95%96%97%share rightQuyet-1.0-LargeQuyet-1.0-Large: 94.5%, interval 93.6 to 95.2 · Model Fatigue, run 5 Oct 21:40 to 6 Oct 2026 00:01 CEST94.5%deck-31Bdeck-31B: 94.0%, interval 93.1 to 94.7 · Model Fatigue, run 5 Oct 21:40 to 6 Oct 2026 00:01 CEST94.0%JevJev: 93.9%, interval 93.0 to 94.7 · Model Fatigue, run 5 Oct 21:40 to 6 Oct 2026 00:01 CEST93.9%plain Gemma (Quyet's prompt)plain Gemma (Quyet's prompt): 93.8%, interval 92.9 to 94.6 · Model Fatigue, run 5 Oct 21:40 to 6 Oct 2026 00:01 CEST93.8%Clef-FlashClef-Flash: 95.3%, interval 94.5 to 96.0 · Model Fatigue, run 5 Oct 21:40 to 6 Oct 2026 00:01 CEST95.3%
92%93%94%95%96%97%share rightQuyet-1.0-LargeQuyet-1.0-Large: 94.5%, interval 93.6 to 95.2 · Model Fatigue, run 5 Oct 21:40 to 6 Oct 2026 00:01 CEST94.5%deck-31Bdeck-31B: 94.0%, interval 93.1 to 94.7 · Model Fatigue, run 5 Oct 21:40 to 6 Oct 2026 00:01 CEST94.0%JevJev: 93.9%, interval 93.0 to 94.7 · Model Fatigue, run 5 Oct 21:40 to 6 Oct 2026 00:01 CEST93.9%plain Gemma (Quyet's prompt)plain Gemma (Quyet's prompt): 93.8%, interval 92.9 to 94.6 · Model Fatigue, run 5 Oct 21:40 to 6 Oct 2026 00:01 CEST93.8%Clef-FlashClef-Flash: 95.3%, interval 94.5 to 96.0 · Model Fatigue, run 5 Oct 21:40 to 6 Oct 2026 00:01 CEST95.3%
Our run of 5 October. The interval is for one model's share right; the gaps between models have their own intervals, in the table above.
The numbers in this chart
share right95% interval
Quyet-1.0-Large94.5%93.6 to 95.2
deck-31B94.0%93.1 to 94.7
Jev93.9%93.0 to 94.7
plain Gemma (Quyet's prompt)93.8%92.9 to 94.6
Clef-Flash95.3%94.5 to 96.0

The model ranked eleventh▶ 4:51

Clef-Flash, the smaller of Cloudflare's two Clef models, got the most right of the five in our run, 95.3%. That is +1.40 points ahead of Jev, with an interval from +0.84 to +1.98, and +0.81 points ahead of Quyet. We only decided to compare Clef-Flash with Quyet after seeing the results, and a gap picked out after the fact is more likely to be luck, so treat that second gap with more caution. On JevBench's Capability, Clef-Flash is number 11 among the Jev-class systems, with 63.8.

A leaderboard and one task can disagree like this for several reasons. JevBench mixes many kinds of task, from sorting messages like ours to working out how much an insurance policy pays for a leak, and the public questions we read have no sorted examples next to them. It gives the hard questions the most weight, and confidence is half its score. Ours is one task, with five sorted examples next to every message, scored on getting it right. If you count plain right answers on JevBench's multiple-choice questions, Clef-Flash gets 67% and Quyet 87%. On one bank's messages, with those examples, Clef-Flash comes out ahead. We didn't test which of those differences is the one that matters.

Who gets a person▶ 6:02

JevBench had Quyet and Jev about level on confidence. On our test, Quyet's confidence was the best of the five at telling its right answers from its wrong ones: its AUROC is 0.90, against Jev's 0.84. That matters for one job, where the model answers the messages it is sure about and a person checks the rest.

We set each model's cut-off on a separate set of 573 messages, aiming at 2% wrong among the messages it answers alone, then applied it to the test.

Who gets a person

Each model handles the messages above its cut-off and sends the rest to a person; cut-offs set on a separate set, aiming at 2% wrong

modelcut-offhandled alonewrong among thoseAUROCcalibration error
Clef-Flash0.80093.2%2.79%0.860.046
Quyet-1.0-Large0.88091.6%2.45%0.900.018
Jev0.99083.8%1.94%0.840.034
plain Gemma (Quyet's prompt)0.99964.0%2.49%0.790.047
deck-31B0.9961.2%2.63%0.710.040

Quyet handled 91.6% of the messages on its own, and 2.45% of those were wrong, a little over the target. Jev handled 83.8%, with 1.94% wrong, inside it. Clef-Flash handled 93.2%, with 2.79% wrong. Plain Gemma handled only 64.0%: its answers were nearly as good as Quyet's, but its confidence was much worse at sorting them (AUROC 0.79). So Quyet's training, with the confidence setting fitted to it, moved its accuracy by under a point, and the share it could handle on its own from 64.0% to 91.6%.

deck-31B, read with its borrowed settings, handled 1.2%. On the separate set, its 9 most confident answers were right and the next was wrong, and further down its error never came back down to the target: the lowest it reached was 2.7%. So the lowest cut-off that met the target sat near the very top. Its temperature setting isn't the cause. With the temperature removed, the same procedure still leaves deck-31B answering almost nothing on its own.

Price and speed▶ 7:23

Price and speed

Per thousand decisions on this ten-option question; times from a Mac mini in Berlin, one request at a time

modeldollars per thousandhow it is countedmedian time95th percentile
Jev$0.026TypeSafe's price list × tokens229 ms282 ms
Clef-Flash$0.042Workers AI price list × tokens290 ms886 ms
Quyet-1.0-Large$0.602our H100 while the test ran431 ms459 ms
plain Gemma (Quyet's prompt)$0.658our H100 while the test ran461 ms496 ms
deck-31B$0.817our H100 while the test ran539 ms845 ms

On this ten-option question, Jev costs $0.026 per thousand decisions and Clef-Flash $0.042, each at its list price for the tokens counted. Nobody we could find sells Quyet as a service, so it has no price. We ran it on a rented GPU, sending one message at a time, the way a speed test does. Counting only the GPU time while our test messages ran, that came to $0.602 per thousand, 23× what Jev costs. Modal's whole bill for Quyet was $4.35, $2.30 of it for start-up, warm-ups and idle time, and spread over all 3,853 calls we sent it that is $1.13 per thousand. Sending many messages at once would bring both figures down. We didn't measure by how much.

JevBench lists Quyet at $0.045 per thousand. That is JevBench's estimate, based on what plain Gemma costs on OpenRouter, a service that sells it for writing text, times the tokens in an earlier run. On that basis, with the tokens of our run, Quyet would be $0.024. No one sells it at that price.

Dollars per thousand decisions

Our test, ten options

$0$0.225$0.45$0.675$0.9dollars per thousand decisionsJevJev: $0.026 · Model Fatigue, run 5 Oct 21:40 to 6 Oct 2026 00:01 CEST$0.026Clef-FlashClef-Flash: $0.042 · Model Fatigue, run 5 Oct 21:40 to 6 Oct 2026 00:01 CEST$0.042Quyet-1.0-LargeQuyet-1.0-Large: $0.602 · Model Fatigue, run 5 Oct 21:40 to 6 Oct 2026 00:01 CEST$0.602plain Gemma (Quyet's prompt)plain Gemma (Quyet's prompt): $0.658 · Model Fatigue, run 5 Oct 21:40 to 6 Oct 2026 00:01 CEST$0.658deck-31Bdeck-31B: $0.817 · Model Fatigue, run 5 Oct 21:40 to 6 Oct 2026 00:01 CEST$0.817
$0$0.225$0.45$0.675$0.9dollars per thousand decisionsJevJev: $0.026 · Model Fatigue, run 5 Oct 21:40 to 6 Oct 2026 00:01 CEST$0.026Clef-FlashClef-Flash: $0.042 · Model Fatigue, run 5 Oct 21:40 to 6 Oct 2026 00:01 CEST$0.042Quyet-1.0-LargeQuyet-1.0-Large: $0.602 · Model Fatigue, run 5 Oct 21:40 to 6 Oct 2026 00:01 CEST$0.602plain Gemma (Quyet's prompt)plain Gemma (Quyet's prompt): $0.658 · Model Fatigue, run 5 Oct 21:40 to 6 Oct 2026 00:01 CEST$0.658deck-31Bdeck-31B: $0.817 · Model Fatigue, run 5 Oct 21:40 to 6 Oct 2026 00:01 CEST$0.817
Filled: price list times tokens counted. Outlined: our rented GPU time while the test ran, one request at a time, before start-up and idle time.
The numbers in this chart
dollars per thousand decisions
Jev$0.026
Clef-Flash$0.042
Quyet-1.0-Large$0.602
plain Gemma (Quyet's prompt)$0.658
deck-31B$0.817

From Berlin, Jev answered in 229 ms at the median, Clef-Flash in 290 ms and Quyet in 431 ms. Quyet itself took only 66 ms on our GPU; the rest was the trip to the GPU and back, through Modal. For Quyet, as for every model it runs itself, JevBench takes the time it measured (116 ms), doubles it and adds fifteen hundredths of a second, which it calls an assumption, and lists 382 ms.

Time per request, from Berlin

The bar runs to the median; the ticks mark p95 and p99. One request at a time.

median95th and 99th percentile
0 s0.32 s0.64 s0.96 s1.28 s1.6 sseconds per requestJevJev: median 229 ms, 95th percentile 282 ms, 99th 348 ms · Model Fatigue, run 5 Oct 21:40 to 6 Oct 2026 00:01 CEST229 msClef-FlashClef-Flash: median 290 ms, 95th percentile 886 ms, 99th 1,520 ms · Model Fatigue, run 5 Oct 21:40 to 6 Oct 2026 00:01 CEST290 msQuyet-1.0-LargeQuyet-1.0-Large: median 431 ms, 95th percentile 459 ms, 99th 580 ms · Model Fatigue, run 5 Oct 21:40 to 6 Oct 2026 00:01 CEST431 msplain Gemma (Quyet's prompt)plain Gemma (Quyet's prompt): median 461 ms, 95th percentile 496 ms, 99th 831 ms · Model Fatigue, run 5 Oct 21:40 to 6 Oct 2026 00:01 CEST461 msdeck-31Bdeck-31B: median 539 ms, 95th percentile 845 ms, 99th 1,186 ms · Model Fatigue, run 5 Oct 21:40 to 6 Oct 2026 00:01 CEST539 ms
0 s0.32 s0.64 s0.96 s1.28 s1.6 sseconds per requestJevJev: median 229 ms, 95th percentile 282 ms, 99th 348 ms · Model Fatigue, run 5 Oct 21:40 to 6 Oct 2026 00:01 CEST229 msClef-FlashClef-Flash: median 290 ms, 95th percentile 886 ms, 99th 1,520 ms · Model Fatigue, run 5 Oct 21:40 to 6 Oct 2026 00:01 CEST290 msQuyet-1.0-LargeQuyet-1.0-Large: median 431 ms, 95th percentile 459 ms, 99th 580 ms · Model Fatigue, run 5 Oct 21:40 to 6 Oct 2026 00:01 CEST431 msplain Gemma (Quyet's prompt)plain Gemma (Quyet's prompt): median 461 ms, 95th percentile 496 ms, 99th 831 ms · Model Fatigue, run 5 Oct 21:40 to 6 Oct 2026 00:01 CEST461 msdeck-31Bdeck-31B: median 539 ms, 95th percentile 845 ms, 99th 1,186 ms · Model Fatigue, run 5 Oct 21:40 to 6 Oct 2026 00:01 CEST539 ms
Quyet, plain Gemma and deck-31B ran on GPUs we rented on Modal, so their times describe our setup. Another job loaded our calling machine during deck-31B's pass, so only its median compares.
The numbers in this chart
median95th percentile99th percentile
Jev229 ms282 ms348 ms
Clef-Flash290 ms886 ms1,520 ms
Quyet-1.0-Large431 ms459 ms580 ms
plain Gemma (Quyet's prompt)461 ms496 ms831 ms
deck-31B539 ms845 ms1,186 ms

What this doesn't settle▶ 8:47

This was one task, in English, in a ten-option version, on one day, from one city. The open models ran on our rented GPUs, so their speed and price describe our setup. None of the makers we checked says whether this bank dataset was in its training. Quyet's author at least lists public classification datasets in its mix, without naming them.

The second and third on Capability were closer than they looked. JevBench v1.6.0 measured Jev through its own service on 600 questions and the open models on 1,500. When it checked Jev a second time, on new questions, Jev came out at 77.5, level with deck-31B. JevBench says that is within the margin of its first measurement, and it kept the official number.

After our run, JevBench published v1.6.1, scored on 6 October at 00:39 UTC. It tests hosted APIs such as Jev on the full question set, as it does the open models, and no longer adjusts their scores. Its official ranking is now Sage 1.3.0 from Levanto Labs first, Jev second and Quyet third, so the video's line that Quyet is still first on the official ranking and Jev second describes v1.6.0. On Capability, Quyet is still first with 81.7. Sage, which v1.6.0 had left out of the Jev-class list on cost, is now second with 78.6, Jev's score is 77.1, and Clef-Flash is number 12. Our test didn't include Sage. The video's leaderboard numbers are v1.6.0's, and that version is unchanged and still online.

JevBench v1.6.0 and v1.6.1 side by side

v1.6.0 is the version the video reports, unchanged and still online; v1.6.1 was published after our run. JevBench's numbers

modelCapability, v1.6.0Capability, v1.6.1Jev-class place, v1.6.0Jev-class place, v1.6.1official place, v1.6.0official place, v1.6.1
Sage 1.3.077.178.6outside the cost cap291
Quyet-1.0-Large81.781.71113
deck-31B77.677.62357
Jev76.577.13422
Clef-Flash63.863.811122121

Which to use▶ 9:34

If you sort messages and can show the model a few sorted examples, Clef-Flash got the most right in our run. If a person checks whatever the model isn't sure about, Quyet's confidence was the best of the five at telling right from wrong; only Clef-Flash handed fewer messages to a person, and it made more mistakes doing it. For the smallest bill and the quickest answer from Berlin, Jev. To run a model on your own GPU, Quyet, or plain Gemma if you can live with weaker confidence for nearly the same accuracy. Clef-Flash's weights are open too, but we only measured it on Cloudflare's service.

On this test, the new number one does beat Jev, by about half a point. And Clef-Flash, number 11 on JevBench's Jev-class list and 18 Capability points behind Quyet there, got more right than both in our run.

Every number

These are all 186 figures behind the video and this page, grouped by whose they are, with the page each came from and when we read it. Figures marked ⟳ can move. When a re-read finds a change, the new value shows next to the one from the video.

Model Fatigue

Our run of the 10-option question: Jev, Clef-Flash, Quyet-1.0-Large, deck-31B and plain Gemma · run 5 Oct 21:40 to 6 Oct 2026 00:01 CEST
Banking77 test messages every model answered3,080
Jev on the full 78-option question, the same day (its own pass)93.9%
Quyet-1.0-Large, share right94.5%
Quyet-1.0-Large, share right, 95% interval, low93.6
Quyet-1.0-Large, share right, 95% interval, high95.2
Quyet-1.0-Large, messages sorted correctly2,910
Quyet-1.0-Large, times it answered none of these10
Quyet-1.0-Large, median time per request from Berlin431 ms
Quyet-1.0-Large, 95th percentile time per request from Berlin459 ms
Quyet-1.0-Large, 99th percentile time per request from Berlin580 ms
Quyet-1.0-Large, dollars per 1,000 decisions (our GPU time while the messages ran, as Modal billed it, one request at a time)$0.602
Quyet-1.0-Large, calibration error of its probabilities (ECE; lower is closer)0.018
Quyet-1.0-Large, how well its probabilities separate its right answers from its wrong ones (AUROC)0.90
Quyet-1.0-Large, cut-off on its top probability, set on the separate set0.880
Quyet-1.0-Large, share of messages it handled without a person91.6%
Quyet-1.0-Large, wrong among the messages it handled without a person2.45%
deck-31B, share right94.0%
deck-31B, share right, 95% interval, low93.1
deck-31B, share right, 95% interval, high94.7
deck-31B, messages sorted correctly2,894
deck-31B, times it answered none of these24
deck-31B, median time per request from Berlin539 ms
deck-31B, 95th percentile time per request from Berlin845 ms
deck-31B, 99th percentile time per request from Berlin1,186 ms
deck-31B, dollars per 1,000 decisions (our GPU time while the messages ran, as Modal billed it, one request at a time)$0.817
deck-31B, calibration error of its probabilities (ECE; lower is closer)0.040
deck-31B, how well its probabilities separate its right answers from its wrong ones (AUROC)0.71
deck-31B, cut-off on its top probability, set on the separate set0.996
deck-31B, share of messages it handled without a person1.2%
deck-31B, wrong among the messages it handled without a person2.63%
Jev, share right93.9%
Jev, share right, 95% interval, low93.0
Jev, share right, 95% interval, high94.7
Jev, messages sorted correctly2,892
Jev, times it answered none of these19
Jev, median time per request from Berlin229 ms
Jev, 95th percentile time per request from Berlin282 ms
Jev, 99th percentile time per request from Berlin348 ms
Jev, dollars per 1,000 decisions (price list × input tokens counted)$0.026
Jev, calibration error of its probabilities (ECE; lower is closer)0.034
Jev, how well its probabilities separate its right answers from its wrong ones (AUROC)0.84
Jev, cut-off on its top probability, set on the separate set0.990
Jev, share of messages it handled without a person83.8%
Jev, wrong among the messages it handled without a person1.94%
Gemma-4-31B-it through Quyet's prompt, share right93.8%
Gemma-4-31B-it through Quyet's prompt, share right, 95% interval, low92.9
Gemma-4-31B-it through Quyet's prompt, share right, 95% interval, high94.6
Gemma-4-31B-it through Quyet's prompt, messages sorted correctly2,888
Gemma-4-31B-it through Quyet's prompt, times it answered none of these31
Gemma-4-31B-it through Quyet's prompt, median time per request from Berlin461 ms
Gemma-4-31B-it through Quyet's prompt, 95th percentile time per request from Berlin496 ms
Gemma-4-31B-it through Quyet's prompt, 99th percentile time per request from Berlin831 ms
Gemma-4-31B-it through Quyet's prompt, dollars per 1,000 decisions (our GPU time while the messages ran, as Modal billed it, one request at a time)$0.658
Gemma-4-31B-it through Quyet's prompt, calibration error of its probabilities (ECE; lower is closer)0.047
Gemma-4-31B-it through Quyet's prompt, how well its probabilities separate its right answers from its wrong ones (AUROC)0.79
Gemma-4-31B-it through Quyet's prompt, cut-off on its top probability, set on the separate set0.999
Gemma-4-31B-it through Quyet's prompt, share of messages it handled without a person64.0%
Gemma-4-31B-it through Quyet's prompt, wrong among the messages it handled without a person2.49%
Clef-Flash, share right95.3%
Clef-Flash, share right, 95% interval, low94.5
Clef-Flash, share right, 95% interval, high96.0
Clef-Flash, messages sorted correctly2,935
Clef-Flash, times it answered none of these1
Clef-Flash, median time per request from Berlin290 ms
Clef-Flash, 95th percentile time per request from Berlin886 ms
Clef-Flash, 99th percentile time per request from Berlin1,520 ms
Clef-Flash, dollars per 1,000 decisions (price list × input tokens counted)$0.042
Clef-Flash, calibration error of its probabilities (ECE; lower is closer)0.046
Clef-Flash, how well its probabilities separate its right answers from its wrong ones (AUROC)0.86
Clef-Flash, cut-off on its top probability, set on the separate set0.800
Clef-Flash, share of messages it handled without a person93.2%
Clef-Flash, wrong among the messages it handled without a person2.79%
Messages Quyet-1.0-Large got right and Jev got wrong36
Messages Jev got right and Quyet-1.0-Large got wrong18
Messages deck-31B got right and Jev got wrong23
Messages Jev got right and deck-31B got wrong21
Messages Gemma-4-31B-it through Quyet's prompt got right and Jev got wrong21
Messages Jev got right and Gemma-4-31B-it through Quyet's prompt got wrong25
Messages Clef-Flash got right and Jev got wrong60
Messages Jev got right and Clef-Flash got wrong17
Quyet per 1,000 on JevBench's basis: OpenRouter's list price for plain Gemma × the tokens of our run (our arithmetic)$0.024
Messages whose right reason is not on the list (none of these counts as right)7
Quyet-1.0-Large, cents per 1,000 decisions (our arithmetic)60.2
deck-31B, cents per 1,000 decisions (our arithmetic)81.7
Jev, cents per 1,000 decisions (our arithmetic)2.6
Gemma-4-31B-it through Quyet's prompt, cents per 1,000 decisions (our arithmetic)65.8
Clef-Flash, cents per 1,000 decisions (our arithmetic)4.2
Quyet-1.0-Large, median time the model itself took on our GPU (its server's count)66 ms
Gemma-4-31B-it through Quyet's prompt, median time the model itself took on our GPU (its server's count)68 ms
deck-31B, median time the model itself took on our GPU (its server's count)169 ms
Quyet-1.0-Large minus Jev, share right, same messages+0.58 points
Quyet-1.0-Large minus Jev, 95% interval, low+0.13
Quyet-1.0-Large minus Jev, 95% interval, high+1.07
deck-31B minus Jev, share right, same messages+0.06 points
deck-31B minus Jev, 95% interval, low−0.36
deck-31B minus Jev, 95% interval, high+0.49
Gemma-4-31B-it through Quyet's prompt minus Jev, share right, same messages−0.13 points
Gemma-4-31B-it through Quyet's prompt minus Jev, 95% interval, low−0.55
Gemma-4-31B-it through Quyet's prompt minus Jev, 95% interval, high+0.29
Clef-Flash minus Jev, share right, same messages+1.40 points
Clef-Flash minus Jev, 95% interval, low+0.84
Clef-Flash minus Jev, 95% interval, high+1.98
Quyet minus plain Gemma through Quyet's prompt: what the training is worth+0.71 points
The training's worth, 95% interval, low+0.26
The training's worth, 95% interval, high+1.20
Clef-Flash minus Quyet, share right (decided after the run: post hoc)+0.81 points
Quyet minus plain Gemma: more messages in 100 handled without a person (our arithmetic)+28
Quyet's cost per decision as we ran it against Jev's (our arithmetic)23×

Model Fatigue

Our protocol addendum for the 10-option question, fixed before any test message was sent · frozen 5 Oct 2026, 21:38 CEST
Options in each question: the 9 nearest reasons plus none of these10
Reasons on each question's list, besides none of these9
Wrong answers allowed among those a model handles alone, the target every cut-off was set for2%

Model Fatigue

Our results write-up (RESULTS.md, section "Who beats Jev now?") · written 6 Oct 2026
Options in the full question (77 reasons plus none of these)78
Messages whose right reason is on its ten-option list3,073

TypeSafe

TypeSafe's models page (Jev's price) · read 6 Oct 2026, 11:00 CEST
Jev, price per million input tokens$0.042

Model Fatigue

Modal's bill for our GPU apps, read back per pass (wbj_gpu_cost.json) · billed 6 Oct 2026, 01:05 CEST
Everything Modal billed for Quyet's GPU app$4.35
Of that, start-up, warm-ups and idle time$2.30
Calls Quyet answered in all (the separate set, the test and a repeat pass)3,853
Quyet's whole bill spread over every call it answered, dollars per 1,000 (our arithmetic)$1.13

Model Fatigue

Our pass over the separate set of 573 messages the cut-offs were set on (dev records, our arithmetic in derived.json) · run 5 Oct 2026, 20:35–21:30 CEST
Messages in the separate set every cut-off was set on573
deck-31B's most confident answers on the separate set that were all right before the first wrong one9
deck-31B, the lowest error rate any cut-off below that point reaches on the separate set2.7%

JevBench (Benchmark Heaven)

Systems JevBench v1.6.0 ranks92
Quyet-1.0-Large, JevBench Capability81.7
Quyet-1.0-Large, JevBench Capability, 95% interval, low78.0
Quyet-1.0-Large, JevBench Capability, 95% interval, high83.6
Quyet-1.0-Large, JevBench's right-answers part (Intelligence)73.4
Quyet-1.0-Large, JevBench's confidence part (Calibration)90.0
Quyet-1.0-Large, place among Jev-class systems on Capability1
Quyet-1.0-Large, place on JevBench's official ranking1
Quyet-1.0-Large, JevBench questions it answered1,500
Quyet-1.0-Large, share right on JevBench's multiple-choice questions87%
deck-31B, JevBench Capability77.6
deck-31B, JevBench Capability, 95% interval, low73.8
deck-31B, JevBench Capability, 95% interval, high80.2
deck-31B, JevBench's right-answers part (Intelligence)73.0
deck-31B, JevBench's confidence part (Calibration)82.2
deck-31B, place among Jev-class systems on Capability2
deck-31B, place on JevBench's official ranking5
deck-31B, JevBench questions it answered1,500
deck-31B, share right on JevBench's multiple-choice questions86%
Jev, JevBench Capability76.5
Jev, JevBench Capability, 95% interval, low69.2
Jev, JevBench Capability, 95% interval, high78.8
Jev, JevBench's right-answers part (Intelligence)62.4
Jev, JevBench's confidence part (Calibration)90.6
Jev, place among Jev-class systems on Capability3
Jev, place on JevBench's official ranking2
Jev, JevBench questions it answered600
Jev, share right on JevBench's multiple-choice questions82%
Clef-Flash, JevBench Capability63.8
Clef-Flash, JevBench Capability, 95% interval, low59.4
Clef-Flash, JevBench Capability, 95% interval, high65.5
Clef-Flash, JevBench's right-answers part (Intelligence)42.2
Clef-Flash, JevBench's confidence part (Calibration)85.5
Clef-Flash, place among Jev-class systems on Capability11
Clef-Flash, place on JevBench's official ranking21
Clef-Flash, JevBench questions it answered1,500
Clef-Flash, share right on JevBench's multiple-choice questions67%
Quyet, JevBench's price per 1,000 decisions (its estimate)$0.045
Quyet, JevBench's median time (after its adjustment)382 ms
Quyet, the median time JevBench measured116 ms
Sage 1.3.0, place on JevBench v1.6.0's official ranking9
Sage 1.3.0, JevBench v1.6.0 Capability (outside the Jev-class cost cap there)77.1
Quyet minus Clef-Flash on JevBench Capability, rounded (our arithmetic)18
Quyet minus Jev on JevBench Capability (our arithmetic)5.2

JevBench (Benchmark Heaven)

Jev on JevBench's second check of v1.6.0, on new questions77.5

JevBench (Benchmark Heaven)

Sage 1.3.0, Capability on JevBench v1.6.178.6
Sage 1.3.0, place on JevBench v1.6.1's official ranking1
Sage 1.3.0, place among Jev-class systems on v1.6.1's Capability2
Sage 1.3.0, JevBench v1.6.1 questions it answered1,500
Quyet-1.0-Large, Capability on JevBench v1.6.181.7
Quyet-1.0-Large, place on JevBench v1.6.1's official ranking3
Quyet-1.0-Large, place among Jev-class systems on v1.6.1's Capability1
Quyet-1.0-Large, JevBench v1.6.1 questions it answered1,500
Jev, Capability on JevBench v1.6.177.1
Jev, place on JevBench v1.6.1's official ranking2
Jev, place among Jev-class systems on v1.6.1's Capability4
Jev, JevBench v1.6.1 questions it answered1,500
deck-31B, Capability on JevBench v1.6.177.6
deck-31B, place on JevBench v1.6.1's official ranking7
deck-31B, place among Jev-class systems on v1.6.1's Capability3
deck-31B, JevBench v1.6.1 questions it answered1,500
Clef-Flash, Capability on JevBench v1.6.163.8
Clef-Flash, place on JevBench v1.6.1's official ranking21
Clef-Flash, place among Jev-class systems on v1.6.1's Capability12
Clef-Flash, JevBench v1.6.1 questions it answered1,500

Sources

These are the pages the video and this page draw on. We keep a copy of each page as we read it, so a figure can be checked against what the page said at the time.

Credits

The narration in the video is an AI voice, made with ElevenLabs.

Music in the video: "Airport Lounge" by Kevin MacLeod (incompetech.com), licensed under Creative Commons: By Attribution 4.0.

Further reading