model fatıgue
Standing page

Decision models, measured

Last measured 2 Oct 2026Our test: bank-support messages

The short answer

On one task, sorting 3,080 bank-support messages into 77 intents with five labelled examples each, Clef-flash (95.2%) and Clef (95.1%) got about a point more right than Jev (93.9%). GPT-6 Luna and Kev-27B were level with Jev.

On the Decision Index's question, the message alone with no examples, the gap is far wider: Clef got 94.2% right and Jev 80.7%. Among the decision models, only Kev-4B cost less per thousand decisions than Jev ($0.078), and from Berlin Clef-flash (196 ms) and Jev (235 ms) were the fastest at the median.

Kev-27B's confidence matched its accuracy most closely. Every figure is dated, comes from our own runs or the page it names, and links to its entry in the table of every number at the foot.

What this page measures

A decision model takes some text and a question with fixed answers, and returns one answer with a probability for each option instead of writing text. Jev made the idea popular; Clef, Kev and others now answer the same kind of request. This page holds our main measurements of them, one question at a time, and grows as new models ship. Each row says when it was measured.

So far all of it comes from one task: 3,080 messages that customers sent to a bank, from the public BANKING77 test split, each to be sorted into one of 77 intents. Every model gets the same question: the message, five labelled examples of similar messages found by a fixed search over the training set, and 78 options, the intents plus none-of-these. We sent the calls one at a time from a Mac mini on a home connection in Berlin.

One task is a narrow base. A model that leads on bank-support intents can trail on your own labels, so read these numbers as a dated measurement on one job. The method section says how they were made, and the files at the foot have every number on this page.

The models

Each row is the newest pass of that model on our test, dated by when it ran.

ModelMade byHow we called itMeasuredServed as
JevTypeSafeTypeSafe's API, direct2 October 2026, and 30 Septemberjev-1.13.0
ClefCloudflare, open weightsCloudflare Workers AI, entering at its Berlin edge2 October 2026clef
Clef-flashCloudflare, open weightsCloudflare Workers AI, entering at its Berlin edge2 October 2026clef-flash
GPT-6 LunaOpenAI; a general model, asked for one label and a confidenceOpenAI's Responses API, direct, reasoning off30 September 2026gpt-6-luna
Kev-27BJared Palmer, open weightsour own H100 on Modal (eu-north), one request at a time2 October 2026jaredpalmer/kev-27b
Kev-9BJared Palmer, open weightsour own H100 on Modal (eu-north), one request at a time2 October 2026jaredpalmer/kev-9b
Kev-4BJared Palmer, open weightsOpenRouter, which sent it to SiliconFlow30 September 2026jaredpalmer/kev-4b-20260924

Which one gets the most right?

With five examples, Clef-flash and Clef sorted about a point more messages correctly than Jev, and the interval on each difference stays above zero. GPT-6 Luna and Kev-27B were level with Jev within the noise of 3,080 messages.

Kev-9B and Kev-4B fell behind mostly by answering none-of-these on messages that had a real intent. Kev-4B did it 252 times. With that option set aside, its top real intent was right 93.3% of the time, close to Jev, and Kev-9B's was right 93.6% of the time. A caller who offers an "other" option should expect these two to take it more often than the rest.

Share of messages sorted correctly

With five labelled examples. Dot: measured. Line: 95% interval.

measured95% interval
86%88%90%92%94%96%share rightClef-flashClef-flash: 95.2%, interval 94.3 to 95.9 · Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST95.2%ClefClef: 95.1%, interval 94.2 to 95.8 · Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST95.1%GPT-6 LunaGPT-6 Luna: 94.1%, interval 93.2 to 94.9 · Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST94.1%JevJev: 93.9%, interval 93.0 to 94.7 · Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST93.9%Kev-27BKev-27B: 93.8%, interval 92.9 to 94.6 · Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST93.8%Kev-9BKev-9B: 91.3%, interval 90.3 to 92.3 · Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST91.3%Kev-4BKev-4B: 87.7%, interval 86.5 to 88.8 · Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST87.7%
86%88%90%92%94%96%share rightClef-flashClef-flash: 95.2%, interval 94.3 to 95.9 · Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST95.2%ClefClef: 95.1%, interval 94.2 to 95.8 · Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST95.1%GPT-6 LunaGPT-6 Luna: 94.1%, interval 93.2 to 94.9 · Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST94.1%JevJev: 93.9%, interval 93.0 to 94.7 · Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST93.9%Kev-27BKev-27B: 93.8%, interval 92.9 to 94.6 · Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST93.8%Kev-9BKev-9B: 91.3%, interval 90.3 to 92.3 · Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST91.3%Kev-4BKev-4B: 87.7%, interval 86.5 to 88.8 · Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST87.7%
Our runs of 30 September and 2 October, BANKING77 test split, from Berlin.
The numbers in this chart
share right95% interval
Clef-flash95.2%94.3 to 95.9
Clef95.1%94.2 to 95.8
GPT-6 Luna94.1%93.2 to 94.9
Jev93.9%93.0 to 94.7
Kev-27B93.8%92.9 to 94.6
Kev-9B91.3%90.3 to 92.3
Kev-4B87.7%86.5 to 88.8

Accuracy with five labelled examples

Share of the 3,080 messages each model sorted into the right intent, with a 95% interval. The difference from Jev is measured message by message against Jev's pass of the same day.

right95% intervalanswered none-of-theseminus Jev, points [95% interval]measured
Clef-flash95.2%[94.3, 95.9]0+1.2 points [+0.7, +1.8]2 October
Clef95.1%[94.2, 95.8]0+1.1 points [+0.6, +1.7]2 October
GPT-6 Luna94.1%[93.2, 94.9]4+0.2 points [−0.3, +0.7]30 September
Jev93.9%[93.0, 94.7]14the reference2 October
Kev-27B93.8%[92.9, 94.6]21−0.1 points [−0.6, +0.4]2 October
Kev-9B91.3%[90.3, 92.3]110−2.6 points [−3.3, −1.9]2 October
Kev-4B87.7%[86.5, 88.8]252−6.2 points [−7.1, −5.3]30 September
A none-of-these answer counts as wrong: every message has a real intent. Our runs, from a Mac mini in Berlin.

Without examples, and with five

Cloudflare launched Clef with a BANKING77 score of 94.20 against Jev's 79.74. Both use the community Decision Index's test: Jev's figure is the Index's own run, Clef's is Cloudflare's run of it. That test asks a different question from ours: the message alone and 77 options described by the intent names, no examples and no none-of-these, scored by macro-F1. We replicated that request on the same messages and got 94.2 for Clef and 80.0 for Jev, so Cloudflare's number holds for the question it asks.

On our question, with five labelled examples, most of the gap closes: Jev scored 13.2 points higher than on the Index's, Clef 0.8 points. The two questions also differ in layout and in offering none-of-these, so not all of that is the examples; on our layout without examples, Jev scored 78.3% in jev-2. If your setup can't show the model any examples, Clef is far ahead on this task. If it can, the models are within a few points of each other, and price and speed decide.

The Decision Index's question against ours

Share right on the same messages: the message alone (the Decision Index's question), and with five labelled examples (ours).

the Decision Index's questionours, with five examples
75%80%85%90%95%100%share rightdifference, pointsClefClef: the Decision Index's question 94.2% (run 2 Oct 2026, 17:03–21:41 CEST), ours, with five examples 95.1% (run 2 Oct 2026, 15:49–16:57 CEST) · Model Fatigue+0.8Clef-flashClef-flash: the Decision Index's question 90.9% (run 2 Oct 2026, 17:03–21:41 CEST), ours, with five examples 95.2% (run 2 Oct 2026, 15:49–16:57 CEST) · Model Fatigue+4.3Kev-27BKev-27B: the Decision Index's question 86.8% (run 2 Oct 2026, 17:03–21:41 CEST), ours, with five examples 93.8% (run 2 Oct 2026, 21:05–22:01 CEST) · Model Fatigue+7.0Kev-9BKev-9B: the Decision Index's question 83.3% (run 2 Oct 2026, 17:03–21:41 CEST), ours, with five examples 91.3% (run 2 Oct 2026, 21:05–22:01 CEST) · Model Fatigue+8.0JevJev: the Decision Index's question 80.7% (run 2 Oct 2026, 17:03–21:41 CEST), ours, with five examples 93.9% (run 2 Oct 2026, 15:49–16:57 CEST) · Model Fatigue+13.2
75%80%85%90%95%100%share rightdifference, pointsClefClef: the Decision Index's question 94.2% (run 2 Oct 2026, 17:03–21:41 CEST), ours, with five examples 95.1% (run 2 Oct 2026, 15:49–16:57 CEST) · Model Fatigue+0.8Clef-flashClef-flash: the Decision Index's question 90.9% (run 2 Oct 2026, 17:03–21:41 CEST), ours, with five examples 95.2% (run 2 Oct 2026, 15:49–16:57 CEST) · Model Fatigue+4.3Kev-27BKev-27B: the Decision Index's question 86.8% (run 2 Oct 2026, 17:03–21:41 CEST), ours, with five examples 93.8% (run 2 Oct 2026, 21:05–22:01 CEST) · Model Fatigue+7.0Kev-9BKev-9B: the Decision Index's question 83.3% (run 2 Oct 2026, 17:03–21:41 CEST), ours, with five examples 91.3% (run 2 Oct 2026, 21:05–22:01 CEST) · Model Fatigue+8.0JevJev: the Decision Index's question 80.7% (run 2 Oct 2026, 17:03–21:41 CEST), ours, with five examples 93.9% (run 2 Oct 2026, 15:49–16:57 CEST) · Model Fatigue+13.2
Our runs of 2 October. The questions also differ in layout and in offering none-of-these, so the difference is not the examples alone.
The numbers in this chart
the Decision Index's questionours, with five examplesdifference, points
Clef94.2%95.1%+0.8 points
Clef-flash90.9%95.2%+4.3 points
Kev-27B86.8%93.8%+7.0 points
Kev-9B83.3%91.3%+8.0 points
Jev80.7%93.9%+13.2 points

The Decision Index's question and ours

The Decision Index's question gives the message alone and 77 numbered options described by the intent names, with no none-of-these; it is scored by macro-F1 and here also by share right. Ours adds five labelled examples and offers the 78 named options. Same messages.

BANKING77 macro-F1, publishedmacro-F1 on that question, measured [95% interval]right on that questionright on oursdifference, points
Clef94.20, Cloudflare's94.2 [93.3, 94.9]94.2%95.1%+0.8 points
Clef-flash90.93, Cloudflare's90.8 [89.8, 91.7]90.9%95.2%+4.3 points
Kev-27Bnot on the board86.4 [85.1, 87.4]86.8%93.8%+7.0 points
Kev-9B84.83, the board's run of the earlier version83.0 [81.6, 84.1]83.3%91.3%+8.0 points
Jev79.74, the board's run, in Cloudflare's table80.0 [78.6, 81.1]80.7%93.9%+13.2 points
Published figures from Cloudflare's launch post and the Decision Index's own data file; Clef's and Clef-flash's are Cloudflare's own runs of the Index's test. The two questions differ in layout and in the none-of-these option as well as in the examples: on our layout without examples, Jev scored 78.3% in jev-2. GPT-6 Luna and Kev-4B were not run on the Index's question.

How fast is it from Berlin?

For a call this short, much of the time is the network, so every time here is from Berlin and comes with the network floor to the same host. Jev's median was 235 ms, and a bare request to TypeSafe's host took 171 ms at the start of that run. Clef-flash had the fastest median, 196 ms, but a longer tail than Jev; its slowest calls bunched in the first part of its pass. Clef took 562 ms at the median and GPT-6 Luna 1,097 ms.

Kev-27B and Kev-9B ran on H100s we rented on Modal. The model itself took 110 ms and 35 ms per request by its server's count, and the rest of each call was the path from Berlin to Modal's machine. These are times of our deployment, which we didn't tune, not of a service you can call.

Time per request, from Berlin

The bar runs to the median; the ticks mark p95 and p99. One request at a time.

median95th and 99th percentile
0 s0.7 s1.4 s2.1 s2.8 s3.5 sseconds per requestClef-flashClef-flash: median 196 ms, 95th percentile 1,014 ms, 99th 2,172 ms · Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST196 msJevJev: median 235 ms, 95th percentile 335 ms, 99th 442 ms · Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST235 msKev-9BKev-9B: median 437 ms, 95th percentile 548 ms, 99th 665 ms · Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST437 msKev-27BKev-27B: median 529 ms, 95th percentile 551 ms, 99th 598 ms · Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST529 msKev-4BKev-4B: median 540 ms, 95th percentile 883 ms, 99th 1,372 ms · Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST540 msClefClef: median 562 ms, 95th percentile 1,147 ms, 99th 2,001 ms · Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST562 msGPT-6 LunaGPT-6 Luna: median 1,097 ms, 95th percentile 2,235 ms, 99th 3,349 ms · Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST1,097 ms
0 s0.7 s1.4 s2.1 s2.8 s3.5 sseconds per requestClef-flashClef-flash: median 196 ms, 95th percentile 1,014 ms, 99th 2,172 ms · Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST196 msJevJev: median 235 ms, 95th percentile 335 ms, 99th 442 ms · Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST235 msKev-9BKev-9B: median 437 ms, 95th percentile 548 ms, 99th 665 ms · Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST437 msKev-27BKev-27B: median 529 ms, 95th percentile 551 ms, 99th 598 ms · Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST529 msKev-4BKev-4B: median 540 ms, 95th percentile 883 ms, 99th 1,372 ms · Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST540 msClefClef: median 562 ms, 95th percentile 1,147 ms, 99th 2,001 ms · Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST562 msGPT-6 LunaGPT-6 Luna: median 1,097 ms, 95th percentile 2,235 ms, 99th 3,349 ms · Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST1,097 ms
Our runs of 30 September and 2 October, from a Mac mini on a Berlin home line. Kev-27B and Kev-9B ran on our own GPUs, so theirs are times of our deployment, not of a public service.
The numbers in this chart

Time per request from Berlin

Wall time of each call, request sent to answer parsed, one call at a time over a kept-alive connection from a Mac mini on a Berlin home line. p95 is the time that one call in twenty took longer than; p99, one in a hundred.

medianp95p99network floorwhat the floor is
Clef-flash196 ms1,014 ms2,172 ms204 msCloudflare's API host, which is not the path model calls take
Jev235 ms335 ms442 ms171 msTypeSafe's API host
Kev-9B437 ms548 ms665 ms132 msModal's edge; the model itself took 35 ms, by its server's count
Kev-27B529 ms551 ms598 ms132 msModal's edge; the model itself took 110 ms, by its server's count
Kev-4B540 ms883 ms1,372 ms21 msOpenRouter's edge only; its hop to SiliconFlow is inside every call
Clef562 ms1,147 ms2,001 ms204 msCloudflare's API host, which is not the path model calls take
GPT-6 Luna1,097 ms2,235 ms3,349 ms170 msOpenAI's API host
The network floor is a bare request to the same host on a kept-alive connection, the median of twenty, taken as each run started. The five examples were found in advance by jev-2's fixed search (10 ms median per message there) and are not in these times.

What a thousand decisions cost

Among the hosted models, Kev-4B through OpenRouter was the cheapest at $0.041 per thousand decisions, then Jev at $0.078. Clef-flash cost $0.180 and Clef $0.480.

Price per token is a poor guide here, because each provider counts the same request differently. Jev counted 1,846 input tokens per request and Clef 1,999, while Kev's server counted 975 for the same body. GPT-6 Luna counted 1,066, and 3,032 of its 3,080 calls also paid OpenAI's charge for writing the prompt to its cache, which no call read from.

Running Kev on our own GPU one request at a time cost $0.730 per thousand for Kev-27B. On the shorter requests without examples, sent six at a time, it cost $0.157, still above Jev. At our volume, running the weights yourself buys control of the model and the data path, not a lower price.

Dollars per thousand decisions

Our test's requests, as each provider counts and charges them.

$0$0.2$0.4$0.6$0.8dollars per thousand decisionsKev-4BKev-4B: $0.041 · Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST$0.041JevJev: $0.078 · Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST$0.078GPT-6 LunaGPT-6 Luna: $0.143 · Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST$0.143Clef-flashClef-flash: $0.180 · Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST$0.180ClefClef: $0.480 · Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST$0.480Kev-9BKev-9B: $0.621 · Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST$0.621Kev-27BKev-27B: $0.730 · Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST$0.730
$0$0.2$0.4$0.6$0.8dollars per thousand decisionsKev-4BKev-4B: $0.041 · Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST$0.041JevJev: $0.078 · Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST$0.078GPT-6 LunaGPT-6 Luna: $0.143 · Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST$0.143Clef-flashClef-flash: $0.180 · Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST$0.180ClefClef: $0.480 · Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST$0.480Kev-9BKev-9B: $0.621 · Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST$0.621Kev-27BKev-27B: $0.730 · Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST$0.730
Outlined: our own GPU time on Modal with one request at a time, the dearest way to run a GPU.
The numbers in this chart
dollars per thousand decisions
Kev-4B$0.041
Jev$0.078
GPT-6 Luna$0.143
Clef-flash$0.180
Clef$0.480
Kev-9B$0.621
Kev-27B$0.730

Cost per thousand decisions

Dollars for a thousand messages of our test, each sent with its five examples and the 78 options.

per thousand decisionsinput tokens counted per requesthow it was counted
Kev-4B$0.041975as OpenRouter billed each call
Jev$0.0781,846TypeSafe's rate card × input tokens counted; output is free
GPT-6 Luna$0.1431,066OpenAI's rate card × tokens, with a cache-write charge on 3,032 of the 3,080 calls
Clef-flash$0.1801,999Cloudflare's rate card × input tokens, which matches the neurons it bills
Clef$0.4801,999Cloudflare's rate card × input tokens, which matches the neurons it bills
Kev-9B$0.621975Modal's bill for our H100 time, one request at a time; $0.112 on the shorter requests without examples, six at a time
Kev-27B$0.730975Modal's bill for our H100 time, one request at a time; $0.157 on the shorter requests without examples, six at a time
The request is the same for every model that takes Jev's request format; GPT-6 Luna's puts the options in its instructions instead. Free daily allowances are not subtracted. The Kev figures share Modal's bill out by each pass's running time; cold starts and idle time, which would add to them, are left out.

How far can you trust the confidence?

Each of these models gives a confidence with its answer, and the practical use is to accept the confident answers automatically and send the rest to a person. That works if the confidence tracks how often the model is right.

Kev-27B's did best: across the range, what it stated was close to how often it was right, with a calibration error of 0.014. Among the models that return probabilities, Clef (0.030) and Jev (0.035) came next. Clef-flash and Kev-4B state numbers well below how often they are right, so a threshold for them has to be set on your own data rather than read off the number. After a simple fit on held-out messages every model's calibration error was 0.021 or lower, and Clef-flash's fell from 0.143 to 0.006.

GPT-6 Luna doesn't return probabilities on this setup. Its confidence is a number it writes in its answer, and it is reasonably well placed, but it is a different kind of number from the others'.

Stated confidence against the share right

Each model's answers grouped by the confidence it gave. On the dashed line a model is right exactly as often as it says.

where stated confidence equals the share righta tenth of the range, sized by how many answers fell in it
Kev-27BECE 0.0140.20.20.60.61.01.0Kev-27B: 9 answers stated about 0.25, 0.22 of them right · Model Fatigue, run 2 Oct 2026, 21:05–22:01 CESTKev-27B: 25 answers stated about 0.37, 0.36 of them right · Model Fatigue, run 2 Oct 2026, 21:05–22:01 CESTKev-27B: 51 answers stated about 0.45, 0.43 of them right · Model Fatigue, run 2 Oct 2026, 21:05–22:01 CESTKev-27B: 55 answers stated about 0.55, 0.45 of them right · Model Fatigue, run 2 Oct 2026, 21:05–22:01 CESTKev-27B: 62 answers stated about 0.65, 0.56 of them right · Model Fatigue, run 2 Oct 2026, 21:05–22:01 CESTKev-27B: 75 answers stated about 0.75, 0.77 of them right · Model Fatigue, run 2 Oct 2026, 21:05–22:01 CESTKev-27B: 204 answers stated about 0.86, 0.86 of them right · Model Fatigue, run 2 Oct 2026, 21:05–22:01 CESTKev-27B: 2,598 answers stated about 0.98, 0.99 of them right · Model Fatigue, run 2 Oct 2026, 21:05–22:01 CESTGPT-6 LunaECE 0.0250.20.20.60.61.01.0GPT-6 Luna: 2 answers stated about 0.38, 0.50 of them right · Model Fatigue, run 30 Sep 2026, 21:42–23:38 CESTGPT-6 Luna: 9 answers stated about 0.46, 0.33 of them right · Model Fatigue, run 30 Sep 2026, 21:42–23:38 CESTGPT-6 Luna: 22 answers stated about 0.56, 0.32 of them right · Model Fatigue, run 30 Sep 2026, 21:42–23:38 CESTGPT-6 Luna: 29 answers stated about 0.65, 0.59 of them right · Model Fatigue, run 30 Sep 2026, 21:42–23:38 CESTGPT-6 Luna: 69 answers stated about 0.75, 0.59 of them right · Model Fatigue, run 30 Sep 2026, 21:42–23:38 CESTGPT-6 Luna: 138 answers stated about 0.86, 0.74 of them right · Model Fatigue, run 30 Sep 2026, 21:42–23:38 CESTGPT-6 Luna: 2,811 answers stated about 0.98, 0.97 of them right · Model Fatigue, run 30 Sep 2026, 21:42–23:38 CESTClefECE 0.0300.20.20.60.61.01.0Clef: 5 answers stated about 0.28, 0.40 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:57 CESTClef: 13 answers stated about 0.36, 0.69 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:57 CESTClef: 23 answers stated about 0.45, 0.57 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:57 CESTClef: 38 answers stated about 0.55, 0.50 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:57 CESTClef: 55 answers stated about 0.65, 0.56 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:57 CESTClef: 69 answers stated about 0.76, 0.70 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:57 CESTClef: 187 answers stated about 0.86, 0.83 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:57 CESTClef: 2,690 answers stated about 0.96, 0.99 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:57 CESTJevECE 0.0350.20.20.60.61.01.0Jev: 1 answers stated about 0.28, 0.00 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:57 CESTJev: 3 answers stated about 0.35, 0.67 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:57 CESTJev: 17 answers stated about 0.45, 0.35 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:57 CESTJev: 37 answers stated about 0.55, 0.49 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:57 CESTJev: 39 answers stated about 0.64, 0.44 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:57 CESTJev: 46 answers stated about 0.75, 0.63 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:57 CESTJev: 81 answers stated about 0.85, 0.72 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:57 CESTJev: 2,856 answers stated about 1.00, 0.97 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:57 CESTKev-9BECE 0.0380.20.20.60.61.01.0Kev-9B: 2 answers stated about 0.29, 0.50 of them right · Model Fatigue, run 2 Oct 2026, 21:05–22:01 CESTKev-9B: 21 answers stated about 0.36, 0.24 of them right · Model Fatigue, run 2 Oct 2026, 21:05–22:01 CESTKev-9B: 88 answers stated about 0.46, 0.38 of them right · Model Fatigue, run 2 Oct 2026, 21:05–22:01 CESTKev-9B: 102 answers stated about 0.55, 0.47 of them right · Model Fatigue, run 2 Oct 2026, 21:05–22:01 CESTKev-9B: 115 answers stated about 0.65, 0.64 of them right · Model Fatigue, run 2 Oct 2026, 21:05–22:01 CESTKev-9B: 180 answers stated about 0.75, 0.81 of them right · Model Fatigue, run 2 Oct 2026, 21:05–22:01 CESTKev-9B: 440 answers stated about 0.86, 0.92 of them right · Model Fatigue, run 2 Oct 2026, 21:05–22:01 CESTKev-9B: 2,132 answers stated about 0.96, 0.99 of them right · Model Fatigue, run 2 Oct 2026, 21:05–22:01 CESTKev-4BECE 0.1130.20.20.60.61.01.0Kev-4B: 16 answers stated about 0.27, 0.19 of them right · Model Fatigue, run 30 Sep 2026, 21:42–23:38 CESTKev-4B: 147 answers stated about 0.36, 0.29 of them right · Model Fatigue, run 30 Sep 2026, 21:42–23:38 CESTKev-4B: 264 answers stated about 0.45, 0.48 of them right · Model Fatigue, run 30 Sep 2026, 21:42–23:38 CESTKev-4B: 201 answers stated about 0.55, 0.71 of them right · Model Fatigue, run 30 Sep 2026, 21:42–23:38 CESTKev-4B: 237 answers stated about 0.65, 0.88 of them right · Model Fatigue, run 30 Sep 2026, 21:42–23:38 CESTKev-4B: 369 answers stated about 0.75, 0.95 of them right · Model Fatigue, run 30 Sep 2026, 21:42–23:38 CESTKev-4B: 797 answers stated about 0.86, 0.98 of them right · Model Fatigue, run 30 Sep 2026, 21:42–23:38 CESTKev-4B: 1,049 answers stated about 0.93, 1.00 of them right · Model Fatigue, run 30 Sep 2026, 21:42–23:38 CESTClef-flashECE 0.1430.20.20.60.61.01.0Clef-flash: 7 answers stated about 0.26, 0.14 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:57 CESTClef-flash: 12 answers stated about 0.35, 0.33 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:57 CESTClef-flash: 59 answers stated about 0.46, 0.68 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:57 CESTClef-flash: 87 answers stated about 0.56, 0.78 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:57 CESTClef-flash: 241 answers stated about 0.66, 0.87 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:57 CESTClef-flash: 736 answers stated about 0.76, 0.95 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:57 CESTClef-flash: 1,339 answers stated about 0.85, 0.98 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:57 CESTClef-flash: 599 answers stated about 0.92, 1.00 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Kev-27BECE 0.0140.20.20.60.61.01.0Kev-27B: 9 answers stated about 0.25, 0.22 of them right · Model Fatigue, run 2 Oct 2026, 21:05–22:01 CESTKev-27B: 25 answers stated about 0.37, 0.36 of them right · Model Fatigue, run 2 Oct 2026, 21:05–22:01 CESTKev-27B: 51 answers stated about 0.45, 0.43 of them right · Model Fatigue, run 2 Oct 2026, 21:05–22:01 CESTKev-27B: 55 answers stated about 0.55, 0.45 of them right · Model Fatigue, run 2 Oct 2026, 21:05–22:01 CESTKev-27B: 62 answers stated about 0.65, 0.56 of them right · Model Fatigue, run 2 Oct 2026, 21:05–22:01 CESTKev-27B: 75 answers stated about 0.75, 0.77 of them right · Model Fatigue, run 2 Oct 2026, 21:05–22:01 CESTKev-27B: 204 answers stated about 0.86, 0.86 of them right · Model Fatigue, run 2 Oct 2026, 21:05–22:01 CESTKev-27B: 2,598 answers stated about 0.98, 0.99 of them right · Model Fatigue, run 2 Oct 2026, 21:05–22:01 CESTGPT-6 LunaECE 0.0250.20.20.60.61.01.0GPT-6 Luna: 2 answers stated about 0.38, 0.50 of them right · Model Fatigue, run 30 Sep 2026, 21:42–23:38 CESTGPT-6 Luna: 9 answers stated about 0.46, 0.33 of them right · Model Fatigue, run 30 Sep 2026, 21:42–23:38 CESTGPT-6 Luna: 22 answers stated about 0.56, 0.32 of them right · Model Fatigue, run 30 Sep 2026, 21:42–23:38 CESTGPT-6 Luna: 29 answers stated about 0.65, 0.59 of them right · Model Fatigue, run 30 Sep 2026, 21:42–23:38 CESTGPT-6 Luna: 69 answers stated about 0.75, 0.59 of them right · Model Fatigue, run 30 Sep 2026, 21:42–23:38 CESTGPT-6 Luna: 138 answers stated about 0.86, 0.74 of them right · Model Fatigue, run 30 Sep 2026, 21:42–23:38 CESTGPT-6 Luna: 2,811 answers stated about 0.98, 0.97 of them right · Model Fatigue, run 30 Sep 2026, 21:42–23:38 CESTClefECE 0.0300.20.20.60.61.01.0Clef: 5 answers stated about 0.28, 0.40 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:57 CESTClef: 13 answers stated about 0.36, 0.69 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:57 CESTClef: 23 answers stated about 0.45, 0.57 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:57 CESTClef: 38 answers stated about 0.55, 0.50 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:57 CESTClef: 55 answers stated about 0.65, 0.56 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:57 CESTClef: 69 answers stated about 0.76, 0.70 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:57 CESTClef: 187 answers stated about 0.86, 0.83 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:57 CESTClef: 2,690 answers stated about 0.96, 0.99 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:57 CESTJevECE 0.0350.20.20.60.61.01.0Jev: 1 answers stated about 0.28, 0.00 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:57 CESTJev: 3 answers stated about 0.35, 0.67 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:57 CESTJev: 17 answers stated about 0.45, 0.35 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:57 CESTJev: 37 answers stated about 0.55, 0.49 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:57 CESTJev: 39 answers stated about 0.64, 0.44 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:57 CESTJev: 46 answers stated about 0.75, 0.63 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:57 CESTJev: 81 answers stated about 0.85, 0.72 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:57 CESTJev: 2,856 answers stated about 1.00, 0.97 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:57 CESTKev-9BECE 0.0380.20.20.60.61.01.0Kev-9B: 2 answers stated about 0.29, 0.50 of them right · Model Fatigue, run 2 Oct 2026, 21:05–22:01 CESTKev-9B: 21 answers stated about 0.36, 0.24 of them right · Model Fatigue, run 2 Oct 2026, 21:05–22:01 CESTKev-9B: 88 answers stated about 0.46, 0.38 of them right · Model Fatigue, run 2 Oct 2026, 21:05–22:01 CESTKev-9B: 102 answers stated about 0.55, 0.47 of them right · Model Fatigue, run 2 Oct 2026, 21:05–22:01 CESTKev-9B: 115 answers stated about 0.65, 0.64 of them right · Model Fatigue, run 2 Oct 2026, 21:05–22:01 CESTKev-9B: 180 answers stated about 0.75, 0.81 of them right · Model Fatigue, run 2 Oct 2026, 21:05–22:01 CESTKev-9B: 440 answers stated about 0.86, 0.92 of them right · Model Fatigue, run 2 Oct 2026, 21:05–22:01 CESTKev-9B: 2,132 answers stated about 0.96, 0.99 of them right · Model Fatigue, run 2 Oct 2026, 21:05–22:01 CESTKev-4BECE 0.1130.20.20.60.61.01.0Kev-4B: 16 answers stated about 0.27, 0.19 of them right · Model Fatigue, run 30 Sep 2026, 21:42–23:38 CESTKev-4B: 147 answers stated about 0.36, 0.29 of them right · Model Fatigue, run 30 Sep 2026, 21:42–23:38 CESTKev-4B: 264 answers stated about 0.45, 0.48 of them right · Model Fatigue, run 30 Sep 2026, 21:42–23:38 CESTKev-4B: 201 answers stated about 0.55, 0.71 of them right · Model Fatigue, run 30 Sep 2026, 21:42–23:38 CESTKev-4B: 237 answers stated about 0.65, 0.88 of them right · Model Fatigue, run 30 Sep 2026, 21:42–23:38 CESTKev-4B: 369 answers stated about 0.75, 0.95 of them right · Model Fatigue, run 30 Sep 2026, 21:42–23:38 CESTKev-4B: 797 answers stated about 0.86, 0.98 of them right · Model Fatigue, run 30 Sep 2026, 21:42–23:38 CESTKev-4B: 1,049 answers stated about 0.93, 1.00 of them right · Model Fatigue, run 30 Sep 2026, 21:42–23:38 CESTClef-flashECE 0.1430.20.20.60.61.01.0Clef-flash: 7 answers stated about 0.26, 0.14 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:57 CESTClef-flash: 12 answers stated about 0.35, 0.33 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:57 CESTClef-flash: 59 answers stated about 0.46, 0.68 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:57 CESTClef-flash: 87 answers stated about 0.56, 0.78 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:57 CESTClef-flash: 241 answers stated about 0.66, 0.87 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:57 CESTClef-flash: 736 answers stated about 0.76, 0.95 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:57 CESTClef-flash: 1,339 answers stated about 0.85, 0.98 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:57 CESTClef-flash: 599 answers stated about 0.92, 1.00 of them right · Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Our runs of 30 September and 2 October; the one answer stated below a fifth, a Kev-27B answer, is left off. GPT-6 Luna's confidence is a number it writes in its answer, not a probability the model computes, so its panel shows a different kind of number.
The numbers in this chart
statedanswersmean statedshare right
Kev-27B0.1 to 0.210.191.00
Kev-27B0.2 to 0.390.250.22
Kev-27B0.3 to 0.4250.370.36
Kev-27B0.4 to 0.5510.450.43
Kev-27B0.5 to 0.6550.550.45
Kev-27B0.6 to 0.7620.650.56
Kev-27B0.7 to 0.8750.750.77
Kev-27B0.8 to 0.92040.860.86
Kev-27B0.9 to 1.02,5980.980.99
GPT-6 Luna0.3 to 0.420.380.50
GPT-6 Luna0.4 to 0.590.460.33
GPT-6 Luna0.5 to 0.6220.560.32
GPT-6 Luna0.6 to 0.7290.650.59
GPT-6 Luna0.7 to 0.8690.750.59
GPT-6 Luna0.8 to 0.91380.860.74
GPT-6 Luna0.9 to 1.02,8110.980.97
Clef0.2 to 0.350.280.40
Clef0.3 to 0.4130.360.69
Clef0.4 to 0.5230.450.57
Clef0.5 to 0.6380.550.50
Clef0.6 to 0.7550.650.56
Clef0.7 to 0.8690.760.70
Clef0.8 to 0.91870.860.83
Clef0.9 to 1.02,6900.960.99
Jev0.2 to 0.310.280.00
Jev0.3 to 0.430.350.67
Jev0.4 to 0.5170.450.35
Jev0.5 to 0.6370.550.49
Jev0.6 to 0.7390.640.44
Jev0.7 to 0.8460.750.63
Jev0.8 to 0.9810.850.72
Jev0.9 to 1.02,8561.000.97
Kev-9B0.2 to 0.320.290.50
Kev-9B0.3 to 0.4210.360.24
Kev-9B0.4 to 0.5880.460.38
Kev-9B0.5 to 0.61020.550.47
Kev-9B0.6 to 0.71150.650.64
Kev-9B0.7 to 0.81800.750.81
Kev-9B0.8 to 0.94400.860.92
Kev-9B0.9 to 1.02,1320.960.99
Kev-4B0.2 to 0.3160.270.19
Kev-4B0.3 to 0.41470.360.29
Kev-4B0.4 to 0.52640.450.48
Kev-4B0.5 to 0.62010.550.71
Kev-4B0.6 to 0.72370.650.88
Kev-4B0.7 to 0.83690.750.95
Kev-4B0.8 to 0.97970.860.98
Kev-4B0.9 to 1.01,0490.931.00
Clef-flash0.2 to 0.370.260.14
Clef-flash0.3 to 0.4120.350.33
Clef-flash0.4 to 0.5590.460.68
Clef-flash0.5 to 0.6870.560.78
Clef-flash0.6 to 0.72410.660.87
Clef-flash0.7 to 0.87360.760.95
Clef-flash0.8 to 0.91,3390.850.98
Clef-flash0.9 to 1.05990.921.00

How well the confidence matches the accuracy

Calibration error (ECE): the average gap between the confidence a model states and the share of those answers that are right, over ten equal bands. Lower is better. AUROC: how well the confidence ranks right answers above wrong ones; one is perfect, a half is a coin flip.

confidence usedcalibration error as returnedafter a fit on held-out messagesAUROC
Kev-27Bthe probability of the answer it chose0.0140.0090.920
GPT-6 Lunathe number it writes in its answer0.0250.0210.876
Clefthe probability of the answer it chose0.0300.0150.900
Jevthe probability of the answer it chose0.0350.0100.840
Kev-9Bthe probability of the answer it chose0.0380.0120.906
Kev-4Bthe probability of the answer it chose0.1130.0090.922
Clef-flashthe probability of the answer it chose0.1430.0060.819
The fit is a two-parameter rescaling (Platt scaling) learned on the 573 dev messages and applied unchanged to the test.

The workload table puts this in a support team's terms. At a budget of 2% wrong among the answers accepted automatically, Clef handled 93.2% of messages alone and Jev 84.4%. Jev's threshold had to sit at 0.990: most of its answers carry a confidence that high, which leaves a threshold little room to separate the right ones from the wrong. GPT-6 Luna's threshold, fixed on the dev messages, let through 2.99% wrong answers on the test, above the budget.

What each model handles alone, within an error budget

The policy: accept the model's answer when its confidence reaches a threshold, send the message to a person otherwise. The threshold is the lowest that kept errors among accepted answers under 2% on the dev messages, fixed before the test.

thresholdaccepted automaticallywrong among the accepted [95% interval]sent to a person, per thousand
Clef0.80393.2%2.40% [1.90, 3.03]68
GPT-6 Luna0.90091.3%2.99% [2.42, 3.68]87
Kev-27B0.79391.0%2.28% [1.79, 2.90]90
Jev0.99084.4%2.04% [1.56, 2.66]156
Kev-9B0.81382.7%2.47% [1.94, 3.15]173
Clef-flash0.73281.7%2.15% [1.65, 2.79]183
Kev-4B0.66974.3%1.75% [1.29, 2.37]257
Six of the seven thresholds land over the budget on the test, GPT-6 Luna's by the most: a dev slice of this size places a threshold coarsely. We report that as it fell and did not refit.

Same answer twice, and over time

Asked the same messages again the same day, Kev-27B, Kev-9B, Clef, Clef-flash and Kev-4B gave identical answers every time. Jev gave the same answer to 199 of 200 and GPT-6 Luna to 197.

Over days, Jev kept the same version string but changed a few answers: 3,075 of 3,080 were the same between 21 September and 30 September, and 3,073 between 30 September and 2 October. Its median time also fell over that week.

The same messages again

A second pass over messages from the test on the same day, answers compared with the first pass.

same answer
Kev-27B200 of 200
Kev-9B200 of 200
Clef200 of 200
Clef-flash200 of 200
Kev-4B200 of 200
Jev (30 September pass)199 of 200
GPT-6 Luna197 of 200

Jev and GPT-6 Luna over time

The same messages, sent again days later.

right, thenright, latersame answermedian time, thenmedian time, later
Jev, 21 September to 30 September94.0%93.9%3,075 of 3,080313 ms253 ms
Jev, 30 September to 2 October93.9%93.9%3,073 of 3,080253 ms235 ms
GPT-6 Luna, 24 September to 30 September94.0%94.1%3,047 of 3,0801,196 ms1,097 ms
Jev reported the same version, jev-1.13.0, on all three days. GPT-6 Luna went through OpenRouter on 24 September and to OpenAI directly on 30 September.

Has it seen the test before?

BANKING77 is public, so a model may have trained on it. Kev's model cards list it among their training data. Cloudflare post-trained Clef on an open Qwen3.8 model with its own synthetic data, and doesn't say whether BANKING77 was in that data or in the base model's. We found nothing from TypeSafe on whether Jev saw it.

We asked five of the models, without examples, about 200 messages from the training split and about paraphrases of the same messages; M37's jev-2 asked Jev and GPT-6 Luna the same. A model that had memorised the originals should do much worse on the paraphrases. Every model dropped, but by less than a simple vote of the nearest labelled examples, which can't memorise anything and still drops 15.0 points. So this probe can't tell a little memorisation from the cost of paraphrasing. For Kev it measures something narrower still, since BANKING77 is in its training data.

Training messages against paraphrases of them

Share right without examples on 200 messages from BANKING77's training split, and on paraphrases of the same messages. A model that had memorised the originals should drop by more than the yardstick does.

originalsparaphrasesdrop, points [95% interval]BANKING77 in its training data?
Kev-27B79.5%65.5%+14.0 points [+9.0, +19.0]yes, its model card lists BANKING77
Clef-flash95.5%83.5%+12.0 points [+7.5, +16.5]not stated; Cloudflare post-trained it on its own synthetic data
Kev-9B74.0%62.5%+11.5 points [+6.5, +17.0]yes, its model card lists BANKING77
Kev-4B63.5%55.5%+8.0 points [+2.5, +14.0]not checked for this version
Clef94.5%87.0%+7.5 points [+4.0, +11.5]not stated; Cloudflare post-trained it on its own synthetic data
GPT-6 Luna (jev-2, 24 September, its own prompt)76.5%71.5%+5.0 points [+1.0, +9.5]not checked
Jev (jev-2, 21 September)72.5%68.0%+4.5 points [+0.5, +8.5]not stated by TypeSafe, as far as we found
The five nearest labelled examples, voting; it cannot memorise anything94.5%79.5%+15.0 points [+10.0, +20.5]the yardstick
Our test itself uses the held-out test split, not these training messages. Jev's and GPT-6 Luna's rows and the yardstick are from M37's jev-2 run of the same probe.

Would a plain classifier do?

Could a trained classifier do the same job? On this task, with nearly all of the training split, it nearly does. A small embedding model with logistic regression on top, trained on 9,405 labelled messages, got 93.1% right in 9 ms per message on a Mac. Jev got 93.9%.

That comparison favours the classifier, which saw thousands of labelled messages. The live question is how few labelled examples a trained classifier needs before it catches up with a decision model, and that is the next thing we measure.

Plain baselines on the same test

From M37's jev-2 run of the same protocol on the same messages, run between 21 September and 24 September, not re-run since.

right95% intervalmedian timeper thousand decisions
A trained classifier: a small embedding model with logistic regression on top93.1%[92.2, 94.0]9 ms, on the Macno API cost
The label of the single nearest labelled example92.6%[91.6, 93.5]10 ms, on the Macno API cost
Ling 3.0 Flash, an LLM, through OpenRouter93.8%[92.9, 94.6]723 ms$0.020
DeepSeek V4.1 Flash, through OpenRouter93.6%[92.7, 94.4]953 ms$0.038
Jev, for comparison (2 October)93.9%[93.0, 94.7]235 ms$0.078
The trained classifier learned from 9,405 labelled training messages, the same pool the five examples are drawn from: the training split less the dev slice and the few messages that also appear in the test.

What we measure next

  • How many labelled examples a plain classifier needs to beat Jev and Clef: the same messages, with the classifier trained on a few, then more, examples per intent.
  • OpenAI's Decisions API, on the day it opens to us. The code that calls it is already written.
  • Latency from other regions, since a Berlin number says little about a caller in Virginia or Singapore.
  • Calibration on other kinds of task, including yes-or-no questions.
  • Accuracy as the list of options grows.

Each will get its own table here, dated, when it has been measured.

How it was measured

The protocol was written down and committed before any test message was sent to a model, and every model since has run it unchanged. One exception is disclosed in the Kev addendum: before Kev's settings were frozen, 40 test messages went to each Kev model to time the round trip; only the timing was read, and those messages were sent again in the scored pass. Thresholds and the recalibration fits were set on 573 held-out training messages, never on the test. Every model call is kept, with a hash of its request, the full response and its timing.

Accuracy intervals are Wilson intervals. Differences between two models are paired, message by message, with a bootstrap over the messages. The calibration error uses ten equal bands of stated confidence. Costs are the provider's bill where it gives one per call, as OpenRouter does, and the rate card times the tokens it reports where it doesn't. The Kev rows are Modal's bill for the GPU time of each pass.

What this doesn't cover yet: other tasks, other places, many requests at once on the hosted models, images, and long documents. The protocol comes from M37's jev-2 test, which has a video of its own. The per-message file at the foot of this page has every model's answer, confidence and time for each message, so the accuracy, the differences between models and the calibration as returned can all be recomputed from it.

Reading this page

One task, one place, one main pass per model (Jev has two, on 30 September and 2 October). Each row carries the date its model ran, and the date at the top is the newest. When a new decision model ships, it gets a row in each table; a new question gets a table of its own.

Corrections

None so far. Every change to a figure already on this page is listed here and on the corrections page; a new model or a new question is added, not corrected.

Every number

These are all 474 figures on this page, grouped by the question they answer, each with the run or page it came from and when. Our own figures are read from the result files of each run; nobody typed them in.

Setup

7 figures
Test messages every model answered (BANKING77 test split)3,080Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
Options in every question: the 77 intents and none-of-these78Model Fatigue, frozen 30 Sep 2026
BANKING77 intents in the test split77Model Fatigue, frozen 30 Sep 2026
Training messages held out as the dev slice, where every threshold and recalibration was fitted573Model Fatigue, frozen 30 Sep 2026
Error budget for the workload figures: share of wrong answers allowed among the accepted ones2%Model Fatigue, frozen 30 Sep 2026
Median time to find the five labelled examples for a message, locally (not in any model's latency)10 msModel Fatigue (M37's jev-2 run), run 21–24 Sep 2026
Test messages sent to each Kev model as a latency probe before its settings were frozen (timing read, answers not scored)40Model Fatigue, frozen 2 Oct 2026

The models

8 figures
Jev, model as servedjev-1.13.0Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Jev (30 September pass), model as servedjev-1.13.0Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
GPT-6 Luna, model as servedgpt-6-lunaModel Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
Kev-4B, model as servedjaredpalmer/kev-4b-20260924Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
Clef, model as servedclefModel Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef-flash, model as servedclef-flashModel Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Kev-27B, model as servedjaredpalmer/kev-27bModel Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-9B, model as servedjaredpalmer/kev-9bModel Fatigue, run 2 Oct 2026, 21:05–22:01 CEST

Accuracy with five examples

35 figures
Jev, share of the messages sorted correctly93.9%Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Jev, accuracy, 95% interval, low93.0Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Jev, accuracy, 95% interval, high94.7Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Jev, times it answered none-of-these (always wrong here)14Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Jev (30 September pass), share of the messages sorted correctly93.9%Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
Jev (30 September pass), accuracy, 95% interval, low93.0Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
Jev (30 September pass), accuracy, 95% interval, high94.7Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
Jev (30 September pass), times it answered none-of-these (always wrong here)13Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
GPT-6 Luna, share of the messages sorted correctly94.1%Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
GPT-6 Luna, accuracy, 95% interval, low93.2Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
GPT-6 Luna, accuracy, 95% interval, high94.9Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
GPT-6 Luna, times it answered none-of-these (always wrong here)4Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
Kev-4B, share of the messages sorted correctly87.7%Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
Kev-4B, accuracy, 95% interval, low86.5Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
Kev-4B, accuracy, 95% interval, high88.8Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
Kev-4B, times it answered none-of-these (always wrong here)252Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
Clef, share of the messages sorted correctly95.1%Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef, accuracy, 95% interval, low94.2Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef, accuracy, 95% interval, high95.8Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef, times it answered none-of-these (always wrong here)0Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef-flash, share of the messages sorted correctly95.2%Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef-flash, accuracy, 95% interval, low94.3Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef-flash, accuracy, 95% interval, high95.9Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef-flash, times it answered none-of-these (always wrong here)0Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Kev-27B, share of the messages sorted correctly93.8%Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-27B, accuracy, 95% interval, low92.9Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-27B, accuracy, 95% interval, high94.6Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-27B, times it answered none-of-these (always wrong here)21Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-9B, share of the messages sorted correctly91.3%Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-9B, accuracy, 95% interval, low90.3Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-9B, accuracy, 95% interval, high92.3Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-9B, times it answered none-of-these (always wrong here)110Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-27B, accuracy with none-of-these set aside (its top real intent)94.1%Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-9B, accuracy with none-of-these set aside (its top real intent)93.6%Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-4B, accuracy with none-of-these set aside (a reading chosen after seeing its dev behaviour)93.3%Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST

Against Jev, same messages

30 figures
Messages Jev got right and GPT-6 Luna got wrong27Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
Messages GPT-6 Luna got right and Jev got wrong33Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
Messages Jev got right and Kev-4B got wrong205Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
Messages Kev-4B got right and Jev got wrong15Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
Messages Jev got right and Clef got wrong19Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Messages Clef got right and Jev got wrong54Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Messages Jev got right and Clef-flash got wrong22Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Messages Clef-flash got right and Jev got wrong60Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Messages Jev got right and Kev-27B got wrong27Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Messages Kev-27B got right and Jev got wrong24Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Messages Jev got right and Kev-9B got wrong97Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Messages Kev-9B got right and Jev got wrong17Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
GPT-6 Luna minus Jev, points (Jev pass of 30 September)+0.2 pointsModel Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
GPT-6 Luna minus Jev, 95% interval, low−0.3Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
GPT-6 Luna minus Jev, 95% interval, high+0.7Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
Kev-4B minus Jev, points (Jev pass of 30 September)−6.2 pointsModel Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
Kev-4B minus Jev, 95% interval, low−7.1Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
Kev-4B minus Jev, 95% interval, high−5.3Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
Clef minus Jev, points (same-day Jev pass)+1.1 pointsModel Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef minus Jev, 95% interval, low+0.6Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef minus Jev, 95% interval, high+1.7Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef-flash minus Jev, points (same-day Jev pass)+1.2 pointsModel Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef-flash minus Jev, 95% interval, low+0.7Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef-flash minus Jev, 95% interval, high+1.8Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Kev-27B minus Jev, points (same-day Jev pass)−0.1 pointsModel Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-27B minus Jev, 95% interval, low−0.6Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-27B minus Jev, 95% interval, high+0.4Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-9B minus Jev, points (same-day Jev pass)−2.6 pointsModel Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-9B minus Jev, 95% interval, low−3.3Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-9B minus Jev, 95% interval, high−1.9Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST

Without examples

30 figures
Jev, macro-F1 on the Decision Index's zero-shot question80.0Model Fatigue, run 2 Oct 2026, 17:03–21:41 CEST
Jev, zero-shot macro-F1, 95% interval, low78.6Model Fatigue, run 2 Oct 2026, 17:03–21:41 CEST
Jev, zero-shot macro-F1, 95% interval, high81.1Model Fatigue, run 2 Oct 2026, 17:03–21:41 CEST
Jev, accuracy on the zero-shot question80.7%Model Fatigue, run 2 Oct 2026, 17:03–21:41 CEST
Clef, macro-F1 on the Decision Index's zero-shot question94.2Model Fatigue, run 2 Oct 2026, 17:03–21:41 CEST
Clef, zero-shot macro-F1, 95% interval, low93.3Model Fatigue, run 2 Oct 2026, 17:03–21:41 CEST
Clef, zero-shot macro-F1, 95% interval, high94.9Model Fatigue, run 2 Oct 2026, 17:03–21:41 CEST
Clef, accuracy on the zero-shot question94.2%Model Fatigue, run 2 Oct 2026, 17:03–21:41 CEST
Clef-flash, macro-F1 on the Decision Index's zero-shot question90.8Model Fatigue, run 2 Oct 2026, 17:03–21:41 CEST
Clef-flash, zero-shot macro-F1, 95% interval, low89.8Model Fatigue, run 2 Oct 2026, 17:03–21:41 CEST
Clef-flash, zero-shot macro-F1, 95% interval, high91.7Model Fatigue, run 2 Oct 2026, 17:03–21:41 CEST
Clef-flash, accuracy on the zero-shot question90.9%Model Fatigue, run 2 Oct 2026, 17:03–21:41 CEST
Kev-27B, macro-F1 on the Decision Index's zero-shot question86.4Model Fatigue, run 2 Oct 2026, 17:03–21:41 CEST
Kev-27B, zero-shot macro-F1, 95% interval, low85.1Model Fatigue, run 2 Oct 2026, 17:03–21:41 CEST
Kev-27B, zero-shot macro-F1, 95% interval, high87.4Model Fatigue, run 2 Oct 2026, 17:03–21:41 CEST
Kev-27B, accuracy on the zero-shot question86.8%Model Fatigue, run 2 Oct 2026, 17:03–21:41 CEST
Kev-9B, macro-F1 on the Decision Index's zero-shot question83.0Model Fatigue, run 2 Oct 2026, 17:03–21:41 CEST
Kev-9B, zero-shot macro-F1, 95% interval, low81.6Model Fatigue, run 2 Oct 2026, 17:03–21:41 CEST
Kev-9B, zero-shot macro-F1, 95% interval, high84.1Model Fatigue, run 2 Oct 2026, 17:03–21:41 CEST
Kev-9B, accuracy on the zero-shot question83.3%Model Fatigue, run 2 Oct 2026, 17:03–21:41 CEST
Jev without examples on our 78-option question, jev-2's pass78.3%Model Fatigue (M37's jev-2 run), run 21–24 Sep 2026
Clef, BANKING77 macro-F1 in Cloudflare's launch table94.20Cloudflare, read 2 Oct 2026, 15:31 CEST
Clef-flash, BANKING77 macro-F1 in Cloudflare's launch table90.93Cloudflare, read 2 Oct 2026, 15:31 CEST
Jev, BANKING77 macro-F1 in Cloudflare's launch table79.74Cloudflare, read 2 Oct 2026, 15:31 CEST
Kev 9B (the v1 checkpoint), BANKING77 macro-F1 on the Decision Index, the board's own run84.83The Decision Index (multimodalart), read 2 Oct 2026, 15:32 CEST
Jev, accuracy on our question minus accuracy on the Decision Index's question, points (our arithmetic)+13.2 pointsModel Fatigue, run 2 Oct 2026, 17:03–21:41 CEST
Clef, accuracy on our question minus accuracy on the Decision Index's question, points (our arithmetic)+0.8 pointsModel Fatigue, run 2 Oct 2026, 17:03–21:41 CEST
Clef-flash, accuracy on our question minus accuracy on the Decision Index's question, points (our arithmetic)+4.3 pointsModel Fatigue, run 2 Oct 2026, 17:03–21:41 CEST
Kev-27B, accuracy on our question minus accuracy on the Decision Index's question, points (our arithmetic)+7.0 pointsModel Fatigue, run 2 Oct 2026, 17:03–21:41 CEST
Kev-9B, accuracy on our question minus accuracy on the Decision Index's question, points (our arithmetic)+8.0 pointsModel Fatigue, run 2 Oct 2026, 17:03–21:41 CEST

Speed from Berlin

33 figures
Jev, median time per request from Berlin235 msModel Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Jev, 95th percentile time per request from Berlin335 msModel Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Jev, 99th percentile time per request from Berlin442 msModel Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Jev (30 September pass), median time per request from Berlin253 msModel Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
Jev (30 September pass), 95th percentile time per request from Berlin306 msModel Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
Jev (30 September pass), 99th percentile time per request from Berlin403 msModel Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
GPT-6 Luna, median time per request from Berlin1,097 msModel Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
GPT-6 Luna, 95th percentile time per request from Berlin2,235 msModel Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
GPT-6 Luna, 99th percentile time per request from Berlin3,349 msModel Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
Kev-4B, median time per request from Berlin540 msModel Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
Kev-4B, 95th percentile time per request from Berlin883 msModel Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
Kev-4B, 99th percentile time per request from Berlin1,372 msModel Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
Clef, median time per request from Berlin562 msModel Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef, 95th percentile time per request from Berlin1,147 msModel Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef, 99th percentile time per request from Berlin2,001 msModel Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef-flash, median time per request from Berlin196 msModel Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef-flash, 95th percentile time per request from Berlin1,014 msModel Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef-flash, 99th percentile time per request from Berlin2,172 msModel Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Kev-27B, median time per request from Berlin529 msModel Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-27B, 95th percentile time per request from Berlin551 msModel Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-27B, 99th percentile time per request from Berlin598 msModel Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-9B, median time per request from Berlin437 msModel Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-9B, 95th percentile time per request from Berlin548 msModel Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-9B, 99th percentile time per request from Berlin665 msModel Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Network floor to typesafe's host from the Mac mini, warm, median of 20 (start of Jev's run)171 msModel Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Network floor to openai's host from the Mac mini, warm, median of 20 (start of GPT-6 Luna's run)170 msModel Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
Network floor to openrouter's host from the Mac mini, warm, median of 20 (start of Kev-4B's run)21 msModel Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
Network floor to cloudflare's host from the Mac mini, warm, median of 20 (start of Clef's run)204 msModel Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Network floor to cloudflare's host from the Mac mini, warm, median of 20 (start of Clef-flash's run)204 msModel Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Network floor to modal's host from the Mac mini, warm, median of 20 (start of Kev-27B's run)132 msModel Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Network floor to modal's host from the Mac mini, warm, median of 20 (start of Kev-9B's run)132 msModel Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-27B, median model time reported by its own server (our arithmetic on the records)110 msModel Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-9B, median model time reported by its own server (our arithmetic on the records)35 msModel Fatigue, run 2 Oct 2026, 21:05–22:01 CEST

Cost

19 figures
Jev, dollars per 1,000 decisions (rate card × tokens counted)$0.078Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Jev (30 September pass), dollars per 1,000 decisions (rate card × tokens counted)$0.078Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
GPT-6 Luna, dollars per 1,000 decisions (rate card × tokens counted)$0.143Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
Kev-4B, dollars per 1,000 decisions (as billed)$0.041Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
Clef, dollars per 1,000 decisions (rate card × tokens counted)$0.480Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef-flash, dollars per 1,000 decisions (rate card × tokens counted)$0.180Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Kev-27B, dollars per 1,000 decisions (our GPU time, as Modal billed it)$0.730Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-9B, dollars per 1,000 decisions (our GPU time, as Modal billed it)$0.621Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
GPT-6 Luna, calls billed for writing tokens to OpenAI's prompt cache (none read from it; our count on the records)3,032Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
Kev-27B, dollars per 1,000 decisions at six requests at a time (the zero-shot pass)$0.157Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-9B, dollars per 1,000 decisions at six requests at a time (the zero-shot pass)$0.112Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Jev, input tokens counted per request (our arithmetic: the pass's total over the messages)1,846Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Jev (30 September pass), input tokens counted per request (our arithmetic: the pass's total over the messages)1,846Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
GPT-6 Luna, input tokens counted per request (our arithmetic: the pass's total over the messages)1,066Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
Kev-4B, input tokens counted per request (our arithmetic: the pass's total over the messages)975Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
Clef, input tokens counted per request (our arithmetic: the pass's total over the messages)1,999Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef-flash, input tokens counted per request (our arithmetic: the pass's total over the messages)1,999Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Kev-27B, input tokens counted per request (our arithmetic: the pass's total over the messages)975Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-9B, input tokens counted per request (our arithmetic: the pass's total over the messages)975Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST

Confidence

21 figures
Jev, calibration error (ECE) of its confidence as returned0.035Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Jev, calibration error after a fit on our dev slice0.010Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Jev, how well its confidence separates its right answers from its wrong ones (AUROC)0.840Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
GPT-6 Luna, calibration error (ECE) of its confidence as returned0.025Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
GPT-6 Luna, calibration error after a fit on our dev slice0.021Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
GPT-6 Luna, how well its confidence separates its right answers from its wrong ones (AUROC)0.876Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
Kev-4B, calibration error (ECE) of its confidence as returned0.113Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
Kev-4B, calibration error after a fit on our dev slice0.009Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
Kev-4B, how well its confidence separates its right answers from its wrong ones (AUROC)0.922Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
Clef, calibration error (ECE) of its confidence as returned0.030Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef, calibration error after a fit on our dev slice0.015Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef, how well its confidence separates its right answers from its wrong ones (AUROC)0.900Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef-flash, calibration error (ECE) of its confidence as returned0.143Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef-flash, calibration error after a fit on our dev slice0.006Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef-flash, how well its confidence separates its right answers from its wrong ones (AUROC)0.819Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Kev-27B, calibration error (ECE) of its confidence as returned0.014Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-27B, calibration error after a fit on our dev slice0.009Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-27B, how well its confidence separates its right answers from its wrong ones (AUROC)0.920Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-9B, calibration error (ECE) of its confidence as returned0.038Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-9B, calibration error after a fit on our dev slice0.012Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-9B, how well its confidence separates its right answers from its wrong ones (AUROC)0.906Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST

Workload at the 2% budget

42 figures
Jev, confidence threshold fixed on the dev slice0.990Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Jev, share of messages accepted automatically84.4%Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Jev, error among the accepted2.04%Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Jev, error among the accepted, 95% interval, low1.56Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Jev, error among the accepted, 95% interval, high2.66Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
GPT-6 Luna, confidence threshold fixed on the dev slice0.900Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
GPT-6 Luna, share of messages accepted automatically91.3%Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
GPT-6 Luna, error among the accepted2.99%Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
GPT-6 Luna, error among the accepted, 95% interval, low2.42Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
GPT-6 Luna, error among the accepted, 95% interval, high3.68Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
Kev-4B, confidence threshold fixed on the dev slice0.669Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
Kev-4B, share of messages accepted automatically74.3%Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
Kev-4B, error among the accepted1.75%Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
Kev-4B, error among the accepted, 95% interval, low1.29Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
Kev-4B, error among the accepted, 95% interval, high2.37Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
Clef, confidence threshold fixed on the dev slice0.803Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef, share of messages accepted automatically93.2%Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef, error among the accepted2.40%Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef, error among the accepted, 95% interval, low1.90Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef, error among the accepted, 95% interval, high3.03Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef-flash, confidence threshold fixed on the dev slice0.732Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef-flash, share of messages accepted automatically81.7%Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef-flash, error among the accepted2.15%Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef-flash, error among the accepted, 95% interval, low1.65Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef-flash, error among the accepted, 95% interval, high2.79Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Kev-27B, confidence threshold fixed on the dev slice0.793Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-27B, share of messages accepted automatically91.0%Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-27B, error among the accepted2.28%Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-27B, error among the accepted, 95% interval, low1.79Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-27B, error among the accepted, 95% interval, high2.90Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-9B, confidence threshold fixed on the dev slice0.813Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-9B, share of messages accepted automatically82.7%Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-9B, error among the accepted2.47%Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-9B, error among the accepted, 95% interval, low1.94Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-9B, error among the accepted, 95% interval, high3.15Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Jev, messages per 1,000 sent to a person (our arithmetic)156Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
GPT-6 Luna, messages per 1,000 sent to a person (our arithmetic)87Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
Kev-4B, messages per 1,000 sent to a person (our arithmetic)257Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
Clef, messages per 1,000 sent to a person (our arithmetic)68Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef-flash, messages per 1,000 sent to a person (our arithmetic)183Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Kev-27B, messages per 1,000 sent to a person (our arithmetic)90Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-9B, messages per 1,000 sent to a person (our arithmetic)173Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST

Calibration bins

168 figures
Jev, answers stated 0.2 to 0.3: how many1Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Jev, answers stated 0.2 to 0.3: mean stated confidence0.28Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Jev, answers stated 0.2 to 0.3: share right0.00Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Jev, answers stated 0.3 to 0.4: how many3Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Jev, answers stated 0.3 to 0.4: mean stated confidence0.35Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Jev, answers stated 0.3 to 0.4: share right0.67Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Jev, answers stated 0.4 to 0.5: how many17Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Jev, answers stated 0.4 to 0.5: mean stated confidence0.45Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Jev, answers stated 0.4 to 0.5: share right0.35Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Jev, answers stated 0.5 to 0.6: how many37Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Jev, answers stated 0.5 to 0.6: mean stated confidence0.55Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Jev, answers stated 0.5 to 0.6: share right0.49Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Jev, answers stated 0.6 to 0.7: how many39Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Jev, answers stated 0.6 to 0.7: mean stated confidence0.64Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Jev, answers stated 0.6 to 0.7: share right0.44Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Jev, answers stated 0.7 to 0.8: how many46Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Jev, answers stated 0.7 to 0.8: mean stated confidence0.75Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Jev, answers stated 0.7 to 0.8: share right0.63Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Jev, answers stated 0.8 to 0.9: how many81Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Jev, answers stated 0.8 to 0.9: mean stated confidence0.85Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Jev, answers stated 0.8 to 0.9: share right0.72Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Jev, answers stated 0.9 to 1.0: how many2,856Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Jev, answers stated 0.9 to 1.0: mean stated confidence1.00Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Jev, answers stated 0.9 to 1.0: share right0.97Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
GPT-6 Luna, answers stated 0.3 to 0.4: how many2Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
GPT-6 Luna, answers stated 0.3 to 0.4: mean stated confidence0.38Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
GPT-6 Luna, answers stated 0.3 to 0.4: share right0.50Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
GPT-6 Luna, answers stated 0.4 to 0.5: how many9Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
GPT-6 Luna, answers stated 0.4 to 0.5: mean stated confidence0.46Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
GPT-6 Luna, answers stated 0.4 to 0.5: share right0.33Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
GPT-6 Luna, answers stated 0.5 to 0.6: how many22Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
GPT-6 Luna, answers stated 0.5 to 0.6: mean stated confidence0.56Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
GPT-6 Luna, answers stated 0.5 to 0.6: share right0.32Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
GPT-6 Luna, answers stated 0.6 to 0.7: how many29Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
GPT-6 Luna, answers stated 0.6 to 0.7: mean stated confidence0.65Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
GPT-6 Luna, answers stated 0.6 to 0.7: share right0.59Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
GPT-6 Luna, answers stated 0.7 to 0.8: how many69Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
GPT-6 Luna, answers stated 0.7 to 0.8: mean stated confidence0.75Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
GPT-6 Luna, answers stated 0.7 to 0.8: share right0.59Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
GPT-6 Luna, answers stated 0.8 to 0.9: how many138Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
GPT-6 Luna, answers stated 0.8 to 0.9: mean stated confidence0.86Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
GPT-6 Luna, answers stated 0.8 to 0.9: share right0.74Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
GPT-6 Luna, answers stated 0.9 to 1.0: how many2,811Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
GPT-6 Luna, answers stated 0.9 to 1.0: mean stated confidence0.98Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
GPT-6 Luna, answers stated 0.9 to 1.0: share right0.97Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
Kev-4B, answers stated 0.2 to 0.3: how many16Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
Kev-4B, answers stated 0.2 to 0.3: mean stated confidence0.27Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
Kev-4B, answers stated 0.2 to 0.3: share right0.19Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
Kev-4B, answers stated 0.3 to 0.4: how many147Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
Kev-4B, answers stated 0.3 to 0.4: mean stated confidence0.36Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
Kev-4B, answers stated 0.3 to 0.4: share right0.29Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
Kev-4B, answers stated 0.4 to 0.5: how many264Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
Kev-4B, answers stated 0.4 to 0.5: mean stated confidence0.45Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
Kev-4B, answers stated 0.4 to 0.5: share right0.48Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
Kev-4B, answers stated 0.5 to 0.6: how many201Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
Kev-4B, answers stated 0.5 to 0.6: mean stated confidence0.55Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
Kev-4B, answers stated 0.5 to 0.6: share right0.71Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
Kev-4B, answers stated 0.6 to 0.7: how many237Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
Kev-4B, answers stated 0.6 to 0.7: mean stated confidence0.65Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
Kev-4B, answers stated 0.6 to 0.7: share right0.88Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
Kev-4B, answers stated 0.7 to 0.8: how many369Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
Kev-4B, answers stated 0.7 to 0.8: mean stated confidence0.75Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
Kev-4B, answers stated 0.7 to 0.8: share right0.95Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
Kev-4B, answers stated 0.8 to 0.9: how many797Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
Kev-4B, answers stated 0.8 to 0.9: mean stated confidence0.86Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
Kev-4B, answers stated 0.8 to 0.9: share right0.98Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
Kev-4B, answers stated 0.9 to 1.0: how many1,049Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
Kev-4B, answers stated 0.9 to 1.0: mean stated confidence0.93Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
Kev-4B, answers stated 0.9 to 1.0: share right1.00Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
Clef, answers stated 0.2 to 0.3: how many5Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef, answers stated 0.2 to 0.3: mean stated confidence0.28Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef, answers stated 0.2 to 0.3: share right0.40Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef, answers stated 0.3 to 0.4: how many13Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef, answers stated 0.3 to 0.4: mean stated confidence0.36Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef, answers stated 0.3 to 0.4: share right0.69Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef, answers stated 0.4 to 0.5: how many23Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef, answers stated 0.4 to 0.5: mean stated confidence0.45Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef, answers stated 0.4 to 0.5: share right0.57Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef, answers stated 0.5 to 0.6: how many38Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef, answers stated 0.5 to 0.6: mean stated confidence0.55Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef, answers stated 0.5 to 0.6: share right0.50Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef, answers stated 0.6 to 0.7: how many55Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef, answers stated 0.6 to 0.7: mean stated confidence0.65Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef, answers stated 0.6 to 0.7: share right0.56Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef, answers stated 0.7 to 0.8: how many69Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef, answers stated 0.7 to 0.8: mean stated confidence0.76Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef, answers stated 0.7 to 0.8: share right0.70Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef, answers stated 0.8 to 0.9: how many187Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef, answers stated 0.8 to 0.9: mean stated confidence0.86Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef, answers stated 0.8 to 0.9: share right0.83Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef, answers stated 0.9 to 1.0: how many2,690Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef, answers stated 0.9 to 1.0: mean stated confidence0.96Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef, answers stated 0.9 to 1.0: share right0.99Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef-flash, answers stated 0.2 to 0.3: how many7Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef-flash, answers stated 0.2 to 0.3: mean stated confidence0.26Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef-flash, answers stated 0.2 to 0.3: share right0.14Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef-flash, answers stated 0.3 to 0.4: how many12Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef-flash, answers stated 0.3 to 0.4: mean stated confidence0.35Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef-flash, answers stated 0.3 to 0.4: share right0.33Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef-flash, answers stated 0.4 to 0.5: how many59Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef-flash, answers stated 0.4 to 0.5: mean stated confidence0.46Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef-flash, answers stated 0.4 to 0.5: share right0.68Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef-flash, answers stated 0.5 to 0.6: how many87Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef-flash, answers stated 0.5 to 0.6: mean stated confidence0.56Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef-flash, answers stated 0.5 to 0.6: share right0.78Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef-flash, answers stated 0.6 to 0.7: how many241Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef-flash, answers stated 0.6 to 0.7: mean stated confidence0.66Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef-flash, answers stated 0.6 to 0.7: share right0.87Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef-flash, answers stated 0.7 to 0.8: how many736Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef-flash, answers stated 0.7 to 0.8: mean stated confidence0.76Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef-flash, answers stated 0.7 to 0.8: share right0.95Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef-flash, answers stated 0.8 to 0.9: how many1,339Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef-flash, answers stated 0.8 to 0.9: mean stated confidence0.85Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef-flash, answers stated 0.8 to 0.9: share right0.98Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef-flash, answers stated 0.9 to 1.0: how many599Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef-flash, answers stated 0.9 to 1.0: mean stated confidence0.92Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef-flash, answers stated 0.9 to 1.0: share right1.00Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Kev-27B, answers stated 0.1 to 0.2: how many1Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-27B, answers stated 0.1 to 0.2: mean stated confidence0.19Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-27B, answers stated 0.1 to 0.2: share right1.00Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-27B, answers stated 0.2 to 0.3: how many9Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-27B, answers stated 0.2 to 0.3: mean stated confidence0.25Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-27B, answers stated 0.2 to 0.3: share right0.22Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-27B, answers stated 0.3 to 0.4: how many25Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-27B, answers stated 0.3 to 0.4: mean stated confidence0.37Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-27B, answers stated 0.3 to 0.4: share right0.36Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-27B, answers stated 0.4 to 0.5: how many51Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-27B, answers stated 0.4 to 0.5: mean stated confidence0.45Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-27B, answers stated 0.4 to 0.5: share right0.43Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-27B, answers stated 0.5 to 0.6: how many55Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-27B, answers stated 0.5 to 0.6: mean stated confidence0.55Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-27B, answers stated 0.5 to 0.6: share right0.45Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-27B, answers stated 0.6 to 0.7: how many62Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-27B, answers stated 0.6 to 0.7: mean stated confidence0.65Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-27B, answers stated 0.6 to 0.7: share right0.56Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-27B, answers stated 0.7 to 0.8: how many75Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-27B, answers stated 0.7 to 0.8: mean stated confidence0.75Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-27B, answers stated 0.7 to 0.8: share right0.77Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-27B, answers stated 0.8 to 0.9: how many204Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-27B, answers stated 0.8 to 0.9: mean stated confidence0.86Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-27B, answers stated 0.8 to 0.9: share right0.86Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-27B, answers stated 0.9 to 1.0: how many2,598Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-27B, answers stated 0.9 to 1.0: mean stated confidence0.98Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-27B, answers stated 0.9 to 1.0: share right0.99Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-9B, answers stated 0.2 to 0.3: how many2Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-9B, answers stated 0.2 to 0.3: mean stated confidence0.29Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-9B, answers stated 0.2 to 0.3: share right0.50Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-9B, answers stated 0.3 to 0.4: how many21Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-9B, answers stated 0.3 to 0.4: mean stated confidence0.36Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-9B, answers stated 0.3 to 0.4: share right0.24Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-9B, answers stated 0.4 to 0.5: how many88Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-9B, answers stated 0.4 to 0.5: mean stated confidence0.46Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-9B, answers stated 0.4 to 0.5: share right0.38Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-9B, answers stated 0.5 to 0.6: how many102Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-9B, answers stated 0.5 to 0.6: mean stated confidence0.55Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-9B, answers stated 0.5 to 0.6: share right0.47Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-9B, answers stated 0.6 to 0.7: how many115Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-9B, answers stated 0.6 to 0.7: mean stated confidence0.65Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-9B, answers stated 0.6 to 0.7: share right0.64Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-9B, answers stated 0.7 to 0.8: how many180Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-9B, answers stated 0.7 to 0.8: mean stated confidence0.75Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-9B, answers stated 0.7 to 0.8: share right0.81Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-9B, answers stated 0.8 to 0.9: how many440Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-9B, answers stated 0.8 to 0.9: mean stated confidence0.86Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-9B, answers stated 0.8 to 0.9: share right0.92Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-9B, answers stated 0.9 to 1.0: how many2,132Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-9B, answers stated 0.9 to 1.0: mean stated confidence0.96Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-9B, answers stated 0.9 to 1.0: share right0.99Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST

Repeat and drift

21 figures
Jev (30 September pass), same answer on the repeat pass199Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
Jev (30 September pass), messages on the repeat pass200Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
GPT-6 Luna, same answer on the repeat pass197Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
GPT-6 Luna, messages on the repeat pass200Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
Kev-4B, same answer on the repeat pass200Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
Kev-4B, messages on the repeat pass200Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
Clef, same answer on the repeat pass200Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef, messages on the repeat pass200Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef-flash, same answer on the repeat pass200Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Clef-flash, messages on the repeat pass200Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Kev-27B, same answer on the repeat pass200Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-27B, messages on the repeat pass200Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-9B, same answer on the repeat pass200Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Kev-9B, messages on the repeat pass200Model Fatigue, run 2 Oct 2026, 21:05–22:01 CEST
Jev's two passes (30 September, 2 October), messages with the same answer3,073Model Fatigue, run 2 Oct 2026, 15:49–16:57 CEST
Jev, accuracy in jev-2's pass94.0%Model Fatigue (M37's jev-2 run), run 21 Sep 2026
Jev, median time per request in jev-2's pass313 msModel Fatigue (M37's jev-2 run), run 21 Sep 2026
Jev, messages with the same answer as in jev-2's pass3,075Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST
GPT-6 Luna, accuracy in jev-2's pass94.0%Model Fatigue (M37's jev-2 run), run 24 Sep 2026
GPT-6 Luna, median time per request in jev-2's pass1,196 msModel Fatigue (M37's jev-2 run), run 24 Sep 2026
GPT-6 Luna, messages with the same answer as in jev-2's pass3,047Model Fatigue, run 30 Sep 2026, 21:42–23:38 CEST

Seen it before?

41 figures
Kev-4B, zero-shot accuracy on training messages63.5%Model Fatigue, run 30 Sep 2026 21:30 to 2 Oct 2026 21:42 CEST
Kev-4B, zero-shot accuracy on paraphrases of them55.5%Model Fatigue, run 30 Sep 2026 21:30 to 2 Oct 2026 21:42 CEST
Clef, zero-shot accuracy on training messages94.5%Model Fatigue, run 30 Sep 2026 21:30 to 2 Oct 2026 21:42 CEST
Clef, zero-shot accuracy on paraphrases of them87.0%Model Fatigue, run 30 Sep 2026 21:30 to 2 Oct 2026 21:42 CEST
Clef-flash, zero-shot accuracy on training messages95.5%Model Fatigue, run 30 Sep 2026 21:30 to 2 Oct 2026 21:42 CEST
Clef-flash, zero-shot accuracy on paraphrases of them83.5%Model Fatigue, run 30 Sep 2026 21:30 to 2 Oct 2026 21:42 CEST
Kev-27B, zero-shot accuracy on training messages79.5%Model Fatigue, run 30 Sep 2026 21:30 to 2 Oct 2026 21:42 CEST
Kev-27B, zero-shot accuracy on paraphrases of them65.5%Model Fatigue, run 30 Sep 2026 21:30 to 2 Oct 2026 21:42 CEST
Kev-9B, zero-shot accuracy on training messages74.0%Model Fatigue, run 30 Sep 2026 21:30 to 2 Oct 2026 21:42 CEST
Kev-9B, zero-shot accuracy on paraphrases of them62.5%Model Fatigue, run 30 Sep 2026 21:30 to 2 Oct 2026 21:42 CEST
Jev, zero-shot accuracy on training messages (jev-2's probe)72.5%Model Fatigue (M37's jev-2 run), run 21 Sep 2026
Jev, zero-shot accuracy on paraphrases (jev-2's probe)68.0%Model Fatigue (M37's jev-2 run), run 21 Sep 2026
Jev, originals minus paraphrases, points (jev-2's probe)+4.5 pointsModel Fatigue (M37's jev-2 run), run 21 Sep 2026
Jev, gap, 95% interval, low (jev-2's probe)+0.5Model Fatigue (M37's jev-2 run), run 21 Sep 2026
Jev, gap, 95% interval, high (jev-2's probe)+8.5Model Fatigue (M37's jev-2 run), run 21 Sep 2026
GPT-6 Luna, zero-shot accuracy on training messages (jev-2's probe)76.5%Model Fatigue (M37's jev-2 run), run 24 Sep 2026
GPT-6 Luna, zero-shot accuracy on paraphrases (jev-2's probe)71.5%Model Fatigue (M37's jev-2 run), run 24 Sep 2026
GPT-6 Luna, originals minus paraphrases, points (jev-2's probe)+5.0 pointsModel Fatigue (M37's jev-2 run), run 24 Sep 2026
GPT-6 Luna, gap, 95% interval, low (jev-2's probe)+1.0Model Fatigue (M37's jev-2 run), run 24 Sep 2026
GPT-6 Luna, gap, 95% interval, high (jev-2's probe)+9.5Model Fatigue (M37's jev-2 run), run 24 Sep 2026
Training messages in the probe (and as many paraphrases)200Model Fatigue, run 30 Sep 2026 21:30 to 2 Oct 2026 21:42 CEST
Five-nearest-neighbour vote (no memory possible), originals94.5%Model Fatigue (M37's jev-2 run), run 21–24 Sep 2026
Five-nearest-neighbour vote, paraphrases79.5%Model Fatigue (M37's jev-2 run), run 21–24 Sep 2026
Five-nearest-neighbour vote, gap, points+15.0 pointsModel Fatigue (M37's jev-2 run), run 21–24 Sep 2026
Five-nearest-neighbour vote, gap, 95% interval, low+10.0Model Fatigue (M37's jev-2 run), run 21–24 Sep 2026
Five-nearest-neighbour vote, gap, 95% interval, high+20.5Model Fatigue (M37's jev-2 run), run 21–24 Sep 2026
Kev-4B, originals minus paraphrases, points+8.0 pointsModel Fatigue, run 30 Sep 2026 21:30 to 2 Oct 2026 21:42 CEST
Kev-4B, gap, 95% interval, low+2.5Model Fatigue, run 30 Sep 2026 21:30 to 2 Oct 2026 21:42 CEST
Kev-4B, gap, 95% interval, high+14.0Model Fatigue, run 30 Sep 2026 21:30 to 2 Oct 2026 21:42 CEST
Clef, originals minus paraphrases, points+7.5 pointsModel Fatigue, run 30 Sep 2026 21:30 to 2 Oct 2026 21:42 CEST
Clef, gap, 95% interval, low+4.0Model Fatigue, run 30 Sep 2026 21:30 to 2 Oct 2026 21:42 CEST
Clef, gap, 95% interval, high+11.5Model Fatigue, run 30 Sep 2026 21:30 to 2 Oct 2026 21:42 CEST
Clef-flash, originals minus paraphrases, points+12.0 pointsModel Fatigue, run 30 Sep 2026 21:30 to 2 Oct 2026 21:42 CEST
Clef-flash, gap, 95% interval, low+7.5Model Fatigue, run 30 Sep 2026 21:30 to 2 Oct 2026 21:42 CEST
Clef-flash, gap, 95% interval, high+16.5Model Fatigue, run 30 Sep 2026 21:30 to 2 Oct 2026 21:42 CEST
Kev-27B, originals minus paraphrases, points+14.0 pointsModel Fatigue, run 30 Sep 2026 21:30 to 2 Oct 2026 21:42 CEST
Kev-27B, gap, 95% interval, low+9.0Model Fatigue, run 30 Sep 2026 21:30 to 2 Oct 2026 21:42 CEST
Kev-27B, gap, 95% interval, high+19.0Model Fatigue, run 30 Sep 2026 21:30 to 2 Oct 2026 21:42 CEST
Kev-9B, originals minus paraphrases, points+11.5 pointsModel Fatigue, run 30 Sep 2026 21:30 to 2 Oct 2026 21:42 CEST
Kev-9B, gap, 95% interval, low+6.5Model Fatigue, run 30 Sep 2026 21:30 to 2 Oct 2026 21:42 CEST
Kev-9B, gap, 95% interval, high+17.0Model Fatigue, run 30 Sep 2026 21:30 to 2 Oct 2026 21:42 CEST

Reference rows from jev-2

19 figures
A trained classifier (bge-small + logistic regression), accuracy93.1%Model Fatigue (M37's jev-2 run), run 21–24 Sep 2026
A trained classifier (bge-small + logistic regression), accuracy, 95% interval, low92.2Model Fatigue (M37's jev-2 run), run 21–24 Sep 2026
A trained classifier (bge-small + logistic regression), accuracy, 95% interval, high94.0Model Fatigue (M37's jev-2 run), run 21–24 Sep 2026
A trained classifier (bge-small + logistic regression), median time per request (on the Mac, no network)9 msModel Fatigue (M37's jev-2 run), run 21–24 Sep 2026
The single nearest labelled example, accuracy92.6%Model Fatigue (M37's jev-2 run), run 21–24 Sep 2026
The single nearest labelled example, accuracy, 95% interval, low91.6Model Fatigue (M37's jev-2 run), run 21–24 Sep 2026
The single nearest labelled example, accuracy, 95% interval, high93.5Model Fatigue (M37's jev-2 run), run 21–24 Sep 2026
The single nearest labelled example, median time per request (on the Mac, no network)10 msModel Fatigue (M37's jev-2 run), run 21–24 Sep 2026
Ling 3.0 Flash, accuracy93.8%Model Fatigue (M37's jev-2 run), run 21–24 Sep 2026
Ling 3.0 Flash, accuracy, 95% interval, low92.9Model Fatigue (M37's jev-2 run), run 21–24 Sep 2026
Ling 3.0 Flash, accuracy, 95% interval, high94.6Model Fatigue (M37's jev-2 run), run 21–24 Sep 2026
Ling 3.0 Flash, median time per request from Berlin723 msModel Fatigue (M37's jev-2 run), run 21–24 Sep 2026
Ling 3.0 Flash, dollars per 1,000 decisions, as billed$0.020Model Fatigue (M37's jev-2 run), run 21–24 Sep 2026
DeepSeek V4.1 Flash, accuracy93.6%Model Fatigue (M37's jev-2 run), run 21–24 Sep 2026
DeepSeek V4.1 Flash, accuracy, 95% interval, low92.7Model Fatigue (M37's jev-2 run), run 21–24 Sep 2026
DeepSeek V4.1 Flash, accuracy, 95% interval, high94.4Model Fatigue (M37's jev-2 run), run 21–24 Sep 2026
DeepSeek V4.1 Flash, median time per request from Berlin953 msModel Fatigue (M37's jev-2 run), run 21–24 Sep 2026
DeepSeek V4.1 Flash, dollars per 1,000 decisions, as billed$0.038Model Fatigue (M37's jev-2 run), run 21–24 Sep 2026
Labelled training messages the trained classifier learned from (and the pool the five examples come from)9,405Model Fatigue (M37's jev-2 run), frozen 21 Sep 2026

Sources

Our own runs, with when each ran, and the pages this page quotes, with when we read them. We keep every model call and a copy of each page as we read it, so a figure can be checked against what the run or the page said at the time.