The short answer
On one task, sorting 3,080 bank-support messages into 77 intents with five labelled examples each, Clef-flash (95.2%) and Clef (95.1%) got about a point more right than Jev (93.9%). GPT-6 Luna and Kev-27B were level with Jev.
On the Decision Index's question, the message alone with no examples, the gap is far wider: Clef got 94.2% right and Jev 80.7%. Among the decision models, only Kev-4B cost less per thousand decisions than Jev ($0.078), and from Berlin Clef-flash (196 ms) and Jev (235 ms) were the fastest at the median.
Kev-27B's confidence matched its accuracy most closely. Every figure is dated, comes from our own runs or the page it names, and links to its entry in the table of every number at the foot.
What this page measures
A decision model takes some text and a question with fixed answers, and returns one answer with a probability for each option instead of writing text. Jev made the idea popular; Clef, Kev and others now answer the same kind of request. This page holds our main measurements of them, one question at a time, and grows as new models ship. Each row says when it was measured.
So far all of it comes from one task: 3,080 messages that customers sent to a bank, from the public BANKING77 test split, each to be sorted into one of 77 intents. Every model gets the same question: the message, five labelled examples of similar messages found by a fixed search over the training set, and 78 options, the intents plus none-of-these. We sent the calls one at a time from a Mac mini on a home connection in Berlin.
One task is a narrow base. A model that leads on bank-support intents can trail on your own labels, so read these numbers as a dated measurement on one job. The method section says how they were made, and the files at the foot have every number on this page.
The models
Each row is the newest pass of that model on our test, dated by when it ran.
| Model | Made by | How we called it | Measured | Served as |
|---|---|---|---|---|
| Jev | TypeSafe | TypeSafe's API, direct | 2 October 2026, and 30 September | jev-1.13.0 |
| Clef | Cloudflare, open weights | Cloudflare Workers AI, entering at its Berlin edge | 2 October 2026 | clef |
| Clef-flash | Cloudflare, open weights | Cloudflare Workers AI, entering at its Berlin edge | 2 October 2026 | clef-flash |
| GPT-6 Luna | OpenAI; a general model, asked for one label and a confidence | OpenAI's Responses API, direct, reasoning off | 30 September 2026 | gpt-6-luna |
| Kev-27B | Jared Palmer, open weights | our own H100 on Modal (eu-north), one request at a time | 2 October 2026 | jaredpalmer/kev-27b |
| Kev-9B | Jared Palmer, open weights | our own H100 on Modal (eu-north), one request at a time | 2 October 2026 | jaredpalmer/kev-9b |
| Kev-4B | Jared Palmer, open weights | OpenRouter, which sent it to SiliconFlow | 30 September 2026 | jaredpalmer/kev-4b-20260924 |
Which one gets the most right?
With five examples, Clef-flash and Clef sorted about a point more messages correctly than Jev, and the interval on each difference stays above zero. GPT-6 Luna and Kev-27B were level with Jev within the noise of 3,080 messages.
Kev-9B and Kev-4B fell behind mostly by answering none-of-these on messages that had a real intent. Kev-4B did it 252 times. With that option set aside, its top real intent was right 93.3% of the time, close to Jev, and Kev-9B's was right 93.6% of the time. A caller who offers an "other" option should expect these two to take it more often than the rest.
Share of messages sorted correctly
With five labelled examples. Dot: measured. Line: 95% interval.
Accuracy with five labelled examples
Share of the 3,080 messages each model sorted into the right intent, with a 95% interval. The difference from Jev is measured message by message against Jev's pass of the same day.
| right | 95% interval | answered none-of-these | minus Jev, points [95% interval] | measured | |
|---|---|---|---|---|---|
| Clef-flash | 95.2% | [94.3, 95.9] | 0 | +1.2 points [+0.7, +1.8] | 2 October |
| Clef | 95.1% | [94.2, 95.8] | 0 | +1.1 points [+0.6, +1.7] | 2 October |
| GPT-6 Luna | 94.1% | [93.2, 94.9] | 4 | +0.2 points [−0.3, +0.7] | 30 September |
| Jev | 93.9% | [93.0, 94.7] | 14 | the reference | 2 October |
| Kev-27B | 93.8% | [92.9, 94.6] | 21 | −0.1 points [−0.6, +0.4] | 2 October |
| Kev-9B | 91.3% | [90.3, 92.3] | 110 | −2.6 points [−3.3, −1.9] | 2 October |
| Kev-4B | 87.7% | [86.5, 88.8] | 252 | −6.2 points [−7.1, −5.3] | 30 September |
Without examples, and with five
Cloudflare launched Clef with a BANKING77 score of 94.20 against Jev's 79.74. Both use the community Decision Index's test: Jev's figure is the Index's own run, Clef's is Cloudflare's run of it. That test asks a different question from ours: the message alone and 77 options described by the intent names, no examples and no none-of-these, scored by macro-F1. We replicated that request on the same messages and got 94.2 for Clef and 80.0 for Jev, so Cloudflare's number holds for the question it asks.
On our question, with five labelled examples, most of the gap closes: Jev scored 13.2 points higher than on the Index's, Clef 0.8 points. The two questions also differ in layout and in offering none-of-these, so not all of that is the examples; on our layout without examples, Jev scored 78.3% in jev-2. If your setup can't show the model any examples, Clef is far ahead on this task. If it can, the models are within a few points of each other, and price and speed decide.
The Decision Index's question against ours
Share right on the same messages: the message alone (the Decision Index's question), and with five labelled examples (ours).
The numbers in this chart
| the Decision Index's question | ours, with five examples | difference, points | |
|---|---|---|---|
| Clef | 94.2% | 95.1% | +0.8 points |
| Clef-flash | 90.9% | 95.2% | +4.3 points |
| Kev-27B | 86.8% | 93.8% | +7.0 points |
| Kev-9B | 83.3% | 91.3% | +8.0 points |
| Jev | 80.7% | 93.9% | +13.2 points |
The Decision Index's question and ours
The Decision Index's question gives the message alone and 77 numbered options described by the intent names, with no none-of-these; it is scored by macro-F1 and here also by share right. Ours adds five labelled examples and offers the 78 named options. Same messages.
| BANKING77 macro-F1, published | macro-F1 on that question, measured [95% interval] | right on that question | right on ours | difference, points | |
|---|---|---|---|---|---|
| Clef | 94.20, Cloudflare's | 94.2 [93.3, 94.9] | 94.2% | 95.1% | +0.8 points |
| Clef-flash | 90.93, Cloudflare's | 90.8 [89.8, 91.7] | 90.9% | 95.2% | +4.3 points |
| Kev-27B | not on the board | 86.4 [85.1, 87.4] | 86.8% | 93.8% | +7.0 points |
| Kev-9B | 84.83, the board's run of the earlier version | 83.0 [81.6, 84.1] | 83.3% | 91.3% | +8.0 points |
| Jev | 79.74, the board's run, in Cloudflare's table | 80.0 [78.6, 81.1] | 80.7% | 93.9% | +13.2 points |
How fast is it from Berlin?
For a call this short, much of the time is the network, so every time here is from Berlin and comes with the network floor to the same host. Jev's median was 235 ms, and a bare request to TypeSafe's host took 171 ms at the start of that run. Clef-flash had the fastest median, 196 ms, but a longer tail than Jev; its slowest calls bunched in the first part of its pass. Clef took 562 ms at the median and GPT-6 Luna 1,097 ms.
Kev-27B and Kev-9B ran on H100s we rented on Modal. The model itself took 110 ms and 35 ms per request by its server's count, and the rest of each call was the path from Berlin to Modal's machine. These are times of our deployment, which we didn't tune, not of a service you can call.
Time per request, from Berlin
The bar runs to the median; the ticks mark p95 and p99. One request at a time.
Time per request from Berlin
Wall time of each call, request sent to answer parsed, one call at a time over a kept-alive connection from a Mac mini on a Berlin home line. p95 is the time that one call in twenty took longer than; p99, one in a hundred.
| median | p95 | p99 | network floor | what the floor is | |
|---|---|---|---|---|---|
| Clef-flash | 196 ms | 1,014 ms | 2,172 ms | 204 ms | Cloudflare's API host, which is not the path model calls take |
| Jev | 235 ms | 335 ms | 442 ms | 171 ms | TypeSafe's API host |
| Kev-9B | 437 ms | 548 ms | 665 ms | 132 ms | Modal's edge; the model itself took 35 ms, by its server's count |
| Kev-27B | 529 ms | 551 ms | 598 ms | 132 ms | Modal's edge; the model itself took 110 ms, by its server's count |
| Kev-4B | 540 ms | 883 ms | 1,372 ms | 21 ms | OpenRouter's edge only; its hop to SiliconFlow is inside every call |
| Clef | 562 ms | 1,147 ms | 2,001 ms | 204 ms | Cloudflare's API host, which is not the path model calls take |
| GPT-6 Luna | 1,097 ms | 2,235 ms | 3,349 ms | 170 ms | OpenAI's API host |
What a thousand decisions cost
Among the hosted models, Kev-4B through OpenRouter was the cheapest at $0.041 per thousand decisions, then Jev at $0.078. Clef-flash cost $0.180 and Clef $0.480.
Price per token is a poor guide here, because each provider counts the same request differently. Jev counted 1,846 input tokens per request and Clef 1,999, while Kev's server counted 975 for the same body. GPT-6 Luna counted 1,066, and 3,032 of its 3,080 calls also paid OpenAI's charge for writing the prompt to its cache, which no call read from.
Running Kev on our own GPU one request at a time cost $0.730 per thousand for Kev-27B. On the shorter requests without examples, sent six at a time, it cost $0.157, still above Jev. At our volume, running the weights yourself buys control of the model and the data path, not a lower price.
Dollars per thousand decisions
Our test's requests, as each provider counts and charges them.
Cost per thousand decisions
Dollars for a thousand messages of our test, each sent with its five examples and the 78 options.
| per thousand decisions | input tokens counted per request | how it was counted | |
|---|---|---|---|
| Kev-4B | $0.041 | 975 | as OpenRouter billed each call |
| Jev | $0.078 | 1,846 | TypeSafe's rate card × input tokens counted; output is free |
| GPT-6 Luna | $0.143 | 1,066 | OpenAI's rate card × tokens, with a cache-write charge on 3,032 of the 3,080 calls |
| Clef-flash | $0.180 | 1,999 | Cloudflare's rate card × input tokens, which matches the neurons it bills |
| Clef | $0.480 | 1,999 | Cloudflare's rate card × input tokens, which matches the neurons it bills |
| Kev-9B | $0.621 | 975 | Modal's bill for our H100 time, one request at a time; $0.112 on the shorter requests without examples, six at a time |
| Kev-27B | $0.730 | 975 | Modal's bill for our H100 time, one request at a time; $0.157 on the shorter requests without examples, six at a time |
How far can you trust the confidence?
Each of these models gives a confidence with its answer, and the practical use is to accept the confident answers automatically and send the rest to a person. That works if the confidence tracks how often the model is right.
Kev-27B's did best: across the range, what it stated was close to how often it was right, with a calibration error of 0.014. Among the models that return probabilities, Clef (0.030) and Jev (0.035) came next. Clef-flash and Kev-4B state numbers well below how often they are right, so a threshold for them has to be set on your own data rather than read off the number. After a simple fit on held-out messages every model's calibration error was 0.021 or lower, and Clef-flash's fell from 0.143 to 0.006.
GPT-6 Luna doesn't return probabilities on this setup. Its confidence is a number it writes in its answer, and it is reasonably well placed, but it is a different kind of number from the others'.
Stated confidence against the share right
Each model's answers grouped by the confidence it gave. On the dashed line a model is right exactly as often as it says.
The numbers in this chart
| stated | answers | mean stated | share right | |
|---|---|---|---|---|
| Kev-27B | 0.1 to 0.2 | 1 | 0.19 | 1.00 |
| Kev-27B | 0.2 to 0.3 | 9 | 0.25 | 0.22 |
| Kev-27B | 0.3 to 0.4 | 25 | 0.37 | 0.36 |
| Kev-27B | 0.4 to 0.5 | 51 | 0.45 | 0.43 |
| Kev-27B | 0.5 to 0.6 | 55 | 0.55 | 0.45 |
| Kev-27B | 0.6 to 0.7 | 62 | 0.65 | 0.56 |
| Kev-27B | 0.7 to 0.8 | 75 | 0.75 | 0.77 |
| Kev-27B | 0.8 to 0.9 | 204 | 0.86 | 0.86 |
| Kev-27B | 0.9 to 1.0 | 2,598 | 0.98 | 0.99 |
| GPT-6 Luna | 0.3 to 0.4 | 2 | 0.38 | 0.50 |
| GPT-6 Luna | 0.4 to 0.5 | 9 | 0.46 | 0.33 |
| GPT-6 Luna | 0.5 to 0.6 | 22 | 0.56 | 0.32 |
| GPT-6 Luna | 0.6 to 0.7 | 29 | 0.65 | 0.59 |
| GPT-6 Luna | 0.7 to 0.8 | 69 | 0.75 | 0.59 |
| GPT-6 Luna | 0.8 to 0.9 | 138 | 0.86 | 0.74 |
| GPT-6 Luna | 0.9 to 1.0 | 2,811 | 0.98 | 0.97 |
| Clef | 0.2 to 0.3 | 5 | 0.28 | 0.40 |
| Clef | 0.3 to 0.4 | 13 | 0.36 | 0.69 |
| Clef | 0.4 to 0.5 | 23 | 0.45 | 0.57 |
| Clef | 0.5 to 0.6 | 38 | 0.55 | 0.50 |
| Clef | 0.6 to 0.7 | 55 | 0.65 | 0.56 |
| Clef | 0.7 to 0.8 | 69 | 0.76 | 0.70 |
| Clef | 0.8 to 0.9 | 187 | 0.86 | 0.83 |
| Clef | 0.9 to 1.0 | 2,690 | 0.96 | 0.99 |
| Jev | 0.2 to 0.3 | 1 | 0.28 | 0.00 |
| Jev | 0.3 to 0.4 | 3 | 0.35 | 0.67 |
| Jev | 0.4 to 0.5 | 17 | 0.45 | 0.35 |
| Jev | 0.5 to 0.6 | 37 | 0.55 | 0.49 |
| Jev | 0.6 to 0.7 | 39 | 0.64 | 0.44 |
| Jev | 0.7 to 0.8 | 46 | 0.75 | 0.63 |
| Jev | 0.8 to 0.9 | 81 | 0.85 | 0.72 |
| Jev | 0.9 to 1.0 | 2,856 | 1.00 | 0.97 |
| Kev-9B | 0.2 to 0.3 | 2 | 0.29 | 0.50 |
| Kev-9B | 0.3 to 0.4 | 21 | 0.36 | 0.24 |
| Kev-9B | 0.4 to 0.5 | 88 | 0.46 | 0.38 |
| Kev-9B | 0.5 to 0.6 | 102 | 0.55 | 0.47 |
| Kev-9B | 0.6 to 0.7 | 115 | 0.65 | 0.64 |
| Kev-9B | 0.7 to 0.8 | 180 | 0.75 | 0.81 |
| Kev-9B | 0.8 to 0.9 | 440 | 0.86 | 0.92 |
| Kev-9B | 0.9 to 1.0 | 2,132 | 0.96 | 0.99 |
| Kev-4B | 0.2 to 0.3 | 16 | 0.27 | 0.19 |
| Kev-4B | 0.3 to 0.4 | 147 | 0.36 | 0.29 |
| Kev-4B | 0.4 to 0.5 | 264 | 0.45 | 0.48 |
| Kev-4B | 0.5 to 0.6 | 201 | 0.55 | 0.71 |
| Kev-4B | 0.6 to 0.7 | 237 | 0.65 | 0.88 |
| Kev-4B | 0.7 to 0.8 | 369 | 0.75 | 0.95 |
| Kev-4B | 0.8 to 0.9 | 797 | 0.86 | 0.98 |
| Kev-4B | 0.9 to 1.0 | 1,049 | 0.93 | 1.00 |
| Clef-flash | 0.2 to 0.3 | 7 | 0.26 | 0.14 |
| Clef-flash | 0.3 to 0.4 | 12 | 0.35 | 0.33 |
| Clef-flash | 0.4 to 0.5 | 59 | 0.46 | 0.68 |
| Clef-flash | 0.5 to 0.6 | 87 | 0.56 | 0.78 |
| Clef-flash | 0.6 to 0.7 | 241 | 0.66 | 0.87 |
| Clef-flash | 0.7 to 0.8 | 736 | 0.76 | 0.95 |
| Clef-flash | 0.8 to 0.9 | 1,339 | 0.85 | 0.98 |
| Clef-flash | 0.9 to 1.0 | 599 | 0.92 | 1.00 |
How well the confidence matches the accuracy
Calibration error (ECE): the average gap between the confidence a model states and the share of those answers that are right, over ten equal bands. Lower is better. AUROC: how well the confidence ranks right answers above wrong ones; one is perfect, a half is a coin flip.
| confidence used | calibration error as returned | after a fit on held-out messages | AUROC | |
|---|---|---|---|---|
| Kev-27B | the probability of the answer it chose | 0.014 | 0.009 | 0.920 |
| GPT-6 Luna | the number it writes in its answer | 0.025 | 0.021 | 0.876 |
| Clef | the probability of the answer it chose | 0.030 | 0.015 | 0.900 |
| Jev | the probability of the answer it chose | 0.035 | 0.010 | 0.840 |
| Kev-9B | the probability of the answer it chose | 0.038 | 0.012 | 0.906 |
| Kev-4B | the probability of the answer it chose | 0.113 | 0.009 | 0.922 |
| Clef-flash | the probability of the answer it chose | 0.143 | 0.006 | 0.819 |
The workload table puts this in a support team's terms. At a budget of 2% wrong among the answers accepted automatically, Clef handled 93.2% of messages alone and Jev 84.4%. Jev's threshold had to sit at 0.990: most of its answers carry a confidence that high, which leaves a threshold little room to separate the right ones from the wrong. GPT-6 Luna's threshold, fixed on the dev messages, let through 2.99% wrong answers on the test, above the budget.
What each model handles alone, within an error budget
The policy: accept the model's answer when its confidence reaches a threshold, send the message to a person otherwise. The threshold is the lowest that kept errors among accepted answers under 2% on the dev messages, fixed before the test.
| threshold | accepted automatically | wrong among the accepted [95% interval] | sent to a person, per thousand | |
|---|---|---|---|---|
| Clef | 0.803 | 93.2% | 2.40% [1.90, 3.03] | 68 |
| GPT-6 Luna | 0.900 | 91.3% | 2.99% [2.42, 3.68] | 87 |
| Kev-27B | 0.793 | 91.0% | 2.28% [1.79, 2.90] | 90 |
| Jev | 0.990 | 84.4% | 2.04% [1.56, 2.66] | 156 |
| Kev-9B | 0.813 | 82.7% | 2.47% [1.94, 3.15] | 173 |
| Clef-flash | 0.732 | 81.7% | 2.15% [1.65, 2.79] | 183 |
| Kev-4B | 0.669 | 74.3% | 1.75% [1.29, 2.37] | 257 |
Same answer twice, and over time
Asked the same messages again the same day, Kev-27B, Kev-9B, Clef, Clef-flash and Kev-4B gave identical answers every time. Jev gave the same answer to 199 of 200 and GPT-6 Luna to 197.
Over days, Jev kept the same version string but changed a few answers: 3,075 of 3,080 were the same between 21 September and 30 September, and 3,073 between 30 September and 2 October. Its median time also fell over that week.
The same messages again
A second pass over messages from the test on the same day, answers compared with the first pass.
Jev and GPT-6 Luna over time
The same messages, sent again days later.
| right, then | right, later | same answer | median time, then | median time, later | |
|---|---|---|---|---|---|
| Jev, 21 September to 30 September | 94.0% | 93.9% | 3,075 of 3,080 | 313 ms | 253 ms |
| Jev, 30 September to 2 October | 93.9% | 93.9% | 3,073 of 3,080 | 253 ms | 235 ms |
| GPT-6 Luna, 24 September to 30 September | 94.0% | 94.1% | 3,047 of 3,080 | 1,196 ms | 1,097 ms |
Has it seen the test before?
BANKING77 is public, so a model may have trained on it. Kev's model cards list it among their training data. Cloudflare post-trained Clef on an open Qwen3.8 model with its own synthetic data, and doesn't say whether BANKING77 was in that data or in the base model's. We found nothing from TypeSafe on whether Jev saw it.
We asked five of the models, without examples, about 200 messages from the training split and about paraphrases of the same messages; M37's jev-2 asked Jev and GPT-6 Luna the same. A model that had memorised the originals should do much worse on the paraphrases. Every model dropped, but by less than a simple vote of the nearest labelled examples, which can't memorise anything and still drops 15.0 points. So this probe can't tell a little memorisation from the cost of paraphrasing. For Kev it measures something narrower still, since BANKING77 is in its training data.
Training messages against paraphrases of them
Share right without examples on 200 messages from BANKING77's training split, and on paraphrases of the same messages. A model that had memorised the originals should drop by more than the yardstick does.
| originals | paraphrases | drop, points [95% interval] | BANKING77 in its training data? | |
|---|---|---|---|---|
| Kev-27B | 79.5% | 65.5% | +14.0 points [+9.0, +19.0] | yes, its model card lists BANKING77 |
| Clef-flash | 95.5% | 83.5% | +12.0 points [+7.5, +16.5] | not stated; Cloudflare post-trained it on its own synthetic data |
| Kev-9B | 74.0% | 62.5% | +11.5 points [+6.5, +17.0] | yes, its model card lists BANKING77 |
| Kev-4B | 63.5% | 55.5% | +8.0 points [+2.5, +14.0] | not checked for this version |
| Clef | 94.5% | 87.0% | +7.5 points [+4.0, +11.5] | not stated; Cloudflare post-trained it on its own synthetic data |
| GPT-6 Luna (jev-2, 24 September, its own prompt) | 76.5% | 71.5% | +5.0 points [+1.0, +9.5] | not checked |
| Jev (jev-2, 21 September) | 72.5% | 68.0% | +4.5 points [+0.5, +8.5] | not stated by TypeSafe, as far as we found |
| The five nearest labelled examples, voting; it cannot memorise anything | 94.5% | 79.5% | +15.0 points [+10.0, +20.5] | the yardstick |
Would a plain classifier do?
Could a trained classifier do the same job? On this task, with nearly all of the training split, it nearly does. A small embedding model with logistic regression on top, trained on 9,405 labelled messages, got 93.1% right in 9 ms per message on a Mac. Jev got 93.9%.
That comparison favours the classifier, which saw thousands of labelled messages. The live question is how few labelled examples a trained classifier needs before it catches up with a decision model, and that is the next thing we measure.
Plain baselines on the same test
From M37's jev-2 run of the same protocol on the same messages, run between 21 September and 24 September, not re-run since.
| right | 95% interval | median time | per thousand decisions | |
|---|---|---|---|---|
| A trained classifier: a small embedding model with logistic regression on top | 93.1% | [92.2, 94.0] | 9 ms, on the Mac | no API cost |
| The label of the single nearest labelled example | 92.6% | [91.6, 93.5] | 10 ms, on the Mac | no API cost |
| Ling 3.0 Flash, an LLM, through OpenRouter | 93.8% | [92.9, 94.6] | 723 ms | $0.020 |
| DeepSeek V4.1 Flash, through OpenRouter | 93.6% | [92.7, 94.4] | 953 ms | $0.038 |
| Jev, for comparison (2 October) | 93.9% | [93.0, 94.7] | 235 ms | $0.078 |
What we measure next
- How many labelled examples a plain classifier needs to beat Jev and Clef: the same messages, with the classifier trained on a few, then more, examples per intent.
- OpenAI's Decisions API, on the day it opens to us. The code that calls it is already written.
- Latency from other regions, since a Berlin number says little about a caller in Virginia or Singapore.
- Calibration on other kinds of task, including yes-or-no questions.
- Accuracy as the list of options grows.
Each will get its own table here, dated, when it has been measured.
How it was measured
The protocol was written down and committed before any test message was sent to a model, and every model since has run it unchanged. One exception is disclosed in the Kev addendum: before Kev's settings were frozen, 40 test messages went to each Kev model to time the round trip; only the timing was read, and those messages were sent again in the scored pass. Thresholds and the recalibration fits were set on 573 held-out training messages, never on the test. Every model call is kept, with a hash of its request, the full response and its timing.
Accuracy intervals are Wilson intervals. Differences between two models are paired, message by message, with a bootstrap over the messages. The calibration error uses ten equal bands of stated confidence. Costs are the provider's bill where it gives one per call, as OpenRouter does, and the rate card times the tokens it reports where it doesn't. The Kev rows are Modal's bill for the GPU time of each pass.
What this doesn't cover yet: other tasks, other places, many requests at once on the hosted models, images, and long documents. The protocol comes from M37's jev-2 test, which has a video of its own. The per-message file at the foot of this page has every model's answer, confidence and time for each message, so the accuracy, the differences between models and the calibration as returned can all be recomputed from it.