model fatıgue
First look

Ember-1 thinks less than Kimi K3. But at what cost?

Published 30 Sep 2026

Pressing play loads the video from YouTube (Google), in privacy-enhanced mode. Privacy

This is the video written out, with every figure in full, each linked to its source in the table below.

Fireworks' new model▶ 0:00

On 23 September, Fireworks released Ember-1. It's Kimi K3, the open-weight model from Moonshot, trained further so that it spends less time thinking before it answers. Fireworks' launch post says it "delivers Kimi K3's quality with 40% fewer tokens", and Fireworks charges the same per token for both: $3.00 per million input tokens, $0.30 per million cached input tokens and $15.00 per million output tokens. So every token Ember-1 doesn't write is money you don't spend.

So what does it give up? The saving looks real. Every source we found that reports cost or tokens has Ember-1 lower than K3. What you give up in quality depends on the kind of work you give it.

The basics▶ 0:40

Kimi K3 is a reasoning model. Before it answers, it writes out its thinking, and you pay for that thinking as output tokens, at $15.00 per million on Fireworks. Moonshot gives it an effort setting with three levels, low, high and max, and max is the default. Artificial Analysis, the independent benchmarking site, counts 48,455 output tokens per task for K3 at max on its Intelligence Index, 32,453 of them thinking and 16,002 the answer.

Fireworks says turning the effort down wasn't the answer: "Lower effort settings gave up too much quality." So it trained the model to think more briefly instead, and that's Ember-1. It keeps an effort setting of its own.

Ember-1 is a research preview. None of Fireworks' pages we saved (the launch post, the model page, the docs) links to weights, a paper or a system card, a Hugging Face search found no Ember-1 weights, and OpenRouter lists no provider other than Fireworks. Fireworks says it gives research models "two-week serverless access", and makes them permanent based on community demand. So everything here comes from Fireworks' own pages and from the few people outside Fireworks who have tested it.

Where the 40% comes from▶ 1:38

The launch post's headline says 40% fewer tokens. The same post puts the saving at "approximately 35% fewer tokens per task" in its section on customer tests, says reasoning traces could be shortened by 35–50% in its section on method, and calls one section "half the tokens, same answers".

The figure in it closest to the headline is a 39% cut in total tokens, from live tests on two customers' production coding traffic, where Ember-1 scored about the same as K3.

Fireworks' customer tests

Live A/B tests on two customers' production coding traffic, as the launch post reports them

scoreoutput tokens
Kimi K30.75149.3K
Ember-10.75329.9K
Fireworks' launch post, read 29 September 2026. The post puts the total-token reduction at 39% and the reasoning-token reduction at 71.3%. It doesn't name the customers or the tasks, or say what the score measures.

Those customer tests are Fireworks' own, and nobody outside Fireworks can check them. The post doesn't name the customers or the tasks, or say what the score measures.

Fireworks' benchmark table▶ 2:11

Further down, the same post has a table of public benchmarks: four coding and terminal tests and one airline customer-service test, with K3 at all three effort settings next to Ember-1. Against K3 at max, Ember-1's cost saving runs from 5.9% on the airline test to 51.9% on Terminal Bench, a set of tasks done in a real terminal. None of the five is 40%, and together they average 25.9%. At the same price per token, those savings come from using fewer tokens.

What Ember-1 saves on each benchmark

Cost saving against Kimi K3 at max, as Fireworks' launch post reports it

0%10%20%30%40%50%60%Terminal Bench 2.1Terminal Bench 2.1: 51.9% · Fireworks, read 29 Sep 2026, 11:10 CEST51.9%SWE-bench VerifiedSWE-bench Verified: 15.5% · Fireworks, read 29 Sep 2026, 11:10 CEST15.5%SWE-InteractSWE-Interact: 32.5% · Fireworks, read 29 Sep 2026, 11:10 CEST32.5%DeepSWE 1.1DeepSWE 1.1: 23.7% · Fireworks, read 29 Sep 2026, 11:10 CEST23.7%τ-2 Bench Airlineτ-2 Bench Airline: 5.9% · Fireworks, read 29 Sep 2026, 11:10 CEST5.9%the headline, 40%mean of the five, 25.9%
0%10%20%30%40%50%60%Terminal Bench 2.1Terminal Bench 2.1: 51.9% · Fireworks, read 29 Sep 2026, 11:10 CEST51.9%SWE-bench VerifiedSWE-bench Verified: 15.5% · Fireworks, read 29 Sep 2026, 11:10 CEST15.5%SWE-InteractSWE-Interact: 32.5% · Fireworks, read 29 Sep 2026, 11:10 CEST32.5%DeepSWE 1.1DeepSWE 1.1: 23.7% · Fireworks, read 29 Sep 2026, 11:10 CEST23.7%τ-2 Bench Airlineτ-2 Bench Airline: 5.9% · Fireworks, read 29 Sep 2026, 11:10 CEST5.9%the headline, 40%mean of the five, 25.9%
Fireworks' launch post, read 29 September 2026. The mean is our arithmetic. The post gives each saving with its dollar amount; at the same price per token, the saving comes from fewer tokens.
The numbers in this chart
cost saving
Terminal Bench 2.151.9%
SWE-bench Verified15.5%
SWE-Interact32.5%
DeepSWE 1.123.7%
τ-2 Bench Airline5.9%

The scores barely move. Ember-1's pass rate against K3 at max is +1.1 points on Terminal Bench, +2.0 points on the airline test, and −1.0 points and −1.3 points on SWE-bench Verified and SWE-Interact. The exception is DeepSWE, a set of deep debugging tasks, where Ember-1 passes 75.2% against K3's 66.4%, a lead of 8.8 points. Keep that one in mind.

The scores barely move, except on DeepSWE

Pass rate on each benchmark, Kimi K3 at max against Ember-1, as Fireworks' launch post reports it

Kimi K3 at maxEmber-1
pass rate, %difference (our arithmetic)Terminal Bench 2.1Terminal Bench 2.1, Kimi K3 at max: 80.9% · Fireworks, read 29 Sep 2026, 11:10 CEST80.9%Terminal Bench 2.1, Ember-1: 82.0% · Fireworks, read 29 Sep 2026, 11:10 CEST82.0%+1.1SWE-bench VerifiedSWE-bench Verified, Kimi K3 at max: 93.2% · Fireworks, read 29 Sep 2026, 11:10 CEST93.2%SWE-bench Verified, Ember-1: 92.2% · Fireworks, read 29 Sep 2026, 11:10 CEST92.2%−1.0SWE-InteractSWE-Interact, Kimi K3 at max: 21.3% · Fireworks, read 29 Sep 2026, 11:10 CEST21.3%SWE-Interact, Ember-1: 20.0% · Fireworks, read 29 Sep 2026, 11:10 CEST20.0%−1.3DeepSWE 1.1DeepSWE 1.1, Kimi K3 at max: 66.4% · Fireworks, read 29 Sep 2026, 11:10 CEST66.4%DeepSWE 1.1, Ember-1: 75.2% · Fireworks, read 29 Sep 2026, 11:10 CEST75.2%+8.8τ-2 Bench Airlineτ-2 Bench Airline, Kimi K3 at max: 64% · Fireworks, read 29 Sep 2026, 11:10 CEST64%τ-2 Bench Airline, Ember-1: 66% · Fireworks, read 29 Sep 2026, 11:10 CEST66%+2.0
pass rate, %difference (our arithmetic)Terminal Bench 2.1Terminal Bench 2.1, Kimi K3 at max: 80.9% · Fireworks, read 29 Sep 2026, 11:10 CEST80.9%Terminal Bench 2.1, Ember-1: 82.0% · Fireworks, read 29 Sep 2026, 11:10 CEST82.0%+1.1SWE-bench VerifiedSWE-bench Verified, Kimi K3 at max: 93.2% · Fireworks, read 29 Sep 2026, 11:10 CEST93.2%SWE-bench Verified, Ember-1: 92.2% · Fireworks, read 29 Sep 2026, 11:10 CEST92.2%−1.0SWE-InteractSWE-Interact, Kimi K3 at max: 21.3% · Fireworks, read 29 Sep 2026, 11:10 CEST21.3%SWE-Interact, Ember-1: 20.0% · Fireworks, read 29 Sep 2026, 11:10 CEST20.0%−1.3DeepSWE 1.1DeepSWE 1.1, Kimi K3 at max: 66.4% · Fireworks, read 29 Sep 2026, 11:10 CEST66.4%DeepSWE 1.1, Ember-1: 75.2% · Fireworks, read 29 Sep 2026, 11:10 CEST75.2%+8.8τ-2 Bench Airlineτ-2 Bench Airline, Kimi K3 at max: 64% · Fireworks, read 29 Sep 2026, 11:10 CEST64%τ-2 Bench Airline, Ember-1: 66% · Fireworks, read 29 Sep 2026, 11:10 CEST66%+2.0
Fireworks' launch post, read 29 September 2026. The full table, with Kimi K3 at low and high as well, is below.
The numbers in this chart
Kimi K3 at maxEmber-1difference (our arithmetic)
Terminal Bench 2.180.9%82.0%+1.1 points
SWE-bench Verified93.2%92.2%−1.0 points
SWE-Interact21.3%20.0%−1.3 points
DeepSWE 1.166.4%75.2%+8.8 points
τ-2 Bench Airline64%66%+2.0 points

Fireworks' benchmark table

Pass rates for Kimi K3 at each effort setting and for Ember-1, and Ember-1's cost saving against Kimi K3 at max, from Fireworks' launch post

benchmarktasksK3 lowK3 highK3 maxEmber-1saving vs K3 max
Terminal Bench 2.18976.4%77.6%80.9%82.0%51.9%, $23.1
SWE-bench Verified50080.4%86.0%93.2%92.2%15.5%, $68.1
SWE-Interact756.7%13.3%21.3%20.0%32.5%, $60.8
DeepSWE 1.111355.8%62.8%66.4%75.2%23.7%, $126.9
τ-2 Bench Airline5064%64%64%66%5.9%, $0.3
Fireworks' launch post, read 29 September 2026. The post gives each saving as a share and in dollars over the whole benchmark.

Fireworks' own leaderboard▶ 3:09

The same week, Fireworks launched a leaderboard of its own, the Specialized Intelligence Index, built from benchmarks it says were "contributed by industry practitioners". It has boards for different kinds of work, and four of them list both Ember-1 and K3. The index doesn't say which effort setting either model ran at.

On all four boards, Ember-1 costs less than K3, between 11% and 35% less (our arithmetic on the index's costs). And on all four it scores lower. On Bedside Bench, clinical cases written by physicians, it's 0.6 points lower; on Big Finance, 2.3 points; and on DeepSWE, 3.25 points.

Lower on all four of Fireworks' own boards

Score on each board of the Specialized Intelligence Index that lists both models

Kimi K3Ember-1
scoredifference (our arithmetic)Bedside BenchBedside Bench, Kimi K3: 89.7 · Fireworks, read 29 Sep 2026, 11:09 CEST89.7Bedside Bench, Ember-1: 89.1 · Fireworks, read 29 Sep 2026, 11:09 CEST89.1−0.6Big FinanceBig Finance, Kimi K3: 41.3 · Fireworks, read 29 Sep 2026, 11:09 CEST41.3Big Finance, Ember-1: 39.0 · Fireworks, read 29 Sep 2026, 11:09 CEST39.0−2.3DeepSWE v1.1DeepSWE v1.1, Kimi K3: 70.21 · Fireworks, read 29 Sep 2026, 11:09 CEST70.21DeepSWE v1.1, Ember-1: 66.96 · Fireworks, read 29 Sep 2026, 11:09 CEST66.96−3.25τ³-Bankingτ³-Banking, Kimi K3: 51.5 · Fireworks, read 29 Sep 2026, 11:09 CEST51.5τ³-Banking, Ember-1: 38.1 · Fireworks, read 29 Sep 2026, 11:09 CEST38.1−13.4
scoredifference (our arithmetic)Bedside BenchBedside Bench, Kimi K3: 89.7 · Fireworks, read 29 Sep 2026, 11:09 CEST89.7Bedside Bench, Ember-1: 89.1 · Fireworks, read 29 Sep 2026, 11:09 CEST89.1−0.6Big FinanceBig Finance, Kimi K3: 41.3 · Fireworks, read 29 Sep 2026, 11:09 CEST41.3Big Finance, Ember-1: 39.0 · Fireworks, read 29 Sep 2026, 11:09 CEST39.0−2.3DeepSWE v1.1DeepSWE v1.1, Kimi K3: 70.21 · Fireworks, read 29 Sep 2026, 11:09 CEST70.21DeepSWE v1.1, Ember-1: 66.96 · Fireworks, read 29 Sep 2026, 11:09 CEST66.96−3.25τ³-Bankingτ³-Banking, Kimi K3: 51.5 · Fireworks, read 29 Sep 2026, 11:09 CEST51.5τ³-Banking, Ember-1: 38.1 · Fireworks, read 29 Sep 2026, 11:09 CEST38.1−13.4
Fireworks' Specialized Intelligence Index, read 29 September 2026. The index doesn't say which effort setting either model ran at.
The numbers in this chart
Kimi K3Ember-1difference (our arithmetic)
Bedside Bench89.789.1−0.6 points
Big Finance41.339.0−2.3 points
DeepSWE v1.170.2166.96−3.25 points
τ³-Banking51.538.1−13.4 points

That's the same DeepSWE where the launch post had Ember-1 8.8 points ahead. Here it's 3.25 points behind. Both models' DeepSWE scores differ between the two: Ember-1's from 75.2% in the launch post to 66.96 on the leaderboard, K3's from 66.4% to 70.21, and neither says why.

The fourth board is where Ember-1 drops furthest. τ³-Banking tests banking customer-service agents against a simulated customer. K3 scores 51.5 there and Ember-1 38.1.

The four boards that list both models

Score and cost on Fireworks' Specialized Intelligence Index

boardKimi K3Ember-1cost cut (our arithmetic)
Bedside Bench (medicine)89.7 at $0.052989.1 at $0.040623%
Big Finance41.3 at $0.22039.0 at $0.19511%
DeepSWE v1.1 (coding)70.21 at $5.4066.96 at $4.1922%
τ³-Banking (customer service)51.5 at $1.0238.1 at $0.6635%
Costs are per task, and per trial on τ³-Banking, as the index lists them. The index doesn't say which effort setting either model ran at. Read 29 September 2026.

These boards list other models too. On three of the four, a model that costs less than Ember-1 scores higher: OpenAI's GPT-6 Sol on Big Finance and on DeepSWE, and GLM-5.3 on the banking board, among others. On Bedside Bench, no model with a listed cost does better for less.

Models that score higher than Ember-1 for less

On the same boards, every model with a listed cost below Ember-1's and a score above it

boardmodelscore at cost
Big FinanceEmber-139.0 at $0.195
GPT-6 Sol47.1 at $0.085
DeepSeek V4.1 Flash40.6 at $0.107
DeepSeek V4 Pro 081339.3 at $0.141
DeepSWE v1.1Ember-166.96 at $4.19
GPT-6 Astra71.0 at $2.56
GPT-6 Sol67.3 at $2.31
τ³-BankingEmber-138.1 at $0.66
GLM-5.351.6 at $0.49
Gemini 3.8 Flash51.4 at $0.53
DeepSeek V4 Pro 081342.7 at $0.32
Fireworks' Specialized Intelligence Index, read 29 September 2026. On Bedside Bench no model with a listed cost scores higher than Ember-1 for less. Fireworks' changelog says DeepSeek V4 Pro 0813 left its serverless service on 26 September, and that DeepSeek V4.1 Flash's serverless price changes on 1 October.

Tests outside Fireworks▶ 4:41

Outside Fireworks, there's much less. Artificial Analysis, which published numbers on Sonnet 5.5 within about half an hour of its launch, still hadn't listed Ember-1 when we checked on 29 September, six days in. It has Kimi K3 at max at 43.6 on its Intelligence Index.

A small public test suite, AI BENCHY, did run it. Ember-1 at max fully passed 10 of its 22 tests, against 16 for K3 at max. But 8 of Ember-1's failed tests at max were API errors, requests that came back with an error instead of an answer, and it got API errors at every effort setting AI BENCHY tried. AI BENCHY doesn't say what caused them, but every route to Ember-1 we found runs on Fireworks, so they're worth knowing about. Set the API errors aside and Ember-1 passed 10 of the remaining 14, K3 16 of 20, which on a suite this small is too close to tell apart.

AI BENCHY: tests fully passed

A small public suite of tests, run on each model and effort setting

model and settingpassedAPI errors
Kimi K3 at max16 of 222
Ember-1 at max10 of 228
Ember-1 at high12 of 227
Ember-1 at low13 of 226
Ember-1, reasoning off8 of 225
AI BENCHY's model pages, read 29 September 2026. It tested Ember-1 on 28 September and Kimi K3 on 14 August, and has Kimi K3 only at max. An API error is a request that came back with an error instead of an answer. Ember-1 at max failed its other tests with 3 wrong answers and 1 with no answer.

Another outside run is the Bro Frontier Index, from Aquiles, a developer who is building an AI app called Bro. It covers eight categories of practical work, and we couldn't find its tasks published. With both models at max effort, but in different agent tools (OpenCode for Ember-1, Kimi Work for K3), his charts show Ember-1 writing 25K output tokens per task against K3's 47K and costing $1.22 a task against $2.58, while scoring 30.88 overall against K3's 31.91. Five of the eight categories stayed within about a point, coding among them. The other three dropped: law by 3.54 points, healthcare by 4.42 points and cybersecurity by 5.81 points. His data file also gives Ember-1 the same input tokens per task as K3 and names Moonshot as its API provider, though every route to Ember-1 we found runs on Fireworks, so some of the Ember-1 row may have been copied from K3's. It's one run, and it's the only source on law and security.

Bro Frontier Index, by category

Both models at max effort, in different agent tools, redrawn from Aquiles' data file

Kimi K3 at maxEmber-1 at max
scoredifference (our arithmetic)IntelligenceIntelligence, Kimi K3 at max: 14.94 · Aquiles (Bro Frontier Index), read 29 Sep 2026, 11:10 CEST14.94Intelligence, Ember-1 at max: 15.12 · Aquiles (Bro Frontier Index), read 29 Sep 2026, 11:10 CEST15.12+0.18CodingCoding, Kimi K3 at max: 23.48 · Aquiles (Bro Frontier Index), read 29 Sep 2026, 11:10 CEST23.48Coding, Ember-1 at max: 24.30 · Aquiles (Bro Frontier Index), read 29 Sep 2026, 11:10 CEST24.30+0.82AutomationAutomation, Kimi K3 at max: 54.30 · Aquiles (Bro Frontier Index), read 29 Sep 2026, 11:10 CEST54.30Automation, Ember-1 at max: 53.08 · Aquiles (Bro Frontier Index), read 29 Sep 2026, 11:10 CEST53.08−1.22Knowledge WorkKnowledge Work, Kimi K3 at max: 39.30 · Aquiles (Bro Frontier Index), read 29 Sep 2026, 11:10 CEST39.30Knowledge Work, Ember-1 at max: 39.04 · Aquiles (Bro Frontier Index), read 29 Sep 2026, 11:10 CEST39.04−0.26FinanceFinance, Kimi K3 at max: 55.99 · Aquiles (Bro Frontier Index), read 29 Sep 2026, 11:10 CEST55.99Finance, Ember-1 at max: 55.18 · Aquiles (Bro Frontier Index), read 29 Sep 2026, 11:10 CEST55.18−0.81LegalLegal, Kimi K3 at max: 23.74 · Aquiles (Bro Frontier Index), read 29 Sep 2026, 11:10 CEST23.74Legal, Ember-1 at max: 20.20 · Aquiles (Bro Frontier Index), read 29 Sep 2026, 11:10 CEST20.20−3.54HealthcareHealthcare, Kimi K3 at max: 42.23 · Aquiles (Bro Frontier Index), read 29 Sep 2026, 11:10 CEST42.23Healthcare, Ember-1 at max: 37.81 · Aquiles (Bro Frontier Index), read 29 Sep 2026, 11:10 CEST37.81−4.42CybersecurityCybersecurity, Kimi K3 at max: 30.25 · Aquiles (Bro Frontier Index), read 29 Sep 2026, 11:10 CEST30.25Cybersecurity, Ember-1 at max: 24.44 · Aquiles (Bro Frontier Index), read 29 Sep 2026, 11:10 CEST24.44−5.81
scoredifference (our arithmetic)IntelligenceIntelligence, Kimi K3 at max: 14.94 · Aquiles (Bro Frontier Index), read 29 Sep 2026, 11:10 CEST14.94Intelligence, Ember-1 at max: 15.12 · Aquiles (Bro Frontier Index), read 29 Sep 2026, 11:10 CEST15.12+0.18CodingCoding, Kimi K3 at max: 23.48 · Aquiles (Bro Frontier Index), read 29 Sep 2026, 11:10 CEST23.48Coding, Ember-1 at max: 24.30 · Aquiles (Bro Frontier Index), read 29 Sep 2026, 11:10 CEST24.30+0.82AutomationAutomation, Kimi K3 at max: 54.30 · Aquiles (Bro Frontier Index), read 29 Sep 2026, 11:10 CEST54.30Automation, Ember-1 at max: 53.08 · Aquiles (Bro Frontier Index), read 29 Sep 2026, 11:10 CEST53.08−1.22Knowledge WorkKnowledge Work, Kimi K3 at max: 39.30 · Aquiles (Bro Frontier Index), read 29 Sep 2026, 11:10 CEST39.30Knowledge Work, Ember-1 at max: 39.04 · Aquiles (Bro Frontier Index), read 29 Sep 2026, 11:10 CEST39.04−0.26FinanceFinance, Kimi K3 at max: 55.99 · Aquiles (Bro Frontier Index), read 29 Sep 2026, 11:10 CEST55.99Finance, Ember-1 at max: 55.18 · Aquiles (Bro Frontier Index), read 29 Sep 2026, 11:10 CEST55.18−0.81LegalLegal, Kimi K3 at max: 23.74 · Aquiles (Bro Frontier Index), read 29 Sep 2026, 11:10 CEST23.74Legal, Ember-1 at max: 20.20 · Aquiles (Bro Frontier Index), read 29 Sep 2026, 11:10 CEST20.20−3.54HealthcareHealthcare, Kimi K3 at max: 42.23 · Aquiles (Bro Frontier Index), read 29 Sep 2026, 11:10 CEST42.23Healthcare, Ember-1 at max: 37.81 · Aquiles (Bro Frontier Index), read 29 Sep 2026, 11:10 CEST37.81−4.42CybersecurityCybersecurity, Kimi K3 at max: 30.25 · Aquiles (Bro Frontier Index), read 29 Sep 2026, 11:10 CEST30.25Cybersecurity, Ember-1 at max: 24.44 · Aquiles (Bro Frontier Index), read 29 Sep 2026, 11:10 CEST24.44−5.81
Bro Frontier Index by Aquiles (@aquilesfd), bro-evals.vercel.app, redrawn by Model Fatigue with his credit. His data file, read 29 September 2026, runs Ember-1 in OpenCode and Kimi K3 in Kimi Work, and lists the same input tokens per task and the same API provider (Moonshot) for both.
The numbers in this chart
Kimi K3 at maxEmber-1 at maxdifference (our arithmetic)
Intelligence14.9415.12+0.18 points
Coding23.4824.30+0.82 points
Automation54.3053.08−1.22 points
Knowledge Work39.3039.04−0.26 points
Finance55.9955.18−0.81 points
Legal23.7420.20−3.54 points
Healthcare42.2337.81−4.42 points
Cybersecurity30.2524.44−5.81 points

On 28 September, Leo Linsky posted results for Ember-1 from Gert Labs' benchmark, GBENCH. GBENCH has models write programs that play games against each other. When the model writes its program in a single request, Ember-1 and K3 come out level, 0.556 against 0.551, with Gert Labs' published intervals overlapping. In what Gert Labs calls agentic coding, where the model works on its program over many requests (for Ember-1, a median of 60 per game), Ember-1 scores 0.518 against K3's 0.659, and the intervals don't overlap. Overall, GBENCH ranks K3 #15 and Ember-1 #26 of 100 models, though Ember-1 hasn't been run on its decision-making games, which K3 has.

GBENCH: level in one request, lower over many

Gert Labs' score for code that plays games, with the interval Gert Labs publishes for each score

Kimi K3Ember-1
GBENCH scoreOne requestOne request, Kimi K3: 0.551 · Gert Labs, read 29 Sep 2026, 11:14 CEST · interval 0.521 to 0.5800.551One request, Ember-1: 0.556 · Gert Labs, read 29 Sep 2026, 11:14 CEST · interval 0.529 to 0.5820.556Agentic, many requestsAgentic, many requests, Kimi K3: 0.659 · Gert Labs, read 29 Sep 2026, 11:14 CEST · interval 0.615 to 0.6990.659Agentic, many requests, Ember-1: 0.518 · Gert Labs, read 29 Sep 2026, 11:14 CEST · interval 0.465 to 0.5660.518
GBENCH scoreOne requestOne request, Kimi K3: 0.551 · Gert Labs, read 29 Sep 2026, 11:14 CEST · interval 0.521 to 0.5800.551One request, Ember-1: 0.556 · Gert Labs, read 29 Sep 2026, 11:14 CEST · interval 0.529 to 0.5820.556Agentic, many requestsAgentic, many requests, Kimi K3: 0.659 · Gert Labs, read 29 Sep 2026, 11:14 CEST · interval 0.615 to 0.6990.659Agentic, many requests, Ember-1: 0.518 · Gert Labs, read 29 Sep 2026, 11:14 CEST · interval 0.465 to 0.5660.518
Gert Labs' rankings data (gertlabs.com/rankings), read 29 September 2026. The whiskers are Gert Labs' published intervals. Its data doesn't say which effort setting either model ran at.
The numbers in this chart
Kimi K3Ember-1
One request0.551 (0.521 to 0.580)0.556 (0.529 to 0.582)
Agentic, many requests0.659 (0.615 to 0.699)0.518 (0.465 to 0.566)

Each of these is one outside test, set up its own way, and they don't always agree with Fireworks: its medical board showed a gap of 0.6 points, where the Bro index's healthcare category showed 4.42 points. So the drops on law, health and security, and on long coding tasks, are worth checking before you rely on Ember-1 for them.

Who's talking about it▶ 6:47

Most of what's been said about Ember-1 has been in posts on X and Hacker News, and most of the engagement we found came on 27 September, four days after launch, when it reached Hacker News. On X, three of the five most-liked posts came from Fireworks' side: the company's own account, its co-founder and CTO Dmytro Dzhulgakov, and Sonya Huang, who works at Sequoia, one of Fireworks' investors. The other two, from the coding agent Cline and from Elvis Saravia, repeated Fireworks' numbers. None of the five reports a test.

Four people posted what they found with some detail: Aquiles, with his index; Leo Linsky, with the Gert Labs results; a Hacker News commenter whose model router, a tool that picks a model for each request, gave Ember-1 full marks on coding and still ranked GPT-6 Sol above it, partly because the router had no benchmark score for Ember-1 to start from; and Vojtech Rinik, who found it slow on one long voice memo. One more post on X said a single reasoning question ran on about forty percent fewer tokens than on K3 with no loss in quality, and gave no data. And on Hacker News, someone asked the question the video is about: "What, if any, capability is lost by the token reduction?"

So what does it lose?▶ 7:57

On coding, most scores are within a point or two either way, but one outside benchmark found a clear drop when Ember-1 works on code over many requests. On customer service, Fireworks' own banking board shows a drop from 51.5 to 38.1, while the launch post's airline test went 2.0 points the other way. On finance, Fireworks' board has it 2.3 points lower and the Bro index about level. On law and security, only the Bro index has tested it, in its one run, and it's lower on both. On medicine, both sources have it lower: Fireworks' medical board by 0.6 points, the Bro index by 4.42 points.

What does it lose?

Ember-1's score minus Kimi K3's (our arithmetic), in points on each source's own scale, by kind of work and by whose number it is

kind of workFireworks' launch postFireworks' indexBro Frontier IndexGBENCH
CodingTerminal Bench +1.1, SWE-bench Verified −1.0, SWE-Interact −1.3, DeepSWE +8.8DeepSWE −3.25+0.82one request: level; many requests: lower
Customer serviceairline +2.0banking −13.4
Finance−2.3−0.81
Medicine−0.6−4.42
Law−3.54
Security−5.81
GBENCH scores run from zero to one, so its results are given in words: level where Gert Labs' published intervals overlap, lower where they don't. The Bro Frontier Index column is one run, with the two models in different agent tools. An empty cell means that source has no test of that kind. Read 29 September 2026.

What nobody has measured▶ 8:34

Two comparisons would settle most of what's left, and we found no one outside Fireworks who has run either.

The first is Ember-1 against K3 simply set one step lower, at high effort, the option Fireworks says doesn't work. Fireworks' own numbers give a first hint. Averaged over the five benchmarks in its table (our arithmetic), K3 at high passes 60.7%, K3 at max 65.2% and Ember-1 67.1%, and the post's own chart of a five-benchmark average against cost per task (its averages differ a little from ours) has K3 at high a little cheaper than at max and Ember-1 cheaper than both. Ember-1's lead over K3 at max comes almost entirely from DeepSWE: it's +1.9 points across all five benchmarks and +0.2 points across the other four. If Fireworks' result holds on other people's tasks, training the model to think less did better than simply turning K3 down. AI BENCHY, the one outside suite we found that ran Ember-1 at more than one setting, has K3 only at max.

The second is the other claim in Fireworks' section title, "same answers". A score can stay the same while the answers change, and nobody outside Fireworks has checked whether Ember-1 gives the answers K3 would, on your kind of work.

Should you try it this week?▶ 9:31

Until someone measures those, here's how we'd act on what's public. If you run K3 on short coding tasks, Ember-1 is worth trying on some of your own traffic. It costs the same per token, so compare output tokens and pass rates on the same tasks. If your work is long agentic coding, customer-service conversations, or legal, medical or security questions, check more carefully, because that's where the drops so far showed up. And if you aren't tied to K3, cheaper models score higher than Ember-1 on three of Fireworks' four boards.

Fireworks hasn't published an end date for the preview. Two weeks from launch would be 7 October, but that's our arithmetic, not a date Fireworks has given. Its docs promise at least two weeks' advance notice before removing a model, and when we checked on 29 September there was no notice, though the docs don't say whether that covers research releases.

Five days after Ember-1, Anthropic made much the same promise for Sonnet 5.5, fewer tokens for the same work. That's our previous video.

Every number

These are all 163 figures behind the video and this page, grouped by whose they are, with the page each came from and when we read it. Figures marked ⟳ can move. When a re-read finds a change, the new value shows next to the one from the video.

Fireworks

price per million input tokens, Ember-1 and Kimi K3 alike$3.00
price per million cached input tokens, Ember-1 and Kimi K3 alike$0.30
price per million output tokens, Ember-1 and Kimi K3 alike$15.00

Fireworks

Fireworks, Introducing Ember-1 (launch post) · read 29 Sep 2026, 11:10 CEST
the launch post's headline: fewer tokens than Kimi K340%
the launch post's customer section: fewer tokens per task35%
the launch post's method section: how much reasoning was shortened35–50%
customer A/B tests, score, Kimi K3 (Fireworks does not define the score)0.751
customer A/B tests, score, Ember-10.753
customer A/B tests, output tokens, Kimi K349.3K
customer A/B tests, output tokens, Ember-129.9K
customer A/B tests, reasoning-token reduction, Ember-1 against Kimi K371.3%
customer A/B tests, total-token reduction, Ember-1 against Kimi K339%
Terminal Bench 2.1: tasks in the benchmark89
Terminal Bench 2.1: Kimi K3 at low, pass rate76.4%
Terminal Bench 2.1: Kimi K3 at high, pass rate77.6%
Terminal Bench 2.1: Kimi K3 at max, pass rate80.9%
Terminal Bench 2.1: Ember-1, pass rate82.0%
Terminal Bench 2.1: Ember-1's cost saving against Kimi K3 at max51.9%
Terminal Bench 2.1: Ember-1's cost saving against Kimi K3 at max, in dollars over the benchmark$23.1
SWE-bench Verified: tasks in the benchmark500
SWE-bench Verified: Kimi K3 at low, pass rate80.4%
SWE-bench Verified: Kimi K3 at high, pass rate86.0%
SWE-bench Verified: Kimi K3 at max, pass rate93.2%
SWE-bench Verified: Ember-1, pass rate92.2%
SWE-bench Verified: Ember-1's cost saving against Kimi K3 at max15.5%
SWE-bench Verified: Ember-1's cost saving against Kimi K3 at max, in dollars over the benchmark$68.1
SWE-Interact: tasks in the benchmark75
SWE-Interact: Kimi K3 at low, pass rate6.7%
SWE-Interact: Kimi K3 at high, pass rate13.3%
SWE-Interact: Kimi K3 at max, pass rate21.3%
SWE-Interact: Ember-1, pass rate20.0%
SWE-Interact: Ember-1's cost saving against Kimi K3 at max32.5%
SWE-Interact: Ember-1's cost saving against Kimi K3 at max, in dollars over the benchmark$60.8
DeepSWE 1.1: tasks in the benchmark113
DeepSWE 1.1: Kimi K3 at low, pass rate55.8%
DeepSWE 1.1: Kimi K3 at high, pass rate62.8%
DeepSWE 1.1: Kimi K3 at max, pass rate66.4%
DeepSWE 1.1: Ember-1, pass rate75.2%
DeepSWE 1.1: Ember-1's cost saving against Kimi K3 at max23.7%
DeepSWE 1.1: Ember-1's cost saving against Kimi K3 at max, in dollars over the benchmark$126.9
τ-2 Bench Airline: tasks in the benchmark50
τ-2 Bench Airline: Kimi K3 at low, pass rate64%
τ-2 Bench Airline: Kimi K3 at high, pass rate64%
τ-2 Bench Airline: Kimi K3 at max, pass rate64%
τ-2 Bench Airline: Ember-1, pass rate66%
τ-2 Bench Airline: Ember-1's cost saving against Kimi K3 at max5.9%
τ-2 Bench Airline: Ember-1's cost saving against Kimi K3 at max, in dollars over the benchmark$0.3
Terminal Bench 2.1: Ember-1 minus Kimi K3 at max, percentage points (our arithmetic)+1.1 points
SWE-bench Verified: Ember-1 minus Kimi K3 at max, percentage points (our arithmetic)−1.0 points
SWE-Interact: Ember-1 minus Kimi K3 at max, percentage points (our arithmetic)−1.3 points
DeepSWE 1.1: Ember-1 minus Kimi K3 at max, percentage points (our arithmetic)+8.8 points
τ-2 Bench Airline: Ember-1 minus Kimi K3 at max, percentage points (our arithmetic)+2.0 points
mean of the five cost savings (our arithmetic)25.9%
mean pass rate of the five, Kimi K3 at high (our arithmetic)60.7%
mean pass rate of the five, Kimi K3 at max (our arithmetic)65.2%
mean pass rate of the five, Ember-1 (our arithmetic)67.1%
Ember-1's lead over Kimi K3 at max, mean of the five (our arithmetic)+1.9 points
Ember-1's lead over Kimi K3 at max, mean of the four without DeepSWE (our arithmetic)+0.2 points

Artificial Analysis

Intelligence Index v4.3, Kimi K3 at max43.6
Kimi K3 at max, output tokens per task on the Intelligence Index48,455
Kimi K3 at max, of which thinking on the Intelligence Index32,453
Kimi K3 at max, of which answer on the Intelligence Index16,002

Fireworks

Bedside Bench: score, Kimi K389.7
Bedside Bench: cost per task, Kimi K3$0.0529
Bedside Bench: score, Ember-189.1
Bedside Bench: cost per task, Ember-1$0.0406
Big Finance: score, Kimi K341.3
Big Finance: cost per task, Kimi K3$0.220
Big Finance: score, Ember-139.0
Big Finance: cost per task, Ember-1$0.195
Big Finance: score, GPT-6 Sol47.1
Big Finance: cost per task, GPT-6 Sol$0.085
Big Finance: score, DeepSeek V4.1 Flash40.6
Big Finance: cost per task, DeepSeek V4.1 Flash$0.107
Big Finance: score, DeepSeek V4 Pro 081339.3
Big Finance: cost per task, DeepSeek V4 Pro 0813$0.141
DeepSWE v1.1: score, Kimi K370.21
DeepSWE v1.1: cost per task, Kimi K3$5.40
DeepSWE v1.1: score, Ember-166.96
DeepSWE v1.1: cost per task, Ember-1$4.19
DeepSWE v1.1: score, GPT-6 Astra71.0
DeepSWE v1.1: cost per task, GPT-6 Astra$2.56
DeepSWE v1.1: score, GPT-6 Sol67.3
DeepSWE v1.1: cost per task, GPT-6 Sol$2.31
τ³-Banking: score, Kimi K351.5
τ³-Banking: cost per trial, Kimi K3$1.02
τ³-Banking: score, Ember-138.1
τ³-Banking: cost per trial, Ember-1$0.66
τ³-Banking: score, GLM-5.351.6
τ³-Banking: cost per trial, GLM-5.3$0.49
τ³-Banking: score, Gemini 3.8 Flash51.4
τ³-Banking: cost per trial, Gemini 3.8 Flash$0.53
τ³-Banking: score, DeepSeek V4 Pro 081342.7
τ³-Banking: cost per trial, DeepSeek V4 Pro 0813$0.32
Bedside Bench: Ember-1 minus Kimi K3, points (our arithmetic)−0.6 points
Bedside Bench: how much less Ember-1 costs per task than Kimi K3 (our arithmetic)23%
Big Finance: Ember-1 minus Kimi K3, points (our arithmetic)−2.3 points
Big Finance: how much less Ember-1 costs per task than Kimi K3 (our arithmetic)11%
DeepSWE v1.1: Ember-1 minus Kimi K3, points (our arithmetic)−3.25 points
DeepSWE v1.1: how much less Ember-1 costs per task than Kimi K3 (our arithmetic)22%
τ³-Banking: Ember-1 minus Kimi K3, points (our arithmetic)−13.4 points
τ³-Banking: how much less Ember-1 costs per trial than Kimi K3 (our arithmetic)35%

AI BENCHY

Ember-1 at max: tests fully passed10
tests in AI BENCHY22
Ember-1 at max: tests failed with an API error8
Ember-1 at max: tests failed with a wrong answer3
Ember-1 at max: tests failed with no answer1
Ember-1 at high: tests fully passed12
Ember-1 at high: tests failed with an API error7
Ember-1 at low: tests fully passed13
Ember-1 at low: tests failed with an API error6
Ember-1 with reasoning off: tests fully passed8
Ember-1 with reasoning off: tests failed with an API error5
Kimi K3 at max: tests fully passed16
Kimi K3 at max: tests failed with an API error2
Ember-1 at max: tests without an API error (our arithmetic)14
Kimi K3 at max: tests without an API error (our arithmetic)20

Aquiles (Bro Frontier Index)

Bro Frontier Index, Intelligence, Kimi K3 at max14.94
Bro Frontier Index, Intelligence, Ember-1 at max15.12
Bro Frontier Index, Coding, Kimi K3 at max23.48
Bro Frontier Index, Coding, Ember-1 at max24.30
Bro Frontier Index, Automation, Kimi K3 at max54.30
Bro Frontier Index, Automation, Ember-1 at max53.08
Bro Frontier Index, Knowledge Work, Kimi K3 at max39.30
Bro Frontier Index, Knowledge Work, Ember-1 at max39.04
Bro Frontier Index, Finance, Kimi K3 at max55.99
Bro Frontier Index, Finance, Ember-1 at max55.18
Bro Frontier Index, Legal, Kimi K3 at max23.74
Bro Frontier Index, Legal, Ember-1 at max20.20
Bro Frontier Index, Healthcare, Kimi K3 at max42.23
Bro Frontier Index, Healthcare, Ember-1 at max37.81
Bro Frontier Index, Cybersecurity, Kimi K3 at max30.25
Bro Frontier Index, Cybersecurity, Ember-1 at max24.44
Bro Frontier Index, cost per task, Kimi K3 at max$2.58
Bro Frontier Index, cost per task, Ember-1 at max$1.22
Bro Frontier Index, Intelligence: Ember-1 minus Kimi K3, points (our arithmetic)+0.18 points
Bro Frontier Index, Coding: Ember-1 minus Kimi K3, points (our arithmetic)+0.82 points
Bro Frontier Index, Automation: Ember-1 minus Kimi K3, points (our arithmetic)−1.22 points
Bro Frontier Index, Knowledge Work: Ember-1 minus Kimi K3, points (our arithmetic)−0.26 points
Bro Frontier Index, Finance: Ember-1 minus Kimi K3, points (our arithmetic)−0.81 points
Bro Frontier Index, Legal: Ember-1 minus Kimi K3, points (our arithmetic)−3.54 points
Bro Frontier Index, Healthcare: Ember-1 minus Kimi K3, points (our arithmetic)−4.42 points
Bro Frontier Index, Cybersecurity: Ember-1 minus Kimi K3, points (our arithmetic)−5.81 points

Aquiles (Bro Frontier Index)

Bro Frontier Index, overall score, Kimi K3 at max (his chart's label)31.91
Bro Frontier Index, overall score, Ember-1 at max (his chart's label)30.88
Bro Frontier Index, output tokens per task, Kimi K3 at max (his chart's label)47K
Bro Frontier Index, output tokens per task, Ember-1 at max (his chart's label)25K

Gert Labs

GBENCH overall rank, Kimi K3#15
GBENCH one-shot coding (one request), Kimi K3, score0.551
GBENCH one-shot coding (one request), Kimi K3, low end of Gert Labs' published interval0.521
GBENCH one-shot coding (one request), Kimi K3, high end of Gert Labs' published interval0.580
GBENCH agentic coding (many requests), Kimi K3, score0.659
GBENCH agentic coding (many requests), Kimi K3, low end of Gert Labs' published interval0.615
GBENCH agentic coding (many requests), Kimi K3, high end of Gert Labs' published interval0.699
GBENCH overall rank, Ember-1#26
GBENCH one-shot coding (one request), Ember-1, score0.556
GBENCH one-shot coding (one request), Ember-1, low end of Gert Labs' published interval0.529
GBENCH one-shot coding (one request), Ember-1, high end of Gert Labs' published interval0.582
GBENCH agentic coding (many requests), Ember-1, score0.518
GBENCH agentic coding (many requests), Ember-1, low end of Gert Labs' published interval0.465
GBENCH agentic coding (many requests), Ember-1, high end of Gert Labs' published interval0.566
models ranked on GBENCH100
GBENCH agentic coding, median model calls per game, Ember-160

Sources

These are the pages the video and this page draw on. We keep a copy of each page as we read it, so a figure can be checked against what the page said at the time.

Credits

The narration in the video is an AI voice, made with ElevenLabs.

The Bro Frontier Index figures are Aquiles'. The video redraws his charts with his credit, and this page draws them again from his posted charts and his data file.

Fireworks' chart "Average of 5 Industry Benchmarks" appears in the video as a quote from its launch post. This page doesn't reproduce it.

Music in the video: "Airport Lounge" by Kevin MacLeod (incompetech.com), licensed under Creative Commons: By Attribution 4.0.