This is the video written out, with every figure in full, each linked to its source in the table below.
Fireworks' new model▶ 0:00
On 23 September, Fireworks released Ember-1. It's Kimi K3, the open-weight model from Moonshot, trained further so that it spends less time thinking before it answers. Fireworks' launch post says it "delivers Kimi K3's quality with 40% fewer tokens", and Fireworks charges the same per token for both: $3.00 per million input tokens, $0.30 per million cached input tokens and $15.00 per million output tokens. So every token Ember-1 doesn't write is money you don't spend.
So what does it give up? The saving looks real. Every source we found that reports cost or tokens has Ember-1 lower than K3. What you give up in quality depends on the kind of work you give it.
The basics▶ 0:40
Kimi K3 is a reasoning model. Before it answers, it writes out its thinking, and you pay for that thinking as output tokens, at $15.00 per million on Fireworks. Moonshot gives it an effort setting with three levels, low, high and max, and max is the default. Artificial Analysis, the independent benchmarking site, counts 48,455 output tokens per task for K3 at max on its Intelligence Index, 32,453 of them thinking and 16,002 the answer.
Fireworks says turning the effort down wasn't the answer: "Lower effort settings gave up too much quality." So it trained the model to think more briefly instead, and that's Ember-1. It keeps an effort setting of its own.
Ember-1 is a research preview. None of Fireworks' pages we saved (the launch post, the model page, the docs) links to weights, a paper or a system card, a Hugging Face search found no Ember-1 weights, and OpenRouter lists no provider other than Fireworks. Fireworks says it gives research models "two-week serverless access", and makes them permanent based on community demand. So everything here comes from Fireworks' own pages and from the few people outside Fireworks who have tested it.
Where the 40% comes from▶ 1:38
The launch post's headline says 40% fewer tokens. The same post puts the saving at "approximately 35% fewer tokens per task" in its section on customer tests, says reasoning traces could be shortened by 35–50% in its section on method, and calls one section "half the tokens, same answers".
The figure in it closest to the headline is a 39% cut in total tokens, from live tests on two customers' production coding traffic, where Ember-1 scored about the same as K3.
Fireworks' customer tests
Live A/B tests on two customers' production coding traffic, as the launch post reports them
Those customer tests are Fireworks' own, and nobody outside Fireworks can check them. The post doesn't name the customers or the tasks, or say what the score measures.
Fireworks' benchmark table▶ 2:11
Further down, the same post has a table of public benchmarks: four coding and terminal tests and one airline customer-service test, with K3 at all three effort settings next to Ember-1. Against K3 at max, Ember-1's cost saving runs from 5.9% on the airline test to 51.9% on Terminal Bench, a set of tasks done in a real terminal. None of the five is 40%, and together they average 25.9%. At the same price per token, those savings come from using fewer tokens.
What Ember-1 saves on each benchmark
Cost saving against Kimi K3 at max, as Fireworks' launch post reports it
The scores barely move. Ember-1's pass rate against K3 at max is +1.1 points on Terminal Bench, +2.0 points on the airline test, and −1.0 points and −1.3 points on SWE-bench Verified and SWE-Interact. The exception is DeepSWE, a set of deep debugging tasks, where Ember-1 passes 75.2% against K3's 66.4%, a lead of 8.8 points. Keep that one in mind.
The scores barely move, except on DeepSWE
Pass rate on each benchmark, Kimi K3 at max against Ember-1, as Fireworks' launch post reports it
The numbers in this chart
| Kimi K3 at max | Ember-1 | difference (our arithmetic) | |
|---|---|---|---|
| Terminal Bench 2.1 | 80.9% | 82.0% | +1.1 points |
| SWE-bench Verified | 93.2% | 92.2% | −1.0 points |
| SWE-Interact | 21.3% | 20.0% | −1.3 points |
| DeepSWE 1.1 | 66.4% | 75.2% | +8.8 points |
| τ-2 Bench Airline | 64% | 66% | +2.0 points |
Fireworks' benchmark table
Pass rates for Kimi K3 at each effort setting and for Ember-1, and Ember-1's cost saving against Kimi K3 at max, from Fireworks' launch post
| benchmark | tasks | K3 low | K3 high | K3 max | Ember-1 | saving vs K3 max |
|---|---|---|---|---|---|---|
| Terminal Bench 2.1 | 89 | 76.4% | 77.6% | 80.9% | 82.0% | 51.9%, $23.1 |
| SWE-bench Verified | 500 | 80.4% | 86.0% | 93.2% | 92.2% | 15.5%, $68.1 |
| SWE-Interact | 75 | 6.7% | 13.3% | 21.3% | 20.0% | 32.5%, $60.8 |
| DeepSWE 1.1 | 113 | 55.8% | 62.8% | 66.4% | 75.2% | 23.7%, $126.9 |
| τ-2 Bench Airline | 50 | 64% | 64% | 64% | 66% | 5.9%, $0.3 |
Fireworks' own leaderboard▶ 3:09
The same week, Fireworks launched a leaderboard of its own, the Specialized Intelligence Index, built from benchmarks it says were "contributed by industry practitioners". It has boards for different kinds of work, and four of them list both Ember-1 and K3. The index doesn't say which effort setting either model ran at.
On all four boards, Ember-1 costs less than K3, between 11% and 35% less (our arithmetic on the index's costs). And on all four it scores lower. On Bedside Bench, clinical cases written by physicians, it's 0.6 points lower; on Big Finance, 2.3 points; and on DeepSWE, 3.25 points.
Lower on all four of Fireworks' own boards
Score on each board of the Specialized Intelligence Index that lists both models
The numbers in this chart
| Kimi K3 | Ember-1 | difference (our arithmetic) | |
|---|---|---|---|
| Bedside Bench | 89.7 | 89.1 | −0.6 points |
| Big Finance | 41.3 | 39.0 | −2.3 points |
| DeepSWE v1.1 | 70.21 | 66.96 | −3.25 points |
| τ³-Banking | 51.5 | 38.1 | −13.4 points |
That's the same DeepSWE where the launch post had Ember-1 8.8 points ahead. Here it's 3.25 points behind. Both models' DeepSWE scores differ between the two: Ember-1's from 75.2% in the launch post to 66.96 on the leaderboard, K3's from 66.4% to 70.21, and neither says why.
The fourth board is where Ember-1 drops furthest. τ³-Banking tests banking customer-service agents against a simulated customer. K3 scores 51.5 there and Ember-1 38.1.
The four boards that list both models
Score and cost on Fireworks' Specialized Intelligence Index
| board | Kimi K3 | Ember-1 | cost cut (our arithmetic) |
|---|---|---|---|
| Bedside Bench (medicine) | 89.7 at $0.0529 | 89.1 at $0.0406 | 23% |
| Big Finance | 41.3 at $0.220 | 39.0 at $0.195 | 11% |
| DeepSWE v1.1 (coding) | 70.21 at $5.40 | 66.96 at $4.19 | 22% |
| τ³-Banking (customer service) | 51.5 at $1.02 | 38.1 at $0.66 | 35% |
These boards list other models too. On three of the four, a model that costs less than Ember-1 scores higher: OpenAI's GPT-6 Sol on Big Finance and on DeepSWE, and GLM-5.3 on the banking board, among others. On Bedside Bench, no model with a listed cost does better for less.
Models that score higher than Ember-1 for less
On the same boards, every model with a listed cost below Ember-1's and a score above it
| board | model | score at cost |
|---|---|---|
| Big Finance | Ember-1 | 39.0 at $0.195 |
| GPT-6 Sol | 47.1 at $0.085 | |
| DeepSeek V4.1 Flash | 40.6 at $0.107 | |
| DeepSeek V4 Pro 0813 | 39.3 at $0.141 | |
| DeepSWE v1.1 | Ember-1 | 66.96 at $4.19 |
| GPT-6 Astra | 71.0 at $2.56 | |
| GPT-6 Sol | 67.3 at $2.31 | |
| τ³-Banking | Ember-1 | 38.1 at $0.66 |
| GLM-5.3 | 51.6 at $0.49 | |
| Gemini 3.8 Flash | 51.4 at $0.53 | |
| DeepSeek V4 Pro 0813 | 42.7 at $0.32 |
Tests outside Fireworks▶ 4:41
Outside Fireworks, there's much less. Artificial Analysis, which published numbers on Sonnet 5.5 within about half an hour of its launch, still hadn't listed Ember-1 when we checked on 29 September, six days in. It has Kimi K3 at max at 43.6 on its Intelligence Index.
A small public test suite, AI BENCHY, did run it. Ember-1 at max fully passed 10 of its 22 tests, against 16 for K3 at max. But 8 of Ember-1's failed tests at max were API errors, requests that came back with an error instead of an answer, and it got API errors at every effort setting AI BENCHY tried. AI BENCHY doesn't say what caused them, but every route to Ember-1 we found runs on Fireworks, so they're worth knowing about. Set the API errors aside and Ember-1 passed 10 of the remaining 14, K3 16 of 20, which on a suite this small is too close to tell apart.
AI BENCHY: tests fully passed
A small public suite of tests, run on each model and effort setting
| model and setting | passed | API errors |
|---|---|---|
| Kimi K3 at max | 16 of 22 | 2 |
| Ember-1 at max | 10 of 22 | 8 |
| Ember-1 at high | 12 of 22 | 7 |
| Ember-1 at low | 13 of 22 | 6 |
| Ember-1, reasoning off | 8 of 22 | 5 |
Another outside run is the Bro Frontier Index, from Aquiles, a developer who is building an AI app called Bro. It covers eight categories of practical work, and we couldn't find its tasks published. With both models at max effort, but in different agent tools (OpenCode for Ember-1, Kimi Work for K3), his charts show Ember-1 writing 25K output tokens per task against K3's 47K and costing $1.22 a task against $2.58, while scoring 30.88 overall against K3's 31.91. Five of the eight categories stayed within about a point, coding among them. The other three dropped: law by 3.54 points, healthcare by 4.42 points and cybersecurity by 5.81 points. His data file also gives Ember-1 the same input tokens per task as K3 and names Moonshot as its API provider, though every route to Ember-1 we found runs on Fireworks, so some of the Ember-1 row may have been copied from K3's. It's one run, and it's the only source on law and security.
Bro Frontier Index, by category
Both models at max effort, in different agent tools, redrawn from Aquiles' data file
The numbers in this chart
| Kimi K3 at max | Ember-1 at max | difference (our arithmetic) | |
|---|---|---|---|
| Intelligence | 14.94 | 15.12 | +0.18 points |
| Coding | 23.48 | 24.30 | +0.82 points |
| Automation | 54.30 | 53.08 | −1.22 points |
| Knowledge Work | 39.30 | 39.04 | −0.26 points |
| Finance | 55.99 | 55.18 | −0.81 points |
| Legal | 23.74 | 20.20 | −3.54 points |
| Healthcare | 42.23 | 37.81 | −4.42 points |
| Cybersecurity | 30.25 | 24.44 | −5.81 points |
On 28 September, Leo Linsky posted results for Ember-1 from Gert Labs' benchmark, GBENCH. GBENCH has models write programs that play games against each other. When the model writes its program in a single request, Ember-1 and K3 come out level, 0.556 against 0.551, with Gert Labs' published intervals overlapping. In what Gert Labs calls agentic coding, where the model works on its program over many requests (for Ember-1, a median of 60 per game), Ember-1 scores 0.518 against K3's 0.659, and the intervals don't overlap. Overall, GBENCH ranks K3 #15 and Ember-1 #26 of 100 models, though Ember-1 hasn't been run on its decision-making games, which K3 has.
GBENCH: level in one request, lower over many
Gert Labs' score for code that plays games, with the interval Gert Labs publishes for each score
Each of these is one outside test, set up its own way, and they don't always agree with Fireworks: its medical board showed a gap of 0.6 points, where the Bro index's healthcare category showed 4.42 points. So the drops on law, health and security, and on long coding tasks, are worth checking before you rely on Ember-1 for them.
Who's talking about it▶ 6:47
Most of what's been said about Ember-1 has been in posts on X and Hacker News, and most of the engagement we found came on 27 September, four days after launch, when it reached Hacker News. On X, three of the five most-liked posts came from Fireworks' side: the company's own account, its co-founder and CTO Dmytro Dzhulgakov, and Sonya Huang, who works at Sequoia, one of Fireworks' investors. The other two, from the coding agent Cline and from Elvis Saravia, repeated Fireworks' numbers. None of the five reports a test.
Four people posted what they found with some detail: Aquiles, with his index; Leo Linsky, with the Gert Labs results; a Hacker News commenter whose model router, a tool that picks a model for each request, gave Ember-1 full marks on coding and still ranked GPT-6 Sol above it, partly because the router had no benchmark score for Ember-1 to start from; and Vojtech Rinik, who found it slow on one long voice memo. One more post on X said a single reasoning question ran on about forty percent fewer tokens than on K3 with no loss in quality, and gave no data. And on Hacker News, someone asked the question the video is about: "What, if any, capability is lost by the token reduction?"
So what does it lose?▶ 7:57
On coding, most scores are within a point or two either way, but one outside benchmark found a clear drop when Ember-1 works on code over many requests. On customer service, Fireworks' own banking board shows a drop from 51.5 to 38.1, while the launch post's airline test went 2.0 points the other way. On finance, Fireworks' board has it 2.3 points lower and the Bro index about level. On law and security, only the Bro index has tested it, in its one run, and it's lower on both. On medicine, both sources have it lower: Fireworks' medical board by 0.6 points, the Bro index by 4.42 points.
What does it lose?
Ember-1's score minus Kimi K3's (our arithmetic), in points on each source's own scale, by kind of work and by whose number it is
| kind of work | Fireworks' launch post | Fireworks' index | Bro Frontier Index | GBENCH |
|---|---|---|---|---|
| Coding | Terminal Bench +1.1, SWE-bench Verified −1.0, SWE-Interact −1.3, DeepSWE +8.8 | DeepSWE −3.25 | +0.82 | one request: level; many requests: lower |
| Customer service | airline +2.0 | banking −13.4 | ||
| Finance | −2.3 | −0.81 | ||
| Medicine | −0.6 | −4.42 | ||
| Law | −3.54 | |||
| Security | −5.81 |
What nobody has measured▶ 8:34
Two comparisons would settle most of what's left, and we found no one outside Fireworks who has run either.
The first is Ember-1 against K3 simply set one step lower, at high effort, the option Fireworks says doesn't work. Fireworks' own numbers give a first hint. Averaged over the five benchmarks in its table (our arithmetic), K3 at high passes 60.7%, K3 at max 65.2% and Ember-1 67.1%, and the post's own chart of a five-benchmark average against cost per task (its averages differ a little from ours) has K3 at high a little cheaper than at max and Ember-1 cheaper than both. Ember-1's lead over K3 at max comes almost entirely from DeepSWE: it's +1.9 points across all five benchmarks and +0.2 points across the other four. If Fireworks' result holds on other people's tasks, training the model to think less did better than simply turning K3 down. AI BENCHY, the one outside suite we found that ran Ember-1 at more than one setting, has K3 only at max.
The second is the other claim in Fireworks' section title, "same answers". A score can stay the same while the answers change, and nobody outside Fireworks has checked whether Ember-1 gives the answers K3 would, on your kind of work.
Should you try it this week?▶ 9:31
Until someone measures those, here's how we'd act on what's public. If you run K3 on short coding tasks, Ember-1 is worth trying on some of your own traffic. It costs the same per token, so compare output tokens and pass rates on the same tasks. If your work is long agentic coding, customer-service conversations, or legal, medical or security questions, check more carefully, because that's where the drops so far showed up. And if you aren't tied to K3, cheaper models score higher than Ember-1 on three of Fireworks' four boards.
Fireworks hasn't published an end date for the preview. Two weeks from launch would be 7 October, but that's our arithmetic, not a date Fireworks has given. Its docs promise at least two weeks' advance notice before removing a model, and when we checked on 29 September there was no notice, though the docs don't say whether that covers research releases.
Five days after Ember-1, Anthropic made much the same promise for Sonnet 5.5, fewer tokens for the same work. That's our previous video.