AI-generated content. This article was researched and written by an automated AI editorial system and published without prior human review. Every factual claim is checked against cited primary sources before publication, but no journalist read this page before you did — treat it accordingly, and report anything that looks wrong. How this works ›

Some links on this page are affiliate links. We may earn a commission at no extra cost to you.
Updated: Sep 26, 2026
·
fireworkskimiopen-weightspricingreasoninginferenceprocurementbenchmarksagentsdevelopers

Fireworks priced Ember-1 identically to Kimi K3 — the discount is in the token count, and it expires in two weeks unless enough people use it

TL;DR: Fireworks released Ember-1 on 23 September 2026 — a Kimi K3 derivative from Fireworks Research claiming 35-50% shorter reasoning traces at comparable quality. Its rate card is identical to base Kimi K3 on the same platform: $3.00 input / $0.30 cached / $15.00 output per million tokens. So the discount is not on the price list at all — it is a claim about how many tokens you will consume, which only your own bill can confirm. The arithmetic is shape-dependent: a 35% output cut is a 23% total saving on a 10k-in/4k-out task and under 6% on a 100k-in/4k-out task. Benchmarks are mixed and self-reported — up on Terminal Bench 2.1 (80.9→82.0) and DeepSWE 1.1 (66.4→75.2), down on SWE-bench Verified (93.2→92.2) and SWE-Interact (21.3→20.0), with no independent replication. And it is a two-week research preview made permanent only on community demand. Measure it now; do not depend on it.

A price cut with no price change

Every cost reduction this desk has covered this year showed up as a smaller number on a rate card. Ember-1 does not.

Fireworks published Ember-1 on 23 September 2026, describing it as a reasoning model from Fireworks Research built on Kimi K3, designed to “make every token go further” by producing shorter reasoning traces — roughly 40% fewer tokens than the base model across its evaluations, and about 35% fewer tokens per task in live customer A/B tests at comparable quality.

The rate card reads: $3.00 per million input tokens, $0.30 cached, $15.00 output, on a context window of roughly 1.04 million tokens.

Those are exactly the numbers Fireworks charges for base Kimi K3 on serverless — $3.00, $0.30 and $15.00 — and exactly what Moonshot charges on its own API. Not approximately. Identically.

This is an unusual and, on reflection, rather clean way to sell a model. There is no premium to justify, no discount to explain, and no pricing-page comparison to run. The proposition is simply: same rates, same quality, fewer tokens. Which means the only instrument that can tell you whether it worked is your invoice.

Where the saving goes, and where it evaporates

Reasoning tokens bill as output tokens. That single fact determines the entire economics here, because it means Ember-1 shrinks exactly one line of your bill and leaves the other untouched.

Take a task with 10,000 input tokens and 4,000 output tokens on Kimi K3:

Now cut output by 35%, to 2,600 tokens, at identical rates:

That is a 23% saving, not 35%. The headline number applies to a line item, and the line item is a fraction of the bill.

Now change one thing — make the prompt long, which is the shape most retrieval-augmented and long-context agentic work actually takes. 100,000 input tokens, same 4,000 output:

A saving of under 6%. Same model, same claim, same arithmetic — and the benefit fell by a factor of four because of prompt shape alone.

This is the practical headline for buyers. Ember-1’s value is inversely proportional to how input-heavy your workload is. Short prompts with long deliberation — hard reasoning, multi-step agent loops — approach the advertised saving. Long prompts with short answers get close to nothing. No benchmark table can tell you which you are; a single query against your own logs can.

It is worth noting the one case where the logic inverts pleasantly: shorter outputs also mean lower latency and fewer tokens for downstream steps to read, so an agent chain compounds the benefit across hops in a way a single-call cost model understates.

The benchmarks are mixed, and self-reported

Fireworks published a comparison against K3 max that does not flatter itself uniformly:

BenchmarkKimi K3 maxEmber-1
Terminal Bench 2.180.9%82.0%
DeepSWE 1.166.4%75.2%
SWE-bench Verified93.2%92.2%
SWE-Interact21.3%20.0%

Two up, two down. The declines are small, but Fireworks printed them rather than selecting the favourable pair, and the pattern is coherent with the training objective: compressing reasoning helps where over-deliberation wastes tokens and costs a little where the extra thinking was earning its keep.

The caveat that matters more than any individual row is that all of these are Fireworks measuring Fireworks, including the ~35% token reduction from customer A/B tests. No independent replication existed at the time of writing. A serving provider has a fair claim to knowing its own token counts better than anyone — but a token-reduction claim published by the company that bills per token is precisely the figure to reproduce on your own traffic before planning around it. The same scepticism this desk applies to vendor-published launch benchmarks generally applies here without modification.

The two-week clock

Ember-1 ships as a research preview. Fireworks describes a new programme in which research releases get two weeks of serverless access and become permanent based on community demand. Published 23 September, that window closes in the first week of October.

The incentive is reasonable rather than sinister: a serving provider needs evidence that a specialised derivative earns its GPU capacity before committing to it indefinitely, and a short public trial is a sane way to get that evidence. But it moves a specific risk onto the buyer. The cost of migrating onto Ember-1 is yours; the decision about whether it continues to exist is not.

That asymmetry prescribes the usage precisely:

How it sits against the frontier tier

The obvious comparison is GPT-6 Sol, cut to $2/$10 on 22 September. Sol is a flat 50% cheaper per token than Ember-1. On the benchmark they share, Fireworks puts Ember-1 at 75.2% on DeepSWE 1.1 against OpenAI’s published 68.8% for Sol at max effort and 66.6% for Luna at max.

So: half again the token price, more than six points ahead on that benchmark, with a token-volume claim against its own base model that has never been measured against Sol. Anyone deciding between them has to run the comparison themselves, and the decisive variable is once again the input-to-output ratio rather than anything either vendor published.

The structural differences outweigh the arithmetic anyway. Sol is generally available and unversioned; Ember-1 has a survival clock but rests on open weights you can take elsewhere. Neither property appears on a rate card.

The thesis, stated plainly

The open-weights price floor has spent this year moving in public — stealth launches that reset the floor, a price war that inverted direction in August, and rate cards from DeepSeek and Qwen that made per-token comparison the default way buyers shop. Ember-1 is the first serious move that is invisible to that entire apparatus. Every price-comparison table, every cost dashboard keyed to dollars-per-million-tokens, and every procurement spreadsheet that ranks models by rate will show Ember-1 and Kimi K3 as identical products. They are not, and the difference is a number none of those tools measure.

That is not a criticism of Fireworks, which has been unusually transparent — it published its regressions, disclosed the preview terms, and did not dress a token-efficiency claim up as a discount. It is a warning about the instrument. Dollars per million tokens is a unit price, and unit price has never been cost. As soon as vendors start competing on consumption rather than on rate, the buyer who ranks models by rate card is optimising the one variable the market has stopped competing on.

For developers and anyone maintaining a coding-agent stack, the operational change is small: start logging tokens per completed task, not just dollars per million tokens, as a first-class metric alongside quality. That number is the only one that moves when a model like this works — and it is the same number that would have caught a silent routing downgrade at peak hours going the other way. Measure Ember-1 in the next two weeks, on your own traffic, with that metric. Then decide — knowing the vendor is deciding too.

Frequently asked questions

How much does Ember-1 actually save me?

Less than 35%, and the exact figure depends on how input-heavy your workload is — which is why no vendor can quote it for you. Ember-1 and Kimi K3 bill identically on Fireworks serverless: $3.00 per million input tokens, $0.30 cached and $15.00 output. Only the output line shrinks, because reasoning tokens bill as output tokens, so a shorter reasoning trace is a smaller output bill and nothing else. Work an example. A task with 10,000 input tokens and 4,000 output tokens costs $0.03 + $0.06 = $0.09 on K3. Cut output by 35% to 2,600 tokens and it costs $0.03 + $0.039 = $0.069 — a 23% saving on the total, not 35%, because input did not move. Now make it input-heavy, the shape most retrieval and long-context agent work takes: 100,000 input and the same 4,000 output costs $0.36 on K3 and $0.339 on Ember-1, a saving of under 6%. Same model, same claim, same arithmetic, and the benefit fell by a factor of four purely because of prompt shape. If your prompts are long and your answers short, this model is close to free of benefit. If your prompts are short and your reasoning traces are long — agentic loops, hard single-shot reasoning — you approach the headline number.

Is Ember-1 actually better than Kimi K3, or just cheaper to run?

On Fireworks' own numbers it is better on two benchmarks and slightly worse on two others, which is a more honest result than a clean sweep would be. Terminal Bench 2.1 rises from 80.9% to 82.0% and DeepSWE 1.1 rises substantially, from 66.4% to 75.2%. Against that, SWE-bench Verified falls from 93.2% to 92.2% and SWE-Interact falls from 21.3% to 20.0%. Those declines are small but they are real and Fireworks published them rather than omitting them. The pattern is consistent with what the model was trained to do: compress reasoning, which helps on tasks where over-thinking wastes tokens and hurts marginally on tasks where the extra deliberation was earning its keep. The important caveat is that every one of these figures is Fireworks measuring its own model, including the customer A/B result of roughly 35% fewer tokens per task at comparable quality. No independent replication existed at the time of writing. That does not make the numbers wrong — a serving provider has a reasonable claim to knowing its own token counts — but a token-reduction claim from the company that bills you per token is exactly the sort of figure to reproduce yourself before you plan around it.

What happens at the end of the two-week research preview?

Fireworks says the model becomes permanent based on community demand, which means the honest answer is that nobody knows, including Fireworks. The company describes this as a new programme: research releases get two weeks of serverless access, and adoption during that window decides whether they stay. Ember-1 was published on 23 September 2026, so the window closes in the first week of October. Read the incentive structure plainly rather than cynically. A serving provider needs to know whether a specialised derivative earns its capacity before committing GPUs to it indefinitely, and a short public trial is a reasonable way to find out. But it transfers a specific risk to you: the cost of migrating onto Ember-1 is yours, and the decision about whether it continues to exist is not. That asymmetry dictates the correct usage. Evaluate it now, because evaluation is cheap and the rates are identical so there is no premium to justify. Do not put it on a critical path, do not write it into a customer commitment, and do not let it become the only model a production route knows how to call.

How does Ember-1 compare with GPT-6 Sol on price and capability?

Sol is cheaper per token and Ember-1 scores higher on the coding benchmark they share, so the comparison turns entirely on token volume. GPT-6 Sol bills $2.00 input and $10.00 output per million tokens after its 22 September cut; Ember-1 bills $3.00 and $15.00. That is a flat 50% premium per token for Ember-1. On DeepSWE 1.1, Fireworks reports Ember-1 at 75.2%, against OpenAI's published Sol figure of 68.8% at max reasoning effort and Luna at 66.6% at max. So Ember-1 costs half again as much per token and leads by more than six points on that benchmark, while also claiming to use fewer tokens than the base model it was derived from — though not necessarily fewer than Sol, which is a comparison nobody has published. Two structural differences matter more than the arithmetic. Ember-1 is a derivative of open weights, so the underlying Kimi K3 model remains available to self-host if the hosted version disappears, whereas Sol exists only as long as OpenAI serves it. Against that, Sol is a generally available product and Ember-1 has a two-week survival clock. Neither of these is a rate-card property, and neither will show up in a price comparison table.

What is the right way to test this in the two weeks available?

Replay real traffic and measure tokens, not quality alone, because quality parity is the claim you are most likely to confirm and token reduction is the claim that actually pays. Take a representative sample of production requests — a few hundred is enough — and run them through both moonshotai/kimi-k3 and fireworks/ember-1 on Fireworks serverless, capturing three things per request: output token count, reasoning token count where exposed, and your own quality judgement or automated grade. Then compute the saving on your actual input-to-output ratio rather than on Fireworks' benchmark mix. Two traps to avoid. First, do not test with synthetic short prompts if your production prompts are long, because as the arithmetic above shows that single difference can swing the benefit from 23% to under 6%. Second, hold reasoning effort constant between the two models if the parameter is exposed, since a token-count comparison across different effort settings measures the setting rather than the model. Record the date alongside the model name in whatever you keep, because an unversioned hosted model can change under you — the lesson from OpenAI shipping an unannounced vision fix behind an unversioned model ID three days after launch applies to every hosted endpoint, not just that one.

Sources

Related tool reviews

Questions or corrections? Email Pick Right. Want the full list? See all news.