mrkeyoor.com_
Mon 28 Sept 05:40 UTC
AI6 min read

Ember-1 Cut Agent Output From 49,300 to 29,900 Tokens

Fireworks says Ember-1 preserves Kimi K3 quality with shorter reasoning. Its own A/B table shows where the bill shrank and what remains unproved.

One production coding workload fell from 49,300 output tokens to 29,900 when Fireworks replaced Kimi K3 with Ember-1. The score edged from 0.751 to 0.753, while the average number of steps dropped from 23.8 to 21.4. Fireworks reports a 39% reduction in total tokens and a 71.3% cut in reasoning tokens for that A/B test. Those figures make the release more interesting than another model launch: Ember-1 charges the same listed per-token rates as Kimi K3, so its economic case rests on doing less work per answer.

That pitch drew 388 points and 196 comments on Hacker News by 19:30 UTC on September 27, according to the live candidate snapshot for this article. The score shows unusual developer interest, though it does not verify Fireworks' performance claims. The evidence comes from Fireworks' launch report, which says the company trained Ember-1 on Kimi K3 to shorten its reasoning traces while retaining answer quality. The result is available through an API as a research preview, rather than as a newly published set of weights.

The discount comes from output length

The list price clarifies what Fireworks changed. Both model cards show $3 per million uncached input tokens, $0.30 per million cached input tokens, and $15 per million output tokens. At that output rate, 49,300 tokens cost about $0.7395 and 29,900 cost about $0.4485. The difference is 29.1 cents for the output reported in the A/B table, before input charges. Kimi K3's current model card confirms the same three rates, so Ember-1's saving is a usage reduction rather than a cheaper token.

That distinction matters in an agent loop. Fireworks says reasoning can account for more than 90% of Kimi K3's generated tokens on some requests. A later turn then sends earlier reasoning back as context, causing old work to be read and billed again. The company describes context growth across a multi-turn task as roughly quadratic. Its explanation of the problem is specific to the way these reasoning traces accumulate. The bill for any one application will still depend on prompt size, caching, the number of turns, and how much history the client preserves.

The published A/B table does not provide enough input-token detail to reconstruct a full task bill. It does show why reducing output early can affect more than the completion charge on that turn. A shorter trace can also leave fewer tokens for later calls to ingest. Fireworks' serverless model page lists a 1.04-million-token context window, which gives agents room to carry long histories but does not make those histories free.

Five coding results do not move in one direction

Fireworks compares Ember-1 with Kimi K3 at its maximum reasoning setting on five public or industry benchmarks. On Terminal-Bench 2.1, Ember-1 scored 82.0% against 80.9% for K3 max. On DeepSWE 1.1, it scored 75.2% against 66.4%. The other two software engineering tests slightly favored K3 max: 93.2% versus 92.2% on SWE-bench Verified, and 21.3% versus 20.0% on SWE-Interact. Ember-1 led the airline subset of tau2-bench 66% to 64%. The complete table appears in Fireworks' release post.

Those mixed movements fit the narrower claim of comparable quality better than a claim of universal improvement. They also show why the sample counts belong beside the percentages. The table covers 500 SWE-bench Verified tasks, 113 DeepSWE tasks, 89 Terminal-Bench tasks, 75 SWE-Interact tasks, and 50 airline tasks. Fireworks says Ember-1 used 5.9% to 51.9% less money across these comparisons, depending on the benchmark, with the largest percentage reduction reported on Terminal-Bench. The company ran more than 50 training experiments and over 200 evaluations before release.

The evaluator and the vendor are the same company, which limits what the figures can settle. Fireworks' Specialized Intelligence Index says its standard setup uses the same harness, tools, and task count across models, with three runs per task unless an exception is recorded. It also states that the index measures one standardized serving setup, not the best score a team could reach with a model-specific harness. Those boundaries are documented in the index methodology, and independent replications would carry more weight than another vendor-run chart.

The production evidence is thinner. Fireworks says it tested two customers' coding traffic and observed roughly 35% fewer tokens per task at comparable quality. One customer had moved Ember-1 into production and planned to expand it, but the post does not name either customer, state the number of tasks, or publish the grading method behind the 0.751 and 0.753 scores. The A/B section gives useful measurements, yet it cannot tell another team whether a five-point refactor, a repository migration, and a browser agent will save the same proportion.

Fireworks trained away unproductive reasoning

Ember-1 is listed as a 2.78-trillion-parameter mixture-of-experts model with image input, function calling, and the same 1.04-million-token context length shown for Kimi K3. Fireworks says it trained across mathematics, coding, instruction following, conversation, search, tool use, and software engineering. The data included single problems and longer interactions, and the company says it used its own data rather than customer data. The launch account does not disclose the size of that training collection or release the training algorithm in enough detail for an outside team to reproduce it.

The design target was selective restraint. Fireworks says lowering K3's reasoning-effort setting reduced quality too sharply, while training could remove repeated or unproductive thought and keep useful self-correction. Its results also claim shorter unsuccessful attempts, an important detail for agents that can spend many tokens before failing. The practical test is whether Ember-1 stops a dead-end trajectory sooner without cutting off the reconsideration that would have rescued it. That behavior needs trajectory-level evaluation. A final pass rate cannot reveal it, as the company's own account of its training goal makes clear.

Kimi K3 gives the derivative model a large starting point. Moonshot AI describes K3 as an open-weight, native multimodal model with 2.8 trillion parameters and a one-million-token context window. Its official repository publishes the base model's code, report, and license. Fireworks calls Ember-1 its own model, but the Ember launch does not announce a weights download or a license for the specialized version. Developers who need self-hosting should treat access to K3's weights and access to Ember-1 as separate questions.

The preview has two unresolved product details

Availability currently comes with a clock. Fireworks says research releases receive two weeks of serverless access and become permanent according to community demand. The Ember-1 model card marks the endpoint as ready, supports serverless use, and identifies the API path as accounts/fireworks/models/ember-1. A team considering a production switch still needs confirmation that the endpoint will remain after the preview and what migration path applies if it does not.

Fine-tuning is also unclear across Fireworks' own pages. The launch post says training support for Ember-1 is being introduced, while the current model card says fine-tuning is not supported. Kimi K3's card, by comparison, lists fine-tuning as supported. This may reflect a staged rollout, but the two Ember descriptions do not yet give developers one consistent availability status. That matters if the attraction is adapting the shorter-reasoning behavior to a private workload.

A sensible evaluation can stay small. Run Kimi K3 and Ember-1 on the same representative agent tasks, preserve the same prompts and tool definitions, and record completion rate, steps, uncached input, cached input, output, and total cost. Fireworks' index methodology uses comparable harnesses for the same reason. A mean token reduction is useful, but the expensive failures at the tail deserve separate inspection because one looping task can erase the savings from several tidy ones.

After the 388-point discussion fades, Fireworks first has to keep the endpoint beyond its two-week preview and reconcile the fine-tuning status. Outside evaluations then need to reproduce the token savings. Until then, the 49,300-to-29,900 result is a credible reason to test. It remains one vendor's measured workload under a research-preview release, rather than a general law about how much reasoning an agent needs.

We reviewed this

  1. serverless — our honest review
  2. terminal — our honest review
  3. harness — our honest review

Sources

  1. Fireworks AI: Introducing Ember-1
  2. Fireworks AI: Ember-1 model card
  3. Fireworks AI: Kimi K3 model card
  4. Fireworks Specialized Intelligence Index methodology
  5. Moonshot AI: Kimi K3 repository