Fireworks Research has announced the release of Ember-1, a new specialized model designed to provide the same quality as Kimi K3 but with significantly higher token efficiency. According to the company, the model achieves approximately 40% fewer tokens per task by learning to perform efficient reasoning, avoiding the unproductive loops often found in long reasoning traces.
Model ReleasesFireworks ResearchEmber-1
Fireworks Research Launches Ember-1, a Token-Efficient Version of Kimi K3
The development involved over 50 training experiments and 200 evaluations conducted on the Fireworks Serverless Training platform. In live A/B tests with two customers on production coding workloads, Ember-1 delivered approximately 35% token savings per task while maintaining comparable quality. One customer has already begun running the model in live production with plans to scale it further.
Ember-1 was evaluated against industry benchmarks, including the Specialized Intelligence Index (SII) on the Bedside Bench. The company reported that Ember-1 sits on or near the Pareto frontier for cost-to-task efficiency, matching the quality of Kimi K3-max at a fraction of the cost.
The model is currently available as a Research Preview on Fireworks Serverless, alongside the base Kimi K3 model. Fireworks also introduced training support for Ember-1, allowing enterprises to build customized, token-efficient models tailored to their specific workloads.
Sources
- Ember-1 (Hacker News Frontpage, 2026-09-27)