English

Model ReleasesFireworks AIEmber-1

Fireworks AI announces "Ember-1," reducing reasoning tokens based on Kimi K3

This article is a translation. Read the Japanese original

Fireworks AI, a provider of AI fine-tuning services, has announced "Ember-1," a new AI model based on Kimi K3. This model aims to solve the challenge where cutting-edge AI models consume extensive tokens during the reasoning process, leading to increased execution costs.

Kimi K3, a Chinese open model, has shown performance exceeding GPT-5.6 Sol and Claude Fable 5 in some tests. However, it faced challenges in terms of cost and time, consuming many tokens during the reasoning stage before generating the final output and taking nearly an hour per task on average. According to Fireworks AI, approximately 90% of output tokens in such high-performance models are dedicated to reasoning.

Ember-1 is designed to eliminate unnecessary reasoning while retaining useful elements, such as feedback for task execution, by performing fine-tuning on Kimi K3. In measurements using the medical task benchmark "Bedside Bench," Ember-1 achieved lower costs while maintaining scores comparable to Kimi K3.

Furthermore, A/B testing conducted by Fireworks AI demonstrated that Ember-1 can reduce the number of output tokens per task by approximately 35% while maintaining performance equivalent to Kimi K3. The API pricing for Ember-1 is $3 per 1 million tokens for input, $0.3 for cached input, and $15 for output.

Sources

  1. 最先端AIモデルの「思考しすぎてコスト増加」を解決するべくKimi K3ベースで開発されたAIモデル「Ember-1」が登場 (GIGAZINE、2026-09-28)
  2. Fireworks AI