Fireworks launches Ember-1 to cut Kimi K3 reasoning tokens by 40 percent
Fireworks trained Kimi K3 to stop overthinking, cutting reasoning tokens in half without hurting quality for coding and AI agents.
Models · Source: Hacker News
What happened
Fireworks AI just released Ember-1. It is a new specialized model built on top of Kimi K3. The goal is simple but highly requested by developers. Keep the reasoning quality of K3 but cut the massive token bloat. Fireworks ran over fifty training experiments and two hundred evaluations to get this right. They used their own Serverless Training platform to move fast without managing GPUs. They used no customer data for training.
Reasoning models waste a lot of tokens. Kimi K3 sometimes spends up to 90 percent of its output on internal thinking rather than the answer. This gets incredibly expensive in multi-turn agent workloads. Context grows quadratically because every turn replays all prior reasoning back to the model. Fireworks trained Ember-1 to learn efficient reasoning. It keeps useful self-reflection but drops unproductive loops and shortens unsuccessful attempts.
The benchmark results are highly competitive. Across seven public benchmarks and live production traffic, Ember-1 cuts reasoning tokens by 35 to 50 percent. It was tested on Doximity Bedside Bench against GPT-5.6 Sol, GPT-6 Astra, and Claude Opus 5. Ember-1 set a new Pareto frontier for cost versus task performance. Live A/B tests with two enterprise customers showed a 35 percent token drop with improved task completion. Ember-1 is available today as a two-week research preview.
Key facts
- 40% — Reduction in tokens compared to Kimi K3
- 90% — Portion of generated tokens Kimi K3 sometimes spends on internal reasoning
- 35-50% — Reduction in Kimi K3 reasoning possible without sacrificing accuracy
- 500 — Clinical cases in Doximity Bedside Bench used to evaluate Ember-1
Why it matters
Agentic workflows are currently too expensive for most production use cases. Every turn forces the model to re-read and re-bill previous long reasoning traces. By shrinking the reasoning footprint, Ember-1 makes multi-turn agents financially viable for bootstrapped founders. You get the intelligence of a frontier model without the massive token tax. This completely changes the unit economics for automated coding tools and customer support agents. Your margins improve overnight.
The broader industry shift is clear. The era of brute-force reasoning is ending. We are moving from raw thinking models to specialized efficient thinkers. Fireworks is proving that you can train a model to self-reflect faster and stop wasting tokens on dead ends. This shifts the competitive edge for AI startups. It is no longer just about who has the smartest base model. It is about who has the most token-efficient inference stack for specific industry workloads.
For builders
Cheaper multi-turn agent loops
Multi-turn agents replay prior reasoning on every turn. Ember-1 cuts this quadratic cost growth by shortening the reasoning traces. Founders building coding agents pay significantly less for API calls while keeping high success rates.
Custom token-efficient models
Fireworks also launched training support for Ember-1. Enterprises can use their own data to build custom models tailored to their workloads. You pay upfront for training but save massively on inference at scale.
Two week research preview window
Ember-1 is currently a research preview on Serverless. Developers have two weeks to test it before Fireworks decides to make it permanent based on demand. Builders need to evaluate it quickly to lock in the cost savings.
My take
Brute force reasoning is a lazy tax on developers. Fireworks gets it exactly right by optimizing the thinking process itself instead of just turning down the effort slider. If your AI agent spends 90 percent of its time talking to itself, you are burning startup cash for absolutely no reason. Stop paying for unproductive AI loops and start optimizing your token spend.
Original reporting: Hacker News. This is my rewrite and opinion.