Fireworks launches Ember-1 to cut Kimi K3 reasoning tokens by 40 percent

Fireworks trained Kimi K3 to stop overthinking, cutting reasoning tokens in half without hurting quality for coding and AI agents.

Models · Source: Hacker News

What happened

Fireworks AI just released Ember-1. It is a new specialized model built on top of Kimi K3. The goal is simple but highly requested by developers. Keep the reasoning quality of K3 but cut the massive token bloat. Fireworks ran over fifty training experiments and two hundred evaluations to get this right. They used their own Serverless Training platform to move fast without managing GPUs. They used no customer data for training.

Reasoning models waste a lot of tokens. Kimi K3 sometimes spends up to 90 percent of its output on internal thinking rather than the answer. This gets incredibly expensive in multi-turn agent workloads. Context grows quadratically because every turn replays all prior reasoning back to the model. Fireworks trained Ember-1 to learn efficient reasoning. It keeps useful self-reflection but drops unproductive loops and shortens unsuccessful attempts.

The benchmark results are highly competitive. Across seven public benchmarks and live production traffic, Ember-1 cuts reasoning tokens by 35 to 50 percent. It was tested on Doximity Bedside Bench against GPT-5.6 Sol, GPT-6 Astra, and Claude Opus 5. Ember-1 set a new Pareto frontier for cost versus task performance. Live A/B tests with two enterprise customers showed a 35 percent token drop with improved task completion. Ember-1 is available today as a two-week research preview.

Key facts

Why it matters

Agentic workflows are currently too expensive for most production use cases. Every turn forces the model to re-read and re-bill previous long reasoning traces. By shrinking the reasoning footprint, Ember-1 makes multi-turn agents financially viable for bootstrapped founders. You get the intelligence of a frontier model without the massive token tax. This completely changes the unit economics for automated coding tools and customer support agents. Your margins improve overnight.

The broader industry shift is clear. The era of brute-force reasoning is ending. We are moving from raw thinking models to specialized efficient thinkers. Fireworks is proving that you can train a model to self-reflect faster and stop wasting tokens on dead ends. This shifts the competitive edge for AI startups. It is no longer just about who has the smartest base model. It is about who has the most token-efficient inference stack for specific industry workloads.

For builders

Cheaper multi-turn agent loops

Multi-turn agents replay prior reasoning on every turn. Ember-1 cuts this quadratic cost growth by shortening the reasoning traces. Founders building coding agents pay significantly less for API calls while keeping high success rates.

Custom token-efficient models

Fireworks also launched training support for Ember-1. Enterprises can use their own data to build custom models tailored to their workloads. You pay upfront for training but save massively on inference at scale.

Two week research preview window

Ember-1 is currently a research preview on Serverless. Developers have two weeks to test it before Fireworks decides to make it permanent based on demand. Builders need to evaluate it quickly to lock in the cost savings.

My take

Brute force reasoning is a lazy tax on developers. Fireworks gets it exactly right by optimizing the thinking process itself instead of just turning down the effort slider. If your AI agent spends 90 percent of its time talking to itself, you are burning startup cash for absolutely no reason. Stop paying for unproductive AI loops and start optimizing your token spend.

Original reporting: Hacker News. This is my rewrite and opinion.

More AI news for builders