Microsoft launches Decision-1: The AI that can't talk, only decides
Microsoft just dropped a new AI model that refuses to chat. It only makes structured decisions, and it does it 35x faster than GPT-6 Sol.
Models · Source: TipRanks
What happened
Microsoft just released Microsoft-Decision-1. It is a specialized artificial intelligence model that skips text generation entirely. Instead, it focuses purely on high-speed, calibrated scoring for automated classification, routing, and workflow governance. Think of it as a referee that only outputs yes, no, or a specific category with a probability score.
The speed of this system is the main feature. Evaluated across 36 blind benchmarks and nearly 150,000 questions, it ran 35 times faster than traditional large language models like GPT-6 Sol. It also beat its closest decision-model competitor, Quyet-1.0-Large, by 4.5 times. Microsoft achieved this performance by post-training a 9-billion-parameter base model called Qwen3.5-9B.
Internal Microsoft teams are already using this model to cut latency in their applications. Xbox Research categorized over 10,000 player reviews into structured categories 14 times faster than their previous setup. The Microsoft Copilot team used it for quality assurance checks on interactive agent responses, hitting speeds 100 times faster than before. The model is available now on Microsoft Foundry and is coming soon to OpenRouter.
Key facts
- 35x — faster execution than flagship text models like GPT-6 Sol
- $0.042 — cost per million input tokens, with zero cost for output tokens
- 1.3% — rate of output change when requests are paraphrased or re-sorted
- 100x — speed increase for the Copilot team running quality assurance checks
- Qwen3.5-9B — the 9-billion-parameter base model post-trained by Microsoft engineers
Why it matters
AI agents make hundreds of micro-decisions per task to function properly. Using a massive conversational model like GPT-6 Sol to decide if a payment should be approved or where a support ticket goes is a massive waste of time and money. Accumulated delay kills multi-step AI operations because slight lags at each step slow down entire applications. By offloading these routing and classification tasks to a specialized decision model, developers can eliminate latency bottlenecks and drastically reduce compute costs.
The era of using one giant brain for every single task is ending. We are moving toward systems where a network of tiny, ultra-fast specialist models handles the workflow, while the heavy language models are saved for actual reasoning or text generation. Interestingly, Microsoft built this first iteration on Qwen3.5-9B rather than their own in-house foundation models. They plan to add OpenAI technology in future versions, signaling that the fast-decision architecture matters more than the base model provider.
For builders
Cut agent latency with specialized routing
Stop using conversational large language models for multiple-choice or pass-fail logic. Swap them out for Decision-1 to speed up your multi-step workflows. Your users pay with their attention when apps lag, and you lose money on unnecessary compute.
Free output tokens lower your API bills
The model costs $0.042 per million input tokens, and output tokens are completely free. This pricing structure heavily favors developers building high-volume classification systems. You win by routing massive datasets through this instead of standard models.
Rethink prompt engineering for stability
Decision-1 changed its output in only 1.3 percent of cases when prompts were paraphrased or options were re-sorted. This means you spend less time tweaking prompts to get consistent structured outputs. Builders save engineering hours previously lost to prompt manipulation.
My take
I love that Microsoft swallowed their pride and used an open model like Qwen3.5-9B to build this instead of forcing an OpenAI wrapper. It proves that for agentic workflows, speed and structure matter way more than having the smartest conversational brain. Stop using massive models to make simple yes-or-no decisions in your apps.
Original reporting: TipRanks. This is my rewrite and opinion.