OpenAI launches Ultrafast tier for GPT-6.1 Sol at six times the price

OpenAI just turned speed into a premium lane. You get 8x faster output for GPT-6.1 Sol, but you pay 6x the standard API price.

Business · Source: OpenAI Developer Forum

What happened

OpenAI rolled out an Ultrafast service tier for GPT-6.1 Sol across its API, Codex, and ChatGPT Work platforms. The new mode delivers token generation up to eight times faster than standard speeds in Codex and six times faster in the API. It targets latency-sensitive tasks where every second counts. OpenAI specifically calls out debugging system outages, autonomous agents navigating applications, and live user experiences. The tier also supports voice inputs and computer use capabilities.

Speed comes with a massive markup. API pricing for the Ultrafast tier sits at $12 per million input tokens and $60 per million output tokens. That is exactly six times the cost of standard GPT-6.1 Sol rates, which are $2 for input and $10 for output. Cached inputs cost $0.60, while cache writes run $15. OpenAI notes this pricing makes Ultrafast just 1.2 times the cost of their heavier Astra model. Regional processing adds another ten percent premium where available.

Access depends entirely on your platform and subscription level. The API tier is available to all developers immediately. However, Codex and ChatGPT Work users face strict gates. You must be on a Pro 500, eligible usage-based Enterprise, or credit-based Edu plan to flip the switch. Enterprise administrators must manually enable the feature. The rollout also includes support for US and EU data residency, alongside global processing.

Key facts

Why it matters

Inference latency is now a premium line item on your cloud bill. You can no longer just pick a model based on its baseline intelligence and call it a day. You have to route traffic based on how long your user is willing to wait. Teams must split their workloads aggressively. Background tasks and batch processing should stay on the standard tier. You must reserve the expensive Ultrafast lane strictly for live user interactions or rapid agent loops where latency directly impacts user retention.

This pricing structure fundamentally changes the math for building agentic applications. Agents making successive tool calls need fast responses to function smoothly without frustrating the user. But running those loops at a six times markup will burn through startup budgets rapidly. Developers will need to optimize their network overhead just to ensure they actually feel the speed they are paying for. If your infrastructure is slow, paying OpenAI for faster token generation is a complete waste of money.

For builders

Route traffic to control API costs

Do not send all your GPT-6.1 Sol traffic through the Ultrafast tier. Background processing, data extraction, and batch jobs should stay on the standard tier to save money. Only pay the six times premium for live user chats or time-critical agent actions.

Use WebSockets for agentic applications

Network overhead can kill the latency gains you just paid a premium to get. OpenAI explicitly recommends using persistent WebSocket connections for agents making successive tool calls. If you stick to standard HTTP requests, network round trips might erase the speed advantage.

Check your workspace plan limits

API users get immediate access, but interface users are heavily gated. If you want Ultrafast in Codex or ChatGPT Work, you need a Pro 500, usage-based Enterprise, or credit-based Edu plan. Smaller teams on lower tiers are locked out of the interface speed boost.

My take

OpenAI is finally admitting that speed is a distinct feature worth charging for. Selling the exact same model at a 600 percent markup just for faster token generation is a bold move. It forces founders to decide exactly how much a single second of user attention is worth.

Original reporting: OpenAI Developer Forum. This is my rewrite and opinion.

More AI news for builders