OpenAI unveils Ultrafast mode for GPT-5.6 Sol in preview
- Maxime Hiez
- OpenAI
- 31 Aug, 2026
Introduction
OpenAI announced on August 13, 2026 Ultrafast, a new API service tier that runs GPT-5.6 Sol up to 14 times faster than standard processing. The model is the same, with the same capabilities, only the inference speed changes.
Check the GPT-5.6 preview article HERE.
The announced figures
Ultrafast reaches up to 750 output tokens per second, against a baseline of around 53 tokens per second in standard processing. In the published comparisons, a task is completed around 7 times faster than with Fable 5 in fast mode, and around 11 times faster than with Fable 5 at standard speed.
The performance does not come from a lighter model but from the hardware. OpenAI relies on a partnership with Cerebras and its wafer-scale architecture, an approach where the processor occupies an entire silicon wafer instead of being cut into individual chips.
What it changes in practice
OpenAI’s argument fits in one sentence from the announcement : “Until now, getting real-time speed typically meant choosing a smaller or more specialized model. Ultrafast points to progress in a new direction : more useful work per second.”
This is the usual trade-off disappearing. Until now, getting a real-time answer meant moving down the range, with the associated loss of quality. Ultrafast keeps the intelligence of the flagship model while removing the wait.
The use cases put forward are those where latency is blocking :
- Incident response : Analysis and recommendation while the incident is still ongoing
- Customer service and support : Conversation with no perceptible dead time
- Financial market analysis : Processing at the speed information circulates
- Online commerce : Personalization computed during the purchase journey
Availability and pricing
Ultrafast starts in preview in the API only, with a restricted group of customers selected in the development, commerce, financial research and customer support fields. OpenAI announces a broadening as capacity ramps up, with no timeline.
No pricing is published. OpenAI positions Ultrafast above the Fast mode, which already applies a per-token premium over standard processing. The actual additional cost therefore remains to be determined, and it is the point that will decide whether the service tier is worth it for most projects.
note
Conclusion
Ultrafast moves the competition between providers from a quality battleground to a latency one, on models whose capability gaps are narrowing. The absence of published pricing prevents any decision at this stage, and the closed preview rules out most organizations. The signal remains interesting : relying on specialized hardware for inference, rather than on reduced models, is becoming a credible answer to real-time constraints.
Sources
Did you enjoy this post ? If you have any questions, comments or suggestions, please feel free to send me a message from the contact form.
Don’t forget to follow us and share this post.