OpenAI on Thursday unveiled Ultrafast, a new mode for its flagship model GPT 5.6 Sol that the company says operates at 14 times the speed of standard processing, delivering up to 750 output tokens per second. The announcement, made via a company blog post, positions the feature as a significant step toward real-time AI interactions without sacrificing model capability.
“Until now, getting real-time speed typically meant choosing a smaller or more specialized model,” the company said in the post. “Ultrafast points to progress in a new direction: more useful work per second.”
Also read: Writer launches Palmyra X6 and a smarter harness to slash enterprise AI token costs
How Ultrafast works and who gets it first
Ultrafast is powered by OpenAI’s collaboration with Cerebras, a chipmaker known for its wafer-scale engines designed to accelerate AI inference. By employing Cerebras’ specialized hardware, OpenAI aims to deliver near-instantaneous responses from GPT 5.6 Sol, its most advanced model to date.
The preview is currently limited to a small group of customers, though OpenAI indicated it will broaden availability “as capacity grows.” The company did not specify a timeline for wider rollout or pricing details.
Also read: IBM expands OpenAI partnership to push enterprise AI adoption
According to OpenAI, the high-speed mode is tailored for several enterprise workflows where latency is critical:
- Incident response — rapid triage and analysis during security or operational events
- Customer service and support — faster, more natural conversational interactions
- Financial market analysis — real-time data interpretation and alerts
- E-commerce — instant product recommendations and query handling
These use cases underscore OpenAI’s push to position GPT 5.6 Sol not just as a chatbot but as a low-latency engine for mission-critical business applications.
Competitive pressure and industry context
OpenAI is not alone in chasing faster inference. Anthropic, a key rival, has introduced a “fast mode” for its Claude models, though the company has not disclosed comparable speed metrics. The race to reduce response latency has become a central battleground as enterprises increasingly deploy AI agents in real-time environments where every second matters.
The move also signals a broader shift in AI infrastructure. While model quality remains a differentiator, the ability to run large models at high speed on specialized silicon is becoming equally important. Cerebras’ partnership with OpenAI highlights how chipmakers are carving out a role in the AI supply chain beyond traditional GPU providers like Nvidia.
For developers and enterprises, Ultrafast could lower the barrier to using GPT 5.6 Sol in time-sensitive applications that previously required smaller, faster models. If the speed claims hold up in production, it may also change expectations for what “real-time” AI means in customer-facing products.
OpenAI’s announcement comes amid a period of executive reshuffling, including the recent hire of a new chief revenue officer. The company continues to expand its enterprise offerings while managing the operational challenges of scaling AI infrastructure.
As Ultrafast moves from preview to broader availability, observers will be watching how it performs under real-world workloads and whether competitors respond with similar speed-focused upgrades. For now, the feature represents a notable technical milestone in the ongoing effort to make large language models faster without compromising their capabilities.
Disclaimer: This article is for informational purposes only and does not constitute financial advice. The AI industry is volatile and rapidly evolving; readers should conduct their own research before making any business or investment decisions.

Be the first to comment