OpenAI and Cerebras Launch 'Ultrafast Mode' for API, Reaching 750 Tokens Per Second
Powered by Cerebras' Wafer-Scale Engine, the new high-speed tier aims to resolve latency bottlenecks for enterprise AI workloads without degrading model quality.

OpenAI and chip startup Cerebras Systems have introduced an early preview of "Ultrafast Mode," a high-speed inference tier integrated into the OpenAI API and driven by specialized wafer-scale silicon, as first reported by Hacker News. The specialized mode powers OpenAI's GPT-5.6 Sol model, achieving output throughput speeds of up to 750 tokens per second without compromising accuracy or output quality.
The announcement directly addresses a longstanding trade-off in artificial intelligence deployment between operational processing speed and reasoning capabilities. As state-of-the-art neural networks increase in parameter scale, the computational overhead and high memory transfer demands slow down execution times, forcing enterprise organizations to choose between delayed responses or reduced task accuracy.
According to performance metrics cited by Cerebras, GPT-5.6 Sol operating on Ultrafast Mode demonstrates substantial speed improvements over competing hardware setups evaluated by third-party benchmark provider Artificial Analysis. The system executes output generation 11 times faster than Anthropic's Claude Fable 5 and achieves a fivefold speed advantage over Opus 4.8 running in its Fast configuration.
In standardized benchmark evaluations performed on Humanity's Last Exam—a rigorous dataset containing 2,500 questions across specialized academic fields such as literature, chemistry, and economics—GPT-5.6 Sol Ultrafast completed the entire test suite in 11 hours and 11 minutes. In comparison, Claude Fable 5 required 78 hours and 27 minutes to solve the identical prompt set during testing conducted in July 2026 using maximum reasoning settings.
Testing on the GDP-Val benchmark, which measures capability on high-value business and professional workflows, demonstrated a 5.6-fold increase in end-to-end task completion speeds while preserving full output fidelity. The companies indicated that the acceleration targeted time-sensitive applications, including root-cause analysis for enterprise IT outages, immediate cyberattack mitigation, complex legal brief generation, and real-time financial modeling.
The hardware architecture driving the service utilizes Cerebras' Wafer-Scale Engine, a specialized processor designed to eliminate memory bandwidth constraints common in traditional graphics processing units. Conventional GPU clusters often bottleneck during frontier model execution due to the continuous shuttling of model weights between external memory modules and compute cores; Cerebras avoids this issue by embedding 44 gigabytes of SRAM directly onto each wafer, allowing parameters to remain on-chip as tokens stream through pipelined processor layers.
Ultrafast Mode is currently accessible in a restricted preview for a limited initial cohort of OpenAI API subscribers, with broader distribution slated to roll out as datacenter computing capacity expands over time. Enterprise customers will be able to designate Ultrafast instances for time-critical automated agents while using standard processing pipelines for secondary administrative workflows.
Sources
Written by
The Company Wire
Inside the companies building what’s next. Reporting on startups, technology, funding and the people shaping them.



