OpenAI on August 13 previewed Ultrafast, a new processing tier for GPT-5.6 Sol that generates up to 750 output tokens per second, the company said — as much as 14 times faster than the model’s standard setting, with no drop in output quality.

Built on Cerebras hardware

The speed comes from chips made by Cerebras, which is powering the new tier as OpenAI’s hardware partner. According to Cerebras, its Wafer-Scale Engine holds an entire model’s weights in 44GB of on-chip memory, letting tokens flow through the model’s layers without the back-and-forth data transfers that slow down chips split across multiple separate GPUs. The site’s own explainer on what a wafer-scale chip is covers how the design differs from a standard GPU.

On the Humanity’s Last Exam benchmark, OpenAI said GPT-5.6 Sol Ultrafast answered all 2,500 questions in just over 11 hours, versus more than three days of continuous compute for Anthropic’s Claude Fable 5 — reaching comparable accuracy roughly seven times faster, according to the company’s own testing. OpenAI also reported a 5.6x end-to-end speedup on the GDP-Val benchmark with no measured quality loss.

Limited preview for now

Ultrafast is available only to a limited set of OpenAI API customers for now, with the company testing it across coding, financial analysis, customer support, commerce, research and incident-response workloads. OpenAI has not disclosed pricing or a timeline for broader access.

The push for faster inference reflects a wider industry shift: as AI models get folded into live workflows — debugging code, answering support tickets, running trading analysis — how quickly a model responds is becoming as important as how well it answers.