OpenAI introduces Ultrafast mode for GPT-5.6 Sol with 14x speed
OpenAI has introduced Ultrafast, a preview API mode for its flagship GPT-5.6 Sol model, delivering up to 750 output tokens per second—roughly 14 times standard speed. Powered by a partnership with chipmaker Cerebras, the mode targets enterprise workflows like incident response, customer service, and financial analysis. The company emphasizes that speed is decoupled from model quality, a shift toward more useful work per second rather than picking smaller models. Currently limited to a select group of customers, access will expand as capacity grows. This reflects a growing trend where latency becomes a product differentiator, though benchmarks and exact performance claims remain vendor-attributed.
OpenAI introduces Ultrafast mode for GPT-5.6 Sol with 14x speed
OpenAI is launching a preview of a sped up version of its latest model, GPT-5.6 Sol, called Ultrafast, which delivers up to 750 output tokens per second. The mode is powered by a partnership with Cerebras and is initially available to a small group of customers.
Key takeaway
Speed has become a distinct product tier, decoupling latency from model intelligence and enabling real-time enterprise apps without sacrificing reasoning quality.
What happened
TechCrunch reports that OpenAI is preparing to launch 'Ultrafast,' a preview mode for its GPT-5.6 Sol model that runs up to 14 times faster than standard processing, generating up to 750 output tokens per second. The mode is designed to court enterprise users, targeting high-stakes workflows such as incident response, customer service, and financial market analysis.
The feature is powered by a partnership with chipmaker Cerebras, which claims the mode delivers 750 tokens per second without quality loss. In a blog post, Cerebras cites benchmarks showing Ultrafast is 11x faster than Fable 5 and 5x faster than Opus 4.8. OpenAI says the preview is initially available to a small group of customers, with expanded access as capacity grows.
Evidence
OpenAI is introducing Ultrafast, a preview mode for GPT-5.6 Sol that delivers up to 750 output tokens per second and runs at 14x speed.
TechCrunch AI · attributed
OpenAI is launching a preview of a sped up version of its latest, most powerful model, in an effort to court enterprise users.
Cerebras claims Ultrafast delivers 750 tokens per second without quality loss, and is 11x faster than Fable 5 and 5x faster than Opus 4.8.
Cerebras Blog · attributed
Cerebras claims GPT-5.6 Sol on Ultrafast mode delivers up to 750 output tokens per second without quality loss, citing benchmarks showing it is 11x faster than Fable 5 and 5x faster than Opus 4.8.
OpenAI says the preview is initially limited to a select group of customers, with access expanding as capacity grows.
OpenAI (X) · attributed
Previewing Ultrafast mode: GPT-5.6 Sol at up to 14x the speed. Launching first in the OpenAI API to a select group of customers with expanded access to more businesses as capacity grows.
Why it matters
This move pushes the frontier from raw accuracy to throughput, signaling that real-time AI applications will no longer be bottlenecked by inference latency, potentially reshaping enterprise deployment strategies and hardware economics.
Limits and uncertainties
Preview is limited to a select group, so real-world performance and pricing remain unverified.
Cerebras benchmarks are vendor-attributed and not independently audited.
Practical implications
Builders should monitor OpenAI's API for the new Ultrafast tier and evaluate latency-sensitive workloads against cost.
The partnership suggests specialized hardware like Cerebras could become a standard for high-speed inference, influencing future infrastructure choices.
What to watch
Expansion of Ultrafast access to more customers and any official pricing structure.
Independent benchmarks or third-party validation of the 750 tokens/sec claim.