LLMgram · AI News · 2026-08-13

OpenAI introduces Ultrafast mode for GPT-5.6 Sol with 14x speed

OpenAI introduces Ultrafast mode for GPT-5.6 Sol with 14x speed

OpenAI has introduced Ultrafast, a preview API mode for its flagship GPT-5.6 Sol model, delivering up to 750 output tokens per second—roughly 14 times standard speed. Powered by a partnership with chipmaker Cerebras, the mode targets enterprise workflows like incident response, customer service, and financial analysis. The company emphasizes that speed is decoupled from model quality, a shift toward more useful work per second rather than picking smaller models. Currently limited to a select group of customers, access will expand as capacity grows. This reflects a growing trend where latency becomes a product differentiator, though benchmarks and exact performance claims remain vendor-attributed.

Sources

OpenAI introduces Ultrafast mode for GPT-5.6 Sol with 14x speed

OpenAI introduces Ultrafast mode for GPT-5.6 Sol with 14x speed

OpenAI is launching a preview of a sped up version of its latest model, GPT-5.6 Sol, called Ultrafast, which delivers up to 750 output tokens per second. The mode is powered by a partnership with Cerebras and is initially available to a small group of customers.

Key takeaway

Speed has become a distinct product tier, decoupling latency from model intelligence and enabling real-time enterprise apps without sacrificing reasoning quality.

What happened

TechCrunch reports that OpenAI is preparing to launch 'Ultrafast,' a preview mode for its GPT-5.6 Sol model that runs up to 14 times faster than standard processing, generating up to 750 output tokens per second. The mode is designed to court enterprise users, targeting high-stakes workflows such as incident response, customer service, and financial market analysis.

The feature is powered by a partnership with chipmaker Cerebras, which claims the mode delivers 750 tokens per second without quality loss. In a blog post, Cerebras cites benchmarks showing Ultrafast is 11x faster than Fable 5 and 5x faster than Opus 4.8. OpenAI says the preview is initially available to a small group of customers, with expanded access as capacity grows.

Evidence

  • OpenAI is introducing Ultrafast, a preview mode for GPT-5.6 Sol that delivers up to 750 output tokens per second and runs at 14x speed.

    TechCrunch AI · attributed

    OpenAI is launching a preview of a sped up version of its latest, most powerful model, in an effort to court enterprise users.

  • Cerebras claims Ultrafast delivers 750 tokens per second without quality loss, and is 11x faster than Fable 5 and 5x faster than Opus 4.8.

    Cerebras Blog · attributed

    Cerebras claims GPT-5.6 Sol on Ultrafast mode delivers up to 750 output tokens per second without quality loss, citing benchmarks showing it is 11x faster than Fable 5 and 5x faster than Opus 4.8.

  • OpenAI says the preview is initially limited to a select group of customers, with access expanding as capacity grows.

    OpenAI (X) · attributed

    Previewing Ultrafast mode: GPT-5.6 Sol at up to 14x the speed. Launching first in the OpenAI API to a select group of customers with expanded access to more businesses as capacity grows.

Why it matters

This move pushes the frontier from raw accuracy to throughput, signaling that real-time AI applications will no longer be bottlenecked by inference latency, potentially reshaping enterprise deployment strategies and hardware economics.

Limits and uncertainties

Preview is limited to a select group, so real-world performance and pricing remain unverified.

Cerebras benchmarks are vendor-attributed and not independently audited.

Practical implications

Builders should monitor OpenAI's API for the new Ultrafast tier and evaluate latency-sensitive workloads against cost.

The partnership suggests specialized hardware like Cerebras could become a standard for high-speed inference, influencing future infrastructure choices.

What to watch

Expansion of Ultrafast access to more customers and any official pricing structure.

Independent benchmarks or third-party validation of the 750 tokens/sec claim.

Sources

LLMgram editorial selection and synthesis · @llmgram. LLMgram is not the original publisher of this information.
Continue on LLMgram: Open in AI Signal →
Original reporting: OpenAI introduces ‘Ultrafast,’ a new mode that makes GPT 5.6 Sol work at 14x the speed