AIO APEX

OpenAI's Ultrafast mode runs GPT-5.6 Sol at 750 tokens per second

TechCrunch
Share:
OpenAI's Ultrafast mode runs GPT-5.6 Sol at 750 tokens per second

OpenAI has launched Ultrafast, a new API service tier that runs GPT-5.6 Sol at up to 750 output tokens per second — 14 times the speed of the current Standard mode, which averages around 53 tokens per second. The mode is available in limited preview to select API customers starting August 13.

The speed gains come from a partnership with Cerebras, the AI chip startup known for its wafer-scale processors. Ultrafast is the first commercial product of that collaboration, using Cerebras hardware to dramatically accelerate inference on OpenAI's flagship model without, according to the company, any reduction in output quality.

What 750 tokens per second actually means

The difference between 53 and 750 tokens per second is the difference between a model that types as you read versus one that delivers a full response almost instantly. A typical 1,000-word article takes roughly 18 seconds at Standard speed and under 2 seconds at Ultrafast. For real-time applications, that gap is enormous.

OpenAI says Ultrafast is designed for workflows that have historically forced a speed-versus-intelligence tradeoff — where developers chose smaller, faster models because frontier models were too slow. The pitch is that you can now run GPT-5.6 Sol, a frontier-grade model, at speeds previously only achievable with stripped-down alternatives.

Target use cases

OpenAI highlights four core applications where sub-second inference unlocks new possibilities: incident response (security analysts getting live answers mid-investigation), customer service (agents that match the pace of a phone conversation), financial market analysis (real-time commentary on price movements), and e-commerce (dynamic product recommendations during active browsing sessions).

These are categories where major enterprises currently route queries away from frontier models specifically because of latency constraints. If Ultrafast holds its quality benchmark claims at scale, it removes that constraint.

The Cerebras connection

Cerebras has spent years building wafer-scale chips as an alternative to Nvidia's GPU approach. Its CS-3 processor is a single 900mm² die rather than a cluster of smaller chips, trading manufacturing complexity for raw inference throughput. This partnership is the highest-profile commercial deployment of Cerebras hardware to date.

OpenAI hasn't announced separate pricing for Ultrafast mode and is expanding access as capacity grows. As first reported by TechCrunch, the announcement came directly from OpenAI's developer blog.

Originally reported by TechCrunch. Read the original article for additional details.

View original source
Share: