OpenAI Sold You Latency: Inside the 14× Cerebras Play
胡新宇
发布于 2026-08-19
On August 13, 2026, OpenAI shipped Ultrafast — GPT-5.6 Sol at 14× Standard speed, 750 tokens per second, running on Cerebras silicon instead of NVIDIA. Five named use cases (incident response, financial security, voice commerce, live research) tell you where the next dollar of AI revenue is going. Speed is no longer a cost line — it is a SKU. The NVIDIA inference monoculture is cracking at the latency-sensitive edges first.
OpenAI Sold You Latency: Inside the 14× Cerebras Play
On August 13, 2026, OpenAI launched Ultrafast — a service tier that runs GPT-5.6 Sol up to 14× faster than Standard, hitting 750 output tokens per second. The chip underneath is not from NVIDIA. It's from Cerebras (OpenAI).
For a decade, OpenAI has sold you capability. Smarter completions, longer context, better reasoning. Now, for the first time, it is selling you latency as a product line — and the five use cases it lists in the announcement (incident response, financial security, voice commerce, live research) tell you exactly where the next dollar of AI revenue is going to come from.
This is a smaller story than "the agent harness is the product." It is also a more important one — because Ultrafast is the cleanest proof in 2026 that the inference stack, not just the model, is now the product.
Speed is no longer a cost line. It is a SKU.
What Ultrafast actually is
The announcement is short. Read it once and the technical content is straightforward:
- GPT-5.6 Sol on Ultrafast mode generates up to 750 output tokens per second.
- That's up to 14× faster than Standard processing for the same model.
- The silicon is Cerebras, not NVIDIA.
- It's a limited preview today, expanding as capacity grows.
That last line is the boring one. The other three lines are the whole story.
Until August 13, "AI is fast" or "AI is slow" was a property of the model. Faster models meant smaller models — GPT-4-class intelligence capped out around 100-200 tokens per second on standard GPUs, and that was fine for chat but unusable for real-time voice or live incident response. The way you got real-time intelligence was by reaching for a smaller model — Haiku, Flash-Lite, GPT-5.4 nano — and accepting the capability hit.
Ultrafast breaks that tradeoff. You get GPT-5.6 Sol — OpenAI's most intelligent production model — at speeds previously reserved for tiny distilled models. The pitch is no longer "smarter" or "smaller." It is "frontier intelligence in real time."
