OpenAI’s New ‘Ultrafast’ Mode Runs GPT-5.6 at 14 Times Normal Speed

The complaint from enterprise customers has been consistent for two years: the models are powerful, but they are slow. OpenAI addressed that complaint head-on on Aug. 13, introducing an Ultrafast preview mode for its flagship model, GPT-5.6 Sol, that generates output at up to 750 tokens per second, roughly 14 times the speed of the regular mode.

The speed comes from a new supplier. Cerebras, the chip maker that builds wafer-scale processors, is providing the compute, a pairing that was described by people familiar with the arrangement as a significant expansion of an existing relationship. The choice is notable because Cerebras’ stock fell sharply the previous day after its own earnings report, and the OpenAI announcement gave investors a reminder of why the company exists in the first place.

The Ultrafast mode is aimed at the market OpenAI most wants to win: enterprises deploying AI in production. For customers running customer-service agents, coding assistants and data-processing pipelines, latency is not a nicety; it is the difference between a tool employees use and a tool they abandon. OpenAI’s pitch is that Ultrafast makes its models feel instant, the way a search engine or a database does, rather than like a clever but slow intern.

The technical setup is as interesting as the product. Cerebras builds a chip the size of a dinner plate, with hundreds of thousands of cores on a single wafer, an architecture that is unconventional but well suited to the specific task of generating tokens quickly. The company has spent years trying to sell that hardware to AI labs, and landing OpenAI as a visible customer validates the approach. The deal gives Cerebras a marquee name at a moment when it needs one.

The timing of the announcement, a day after Cerebras’ disappointing earnings, was either unfortunate or deliberate, depending on whom you ask. The stock’s slide reflected concerns about Cerebras’ own revenue growth and its dependence on a small number of customers, and the OpenAI news is the kind of development that can rebuild confidence. Analysts said the contract, if it expands beyond the preview, would be among the most important in Cerebras’ history.

For OpenAI, the move is part of a broader strategy of diversifying its compute. The company has built most of its infrastructure on Nvidia chips, but it has been clear that relying on a single supplier is both a risk and a negotiating weakness. Cerebras offers an alternative for the specific workload of fast inference, and the Ultrafast mode is a way to test how much of OpenAI’s traffic can run on non-Nvidia hardware.

Inference speed has become the industry’s newest battleground. As models have converged in capability, the differentiators that remain are price, speed and reliability, and every major lab has been racing to make its models faster. OpenAI’s 750 tokens per second is an order of magnitude beyond what most production systems deliver, and rivals, including Google and Anthropic, are expected to answer with speed improvements of their own.

The Ultrafast mode is a preview, which means it is available to a limited set of enterprise customers rather than to everyone. OpenAI has used the preview pattern before, shipping features early to gather feedback and iterate, and the company said pricing for the fast mode will be announced as it scales. The economics matter: faster generation consumes more compute per second, and OpenAI will need to price Ultrafast at a premium that reflects its cost.

The product also signals where OpenAI thinks the AI market is heading. The company has said that the future of AI is agentic, with models performing tasks autonomously rather than answering single questions, and agents are brutally demanding on latency. An agent that must make dozens of model calls to complete one task needs each call to be fast, or the whole workflow crawls. Ultrafast is, in that sense, infrastructure for the agent era.

Enterprise customers have responded with interest, according to people familiar with the company’s sales pipeline, particularly in financial services and customer operations, where speed translates directly into cost savings. A call center that can generate responses in real time handles more interactions per agent, and a trading desk that can summarize documents instantly gains a competitive edge. The use cases that were impossible at yesterday’s speeds become routine at 750 tokens per second.

The announcement closes a week in which OpenAI was in the news for reasons other than speed, including the appointment of a new revenue chief and continued speculation about its path to an IPO. The Ultrafast launch refocuses the conversation on product, which is where OpenAI prefers it. The model was already among the most capable in the world; now it is also among the fastest, and the company’s pitch to enterprise customers has a new number to lead with: 750 tokens per second, fourteen times the old pace, powered by chips that most of the industry said would never matter.

Related Posts

  • September 6, 2026
  • 6 views
Anthropic Moves Its IPO Filing to Late September

The bankers and lawyers running Anthropic’s initial public offering had told investors to expect the company’s registration documents as soon as this week. The calendar has moved. Anthropic now plans…

  • September 6, 2026
  • 6 views
OpenAI Quietly Revises GPT-6 Astra Scores After Launch

When OpenAI released GPT-6 Astra on Sept. 3, the launch post carried the usual furniture of a modern model debut: coding results, speed comparisons and a figure for how often…