OpenAI’s new Ultrafast mode operates GPT-5.6 Sol at a speed 14 times quicker, utilizing Cerebras chips.

OpenAI’s new Ultrafast mode operates GPT-5.6 Sol at a speed 14 times quicker, utilizing Cerebras chips.

      OpenAI aims for its most advanced model to also be its fastest. The company has introduced Ultrafast, a new API tier that operates the flagship GPT-5.6 Sol at speeds up to 14 times higher than usual, achieving approximately 750 output tokens per second, using hardware created by the wafer-scale chip manufacturer Cerebras.

      Ultrafast represents a novel method of delivering an existing model rather than a new model itself. It leverages Cerebras’s large chips to eliminate the latency that has long plagued cutting-edge AI, and OpenAI initiated a limited preview on August 13 for a select group of customers, with intentions to expand access as capacity permits.

      The focus is on a trade-off that OpenAI claims it can finally eliminate. Previously, anyone seeking truly real-time responses had to resort to a smaller, less capable model, sacrificing intelligence for speed.

      Ultrafast is designed to offer high-level reasoning and near-instant responses simultaneously, rather than making users choose between the two.

      This feature is particularly significant for the intelligent software that the entire industry is striving to develop. An AI agent that takes thirty seconds to process each step serves as a demonstration, while one that responds as quickly as a conversation feels more like a viable product.

      In essence, speed is increasingly becoming as crucial a feature as sheer intelligence.

      OpenAI is specifically targeting this tier at time-sensitive applications, including incident response and debugging, financial research and fraud detection, real-time customer support and voice services, as well as e-commerce.

      Early users like Jane Street, Podium, Basis, and Rogo describe the change as qualitative rather than merely incremental, with one stating that the speed “completely transforms the call experience for complex tasks” and enables “synchronous experiences for users that were previously constrained by intelligence.”

      For Cerebras, this partnership serves as a significant endorsement at a pivotal time. The company went public in one of the year’s largest offerings but has struggled to prove to the market that its wafer-scale ambitions can yield sustainable profits, making OpenAI’s fastest tier the validation it needed.

      It also highlights that the innovative chip architectures once regarded as experimental are now providing tangible support for the leading figures in AI.

      The remarkable speed is achieved through an innovative engineering approach. Cerebras manufactures processors the size of dinner plates from a single silicon wafer, allowing an entire model to reside on one chip instead of being divided among multiple Nvidia GPUs that must continually transfer data between them.

      This design minimizes internal data traffic, which significantly reduces the delay between a prompt and a response, and reinforces the company's long-standing assertion that its architecture is more suitable for inference than conventional chips intended for training.

      This development aligns with a broader trend as raw model capability nears a plateau, shifting the competition to those who can execute these models the fastest and most cost-effectively. This competitive landscape has benefited inference specialists like Groq and made low latency a selling point.

      Competitors such as SambaNova and several specialized cloud providers are vying for the same market, and the industry is increasingly rewarding those who can deliver a model’s responses the quickest, rather than just the ones who developed the largest models.

      For OpenAI, relying on Cerebras also represents a gradual move away from complete dependence on Nvidia, complementing its efforts to develop custom silicon.

      OpenAI has yet to disclose pricing, and operating its leading model at 14 times the normal speed on specialized hardware is expected to be costly, suggesting that Ultrafast may remain a premium choice for latency-sensitive applications rather than the standard option.

      It is also still in the preview stage, limited to a select few customers while OpenAI seeks to expand its capacity. Nonetheless, the underlying message is clear: in the next stage of the AI competition, being intelligent won’t matter much if your performance is slow.

Other articles

Goldman Sachs is seeking investors for Nvidia’s $500 billion AI-compute financing arrangement. Goldman Sachs is seeking investors for Nvidia’s $500 billion AI-compute financing arrangement. Goldman Sachs is seeking to attract investors for Nvidia's $500 billion AI-compute financing arrangement after securing a leading position, and it aims to transform chips into a tradable asset class similar to bonds. Apple developed its own AI model specifically for China and handed it over to Alibaba. Apple developed its own AI model specifically for China and handed it over to Alibaba. Apple has developed a large language model specifically for China with assistance from Alibaba, making it the first foreign company authorized to provide a proprietary AI model in mainland China. This startup in Bengaluru is training dogs and artificial intelligence to detect cancer at an early stage. This startup in Bengaluru is training dogs and artificial intelligence to detect cancer at an early stage. Dognosis, a startup based in Bengaluru, combines trained sniffer dogs with artificial intelligence to detect early signs of cancer from a breath sample. It serves as a prescreening tool currently undergoing trials and is not yet a validated diagnostic method. Pixel 11 Pro Fold vs. Galaxy Z Fold 8 Ultra: Google takes a cautious approach while Samsung fully commits. Pixel 11 Pro Fold vs. Galaxy Z Fold 8 Ultra: Google takes a cautious approach while Samsung fully commits. The Pixel 11 Pro Fold and Galaxy Z Fold 8 Ultra pursue different approaches to achieve a similar objective. Here’s a comparison to help you decide which foldable suits you better. Silver Lake is said to be negotiating to acquire Workday for $43 billion in a deal to take the company private. Silver Lake is said to be negotiating to acquire Workday for $43 billion in a deal to take the company private. Silver Lake is said to be negotiating to acquire Workday and take it private for about $43 billion, which would mark one of the largest software buyouts in history and the first major take-private deal involving a prominent SaaS company. Goldman Sachs is seeking investors for Nvidia's $500 billion financing agreement related to AI computing. Goldman Sachs is seeking investors for Nvidia's $500 billion financing agreement related to AI computing. Goldman Sachs is seeking investors for Nvidia's $500 billion AI-compute financing arrangement after securing a leadership position, with the aim of transforming chips into a tradable asset class similar to bonds.

OpenAI’s new Ultrafast mode operates GPT-5.6 Sol at a speed 14 times quicker, utilizing Cerebras chips.

OpenAI has unveiled Ultrafast, an API tier that operates GPT-5.6 Sol at speeds of up to 14 times faster on Cerebras hardware, claiming that latency, in addition to intelligence, is crucial for the usability of AI agents.