OpenAI's latest Ultrafast mode operates GPT-5.6 Sol at a speed 14 times quicker on Cerebras chips.

OpenAI's latest Ultrafast mode operates GPT-5.6 Sol at a speed 14 times quicker on Cerebras chips.

      OpenAI aims to make its most advanced model also the fastest. The company has introduced Ultrafast, a new tier of its API that operates the flagship GPT-5.6 Sol at speeds up to 14 times greater than usual, achieving around 750 output tokens per second, using hardware developed by Cerebras, a chipmaker known for wafer-scale technology.

      Ultrafast is less about a new model and more about a different method of utilizing an existing one. It utilizes Cerebras’s large chips to eliminate the latency that has historically plagued cutting-edge AI, and OpenAI offered a limited preview to a select group of customers on August 13, with intentions to expand availability based on capacity.

      The proposal revolves around a trade-off that OpenAI claims it can now overcome. Previously, those seeking genuinely real-time responses had to opt for smaller, less capable models, sacrificing intelligence for speed. Ultrafast aims to provide top-tier reasoning and near-instant responses simultaneously, rather than forcing users to choose between them.

      This combination is particularly significant for the intelligent software that the entire industry is pursuing. An AI agent that takes thirty seconds to contemplate each action is more of a demonstration, while one that responds quickly like a conversation feels more like a product.

      In this context, speed is becoming as integral a feature as sheer intelligence. OpenAI is targeting this tier specifically for time-sensitive tasks, such as incident response, debugging, financial research, fraud detection, real-time customer support, and e-commerce.

      Initial users, including Jane Street, Podium, Basis, and Rogo, describe the change as qualitatively different rather than just a minor improvement. One tester noted that the increased speed “entirely transforms the call experience for complex tasks” and enables “synchronous experiences for users that were previously hampered by intelligence limitations.”

      For Cerebras, this partnership serves as a significant endorsement at a critical time. The company recently went public in one of the year's largest listings but has faced challenges in demonstrating that its wafer-scale aspirations can lead to consistent profits. Thus, powering OpenAI’s fastest tier is the kind of validation it seeks.

      This also highlights that unconventional chip architectures, once viewed as experimental, are now successfully operating for major AI players. The remarkable speed is attributed to cutting-edge engineering. Cerebras designs processors that are the size of dinner plates, made from a single silicon wafer, allowing an entire model to fit on one chip instead of being distributed across multiple Nvidia GPUs that need to exchange data frequently.

      Eliminating this internal data transfer reduces the delay between prompts and responses, supporting the company's ongoing argument that its architecture is better suited for inference than general-purpose chips designed for training.

      This development aligns with a wider trend. As basic model capabilities begin to level off, the focus is shifting toward who can run those models the fastest and most economically. This competitive landscape has elevated inference specialists like Groq, turning latency into a marketable asset.

      Competitors such as SambaNova and several specialized cloud providers are similarly aiming for this advantage. The market increasingly favors those who can generate responses the quickest, rather than merely the entities that developed the largest models.

      For OpenAI, collaborating with Cerebras also represents a subtle move away from total reliance on Nvidia, in line with its efforts on its own custom silicon. While OpenAI has not disclosed pricing details, running its top model at 14 times the usual speed on specialized hardware is expected to be costly. Therefore, Ultrafast may remain a premium choice for latency-sensitive applications rather than a standard offering.

      Currently, it is only in preview for a limited number of clients while OpenAI seeks to increase its capacity. Nonetheless, the underlying message is clear: in the next stage of the AI race, intelligence alone will not suffice if it comes with slowness.

Other articles

Silver Lake is said to be in discussions to acquire Workday and take it private for $43 billion. Silver Lake is said to be in discussions to acquire Workday and take it private for $43 billion. Silver Lake is said to be negotiating to acquire Workday and take it private for approximately $43 billion, which would make it one of the biggest software buyouts in history and the first significant take-private deal involving a prominent SaaS company. Apple developed its own AI model for China and then provided the technology to Alibaba. Apple developed its own AI model for China and then provided the technology to Alibaba. Apple, with assistance from Alibaba, has developed a large language model exclusively for China, making it the first foreign company authorized to provide a proprietary AI model in mainland China. Investors are suing Selena Gomez, alleging that her wellness startup's app was never developed. Investors are suing Selena Gomez, alleging that her wellness startup's app was never developed. Investors have filed a lawsuit against Selena Gomez and her mother, who is also a co-founder, regarding their mental health startup, Wondermind. They are accusing them of securities fraud, asserting that the promised app was never developed. These allegations have not been substantiated. Apple developed its own AI model specifically for China and handed it over to Alibaba. Apple developed its own AI model specifically for China and handed it over to Alibaba. Apple has developed a large language model specifically for China with assistance from Alibaba, making it the first foreign company authorized to provide a proprietary AI model in mainland China. Goldman Sachs is reaching out to investors regarding Nvidia's $500 billion financing arrangement for AI computing. Goldman Sachs is reaching out to investors regarding Nvidia's $500 billion financing arrangement for AI computing. Goldman Sachs is seeking investors for Nvidia's $500 billion AI-compute financing arrangement after securing a leading position, and it aims to develop chips into a tradable asset class akin to bonds. Silver Lake is said to be negotiating to acquire Workday for $43 billion in a deal to take the company private. Silver Lake is said to be negotiating to acquire Workday for $43 billion in a deal to take the company private. Silver Lake is said to be negotiating to acquire Workday and take it private for about $43 billion, which would mark one of the largest software buyouts in history and the first major take-private deal involving a prominent SaaS company.

OpenAI's latest Ultrafast mode operates GPT-5.6 Sol at a speed 14 times quicker on Cerebras chips.

OpenAI has showcased Ultrafast, an API tier capable of operating GPT-5.6 Sol at speeds up to 14 times faster on Cerebras hardware, asserting that low latency, rather than solely intelligence, contributes to the usability of AI agents.