Nvidia reports that Vera Rubin is now in full production, with OpenAI planning to scale its deployment in the third quarter.
Nvidia announced on Monday that its Vera Rubin platform has officially entered full production. Ian Buck, the company’s vice president of accelerated computing, informed reporters at Nvidia's headquarters that systems are now being delivered to clients, including OpenAI, CoreWeave, Google Cloud, Microsoft Azure, Meta, and Dell. According to Bloomberg, which initially reported the news, OpenAI intends to implement Vera Rubin on a large scale in the third quarter.
CoreWeave, one of the first cloud providers to obtain the hardware, shared with Bloomberg that its NVL72 racks are achieving ten times the token output compared to the previous generation. The NVL72 is a complete rack system that combines 72 Rubin GPUs with Vera CPUs and utilizes liquid cooling to eliminate internal wiring, a design innovation that Nvidia showcased at the event. Buck explained that this cooling method enables them to remove copper connections that previously constrained the compactness of component arrangement.
The Vera CPU was designed by Nvidia to replace the processors it used to source from other suppliers. During the briefing, the company highlighted direct comparisons with AMD, asserting that the Vera CPU performs nearly twice as swiftly as AMD’s Turin chip in Python workloads, a relevant benchmark since Python is the dominant language for many AI inference applications. Among the initial labs to receive this processor are Anthropic, OpenAI, Perplexity, SpaceX, and Oracle.
This production milestone occurs amidst a challenging time for Nvidia’s stock. While the company’s shares have appreciated by nine percent this year, the broader chip index has surged by 66 percent in the same timeframe, according to Bloomberg, with Intel, ARM, and AMD all more than doubling in value.
Analysts predict that Nvidia’s revenue will grow by 82 percent, reaching approximately $393 billion for the fiscal year; however, this growth rate has not led to the sort of share price momentum that investors experienced during the Blackwell cycle.
Over the last two months, Nvidia has consistently developed the narrative around Vera Rubin. Jensen Huang announced that the platform had achieved full production at Computex in early June and identified Anthropic, OpenAI, SpaceX, and Oracle as initial recipients. Monday’s briefing provided performance metrics and customer feedback that were not included in the keynote, converting claims into quantifiable benchmarks.
It is important to note that the figures Nvidia highlighted come from its own assessments and those of its customers, rather than independent evaluations. The tenfold token output reported by CoreWeave and the Python benchmark comparison against AMD originated from the company’s own briefing and not from an external source. Volume shipments to all mentioned customers and independent validation of the performance claims are still forthcoming.
Other articles
Nvidia reports that Vera Rubin is now in full production, with OpenAI planning to scale its deployment in the third quarter.
Nvidia showcased its Vera Rubin systems at its headquarters, with Ian Buck announcing that full production is underway, and OpenAI intends to initiate large-scale deployment this quarter.
