Nvidia is developing an open model with a trillion parameters, which would still be smaller than China's model.

Nvidia is developing an open model with a trillion parameters, which would still be smaller than China's model.

      Nvidia launched Nemotron 3.5 Lightning on Tuesday. This is a mixture-of-experts model with 30 billion parameters, of which three billion are active at any given time. The architecture combines Mamba-2, MoE, and attention layers, featuring a context window of one million tokens. The model weights are available on Hugging Face and ModelScope.

      The licensing aspect is particularly significant. It is distributed under OpenMDW-1.1, allowing for commercial use. Nvidia has also published the training data and recipes along with the weights, enabling companies to download, use, and modify it without needing permission or making any payments to Nvidia, as noted by CNBC.

      Regarding speed, Nvidia claims an output speed up to four times that of similarly sized models. However, a more modest performance figure emerges upon further examination. On PinchBench, it achieved 86% accuracy while completing 10,000 tasks 30% faster than Alibaba’s Qwen 3.6 35B, maintaining similar accuracy levels.

      Both of these performance metrics are included in the same release. The "four times faster" claim pertains to token generation in controlled conditions, while the 30% improvement reflects performance during actual tasks. This discrepancy is important to note whenever vendors cite throughput.

      The benchmark results are commendable rather than exceptional, with scores of 81.94 on MMLU Pro, 75.44 on GPQA Diamond, and 51.56 on SWE-bench Verified. The pre-training phase involved over 20 trillion tokens.

      The true innovation lies in the router system. Nvidia also introduced NeMo Switchyard, an open-source library that directs each step of an agent's workflow to the most suitable model. Plans are routed to a frontier model, while execution is directed to Lightning.

      LangChain evaluated the system across 145 tasks, as compiled by MarkTechPost. Routing between Lightning and Claude Opus 4.8 resulted in a 74% reduction in costs compared to using only the frontier model. Only 7% of the calls were made to the pricier model, which led to a loss of approximately six accuracy points.

      This approach does not position itself as the top model but rather as a means to manage 93% of calls more cost-effectively. Nvidia argues that most agent tasks do not require a frontier model, emphasizing a commoditization narrative aimed at laboratories.

      The rationale behind Nvidia's decision to release models for free is straightforward and has always been transparent. Lower-cost and readily available models lead to increased inference, which relies on GPUs. By commoditizing the software, Nvidia is expanding the market for its underlying hardware.

      We previously mentioned this strategy when Nvidia linked its safety initiatives to chip demand. Furthermore, it has placed models under a signature initiative, with Nvidia advocating against regulatory crackdowns in a letter to Washington. Jensen Huang even made his first post on X to support this initiative, unlike OpenAI, Anthropic, and Google, which did not sign.

      In terms of future developments, The Information has reported that Nvidia is developing Nemotron 4, which is expected to exceed a trillion parameters, an increase from Nemotron 3 Ultra's 550 billion. This model could potentially be ready by late autumn and aims to compete with the best open models globally.

      However, an important point is that even with over a trillion parameters, Nemotron 4 would still be smaller than the leading Chinese open models, with Moonshot’s Kimi K3 currently holding the record for the largest open model released.

      Nvidia’s goal is thus to catch up rather than claim leadership, aiming for a size that China has already surpassed and acknowledging this in its briefings.

      China has set and maintains this record pace, as Nvidia chose to benchmark against a Chinese model, specifically Qwen 3.6, instead of any American open models. Given its parameter size, Qwen represents a significant challenge.

      DeepSeek has worked throughout the year to reduce prices on open models, making it impractical to run anything less capable. Alibaba and Moonshot have exchanged dominance at the top of the open model rankings, while American open weights have largely been a policy discussion rather than a product in circulation.

      Nemotron 3.5 Lightning slightly alters this landscape by providing a legitimate, genuinely open model from the company that supplies the infrastructure.

      Looking ahead, the focus should not solely rest on benchmark results. Nvidia has previously released Nemotron models, including Nemotron Nano Omni for edge agents, but none have shifted the balance of open weights away from Chinese research labs.

      The true test will be Nemotron 4, which has a tentative release timeframe. If it arrives in late autumn and successfully claims the top spot in open charts, Nvidia will secure a position at a table it currently only supports financially. Conversely, if it ranks second to Kimi again, the reality of the strategy will lean toward generating demand for GPUs.

      Both perspectives can hold true simultaneously, highlighting the rationale

Other articles

Google's latest Pixel 11 Pro Fold is more lightweight, slimmer, and smarter, although its battery life has been compromised. Google's latest Pixel 11 Pro Fold is more lightweight, slimmer, and smarter, although its battery life has been compromised. Google's Pixel 11 Pro Fold weighs 239 grams and features a larger main camera sensor, a brighter screen, and HiLight, although its battery capacity decreases more significantly compared to the Pixel 11 Pro and Pro XL. Zoom resolved three vulnerabilities that allowed any participant on a call to gain control of your device. Zoom resolved three vulnerabilities that allowed any participant on a call to gain control of your device. Three bugs in Zoom's annotation feature allow any participant on a call to execute code on other devices. Fixes were released in June and July, two months prior to the announcement. Pixel 11 incorporates Gemini into its operating system, a few weeks ahead of Siri adopting the same models. Pixel 11 incorporates Gemini into its operating system, a few weeks ahead of Siri adopting the same models. Google has introduced the Pixel 11 series, equipped with advanced Gemini features, just weeks ahead of Apple's release of a revamped Siri that operates on Google's models. Gemini boasts a billion monthly users. ChatGPT reached that milestone in June and subsequently raised the bar. Gemini boasts a billion monthly users. ChatGPT reached that milestone in June and subsequently raised the bar. Gemini has surpassed a billion monthly users. ChatGPT achieved this milestone in June and hit a billion weekly users in July. Therefore, Gemini is one metric behind, rather than being on the same level. Beyond algorithms: Booker Loud discusses creating trustworthy social platforms. Beyond algorithms: Booker Loud discusses creating trustworthy social platforms. Booker Loud, the founder of Pin-Social, asserts that the forthcoming phase of social media will be characterized not by algorithm-driven reach but by the platforms that foster the most trusted environments. His approach prioritizes connection instead of outrage. The Pixel 11 finally includes the video enhancements I've been looking forward to, but not all Pixels receive these updates. The Pixel 11 finally includes the video enhancements I've been looking forward to, but not all Pixels receive these updates. This year, Google equipped the Pixel 11 lineup with authentic video enhancements, including 8K recording and an integrated teleprompter.

Nvidia is developing an open model with a trillion parameters, which would still be smaller than China's model.

Nvidia's Nemotron 3.5 Lightning is both free and quick. However, the upcoming trillion-parameter Nemotron 4 would still fall behind the top Chinese open models.