Nvidia is developing an open model with a trillion parameters, which would still be smaller than China's model.
Nvidia launched Nemotron 3.5 Lightning on Tuesday. This is a mixture-of-experts model with 30 billion parameters, of which three billion are active at any given time. The architecture combines Mamba-2, MoE, and attention layers, featuring a context window of one million tokens. The model weights are available on Hugging Face and ModelScope.
The licensing aspect is particularly significant. It is distributed under OpenMDW-1.1, allowing for commercial use. Nvidia has also published the training data and recipes along with the weights, enabling companies to download, use, and modify it without needing permission or making any payments to Nvidia, as noted by CNBC.
Regarding speed, Nvidia claims an output speed up to four times that of similarly sized models. However, a more modest performance figure emerges upon further examination. On PinchBench, it achieved 86% accuracy while completing 10,000 tasks 30% faster than Alibaba’s Qwen 3.6 35B, maintaining similar accuracy levels.
Both of these performance metrics are included in the same release. The "four times faster" claim pertains to token generation in controlled conditions, while the 30% improvement reflects performance during actual tasks. This discrepancy is important to note whenever vendors cite throughput.
The benchmark results are commendable rather than exceptional, with scores of 81.94 on MMLU Pro, 75.44 on GPQA Diamond, and 51.56 on SWE-bench Verified. The pre-training phase involved over 20 trillion tokens.
The true innovation lies in the router system. Nvidia also introduced NeMo Switchyard, an open-source library that directs each step of an agent's workflow to the most suitable model. Plans are routed to a frontier model, while execution is directed to Lightning.
LangChain evaluated the system across 145 tasks, as compiled by MarkTechPost. Routing between Lightning and Claude Opus 4.8 resulted in a 74% reduction in costs compared to using only the frontier model. Only 7% of the calls were made to the pricier model, which led to a loss of approximately six accuracy points.
This approach does not position itself as the top model but rather as a means to manage 93% of calls more cost-effectively. Nvidia argues that most agent tasks do not require a frontier model, emphasizing a commoditization narrative aimed at laboratories.
The rationale behind Nvidia's decision to release models for free is straightforward and has always been transparent. Lower-cost and readily available models lead to increased inference, which relies on GPUs. By commoditizing the software, Nvidia is expanding the market for its underlying hardware.
We previously mentioned this strategy when Nvidia linked its safety initiatives to chip demand. Furthermore, it has placed models under a signature initiative, with Nvidia advocating against regulatory crackdowns in a letter to Washington. Jensen Huang even made his first post on X to support this initiative, unlike OpenAI, Anthropic, and Google, which did not sign.
In terms of future developments, The Information has reported that Nvidia is developing Nemotron 4, which is expected to exceed a trillion parameters, an increase from Nemotron 3 Ultra's 550 billion. This model could potentially be ready by late autumn and aims to compete with the best open models globally.
However, an important point is that even with over a trillion parameters, Nemotron 4 would still be smaller than the leading Chinese open models, with Moonshot’s Kimi K3 currently holding the record for the largest open model released.
Nvidia’s goal is thus to catch up rather than claim leadership, aiming for a size that China has already surpassed and acknowledging this in its briefings.
China has set and maintains this record pace, as Nvidia chose to benchmark against a Chinese model, specifically Qwen 3.6, instead of any American open models. Given its parameter size, Qwen represents a significant challenge.
DeepSeek has worked throughout the year to reduce prices on open models, making it impractical to run anything less capable. Alibaba and Moonshot have exchanged dominance at the top of the open model rankings, while American open weights have largely been a policy discussion rather than a product in circulation.
Nemotron 3.5 Lightning slightly alters this landscape by providing a legitimate, genuinely open model from the company that supplies the infrastructure.
Looking ahead, the focus should not solely rest on benchmark results. Nvidia has previously released Nemotron models, including Nemotron Nano Omni for edge agents, but none have shifted the balance of open weights away from Chinese research labs.
The true test will be Nemotron 4, which has a tentative release timeframe. If it arrives in late autumn and successfully claims the top spot in open charts, Nvidia will secure a position at a table it currently only supports financially. Conversely, if it ranks second to Kimi again, the reality of the strategy will lean toward generating demand for GPUs.
Both perspectives can hold true simultaneously, highlighting the rationale
Other articles
Nvidia is developing an open model with a trillion parameters, which would still be smaller than China's model.
Nvidia's Nemotron 3.5 Lightning is both free and quick. However, the upcoming trillion-parameter Nemotron 4 would still fall behind the top Chinese open models.
