Nvidia is developing an open model with a trillion parameters, which would still be less extensive than China's model.
Nvidia unveiled Nemotron 3.5 Lightning on Tuesday. This model features 30 billion parameters, with three billion active at any given time. Its architecture is a combination of Mamba-2, MoE, and attention layers, along with a context window of one million tokens. The model weights are available on Hugging Face and ModelScope.
The license is the critical aspect. It is released under OpenMDW-1.1, allowing for commercial use, and Nvidia has made the training data and methodologies available alongside the weights. Companies can download, use, and modify it freely without needing permission or paying Nvidia, as noted by CNBC.
Nvidia's speed claims feature two figures: the company asserts that its output speed can be up to four times faster than similar-sized models. However, a more modest figure emerges when examining practical tasks. On PinchBench, it achieved 86% accuracy while completing 10,000 tasks 30% faster than Alibaba’s Qwen 3.6 35B at comparable accuracy.
Both figures are part of the same release: the fourfold speed increase refers to token generation in a laboratory setting, while the thirty percent improvement applies to real-world tasks. This disparity is important to consider when a vendor discusses throughput.
The benchmark results are commendable but not extraordinary. The model scores 81.94 on MMLU Pro, 75.44 on GPQA Diamond, and 51.56 on SWE-bench Verified, with pre-training covering more than 20 trillion tokens.
The true innovation lies in the router. Nvidia also introduced NeMo Switchyard, an open-source library that directs each step of an agent's workflow to the most suitable model. Plans are routed to a frontier model, while execution goes to Lightning.
LangChain utilized it for 145 tasks, with data compiled by MarkTechPost. Routing calls between Lightning and Claude Opus 4.8 resulted in a 74% cost reduction compared to using the frontier model alone, with only 7% of calls directed to the more expensive model, resulting in an approximate loss of six accuracy points.
This is not an assertion of being the top model, but rather an argument for managing 93% of the calls. Nvidia contends that most agent work does not require a frontier model, making this a case for commoditization, targeting lab operations.
As for why a chip company would distribute models for free, the reason is clear and Nvidia has been open about it. Offering cheaper, freely accessible models increases inference, which relies on GPUs. By commoditizing software, it broadens the market for its underlying hardware.
This strategy mirrors a previous report where Nvidia linked its safety initiatives to chip demand. The company has also signed a letter advocating for open weights, urging Washington to reconsider regulatory measures. Jensen Huang made his first post on X to highlight this initiative, unlike OpenAI, Anthropic, and Google, which did not sign.
The Information reports that Nvidia is developing Nemotron 4 with over a trillion parameters, an increase from the 550 billion in Nemotron 3 Ultra. It could be ready by late autumn, aiming to compete with the top open models globally.
However, the notable point is that even at a trillion parameters, Nemotron 4 would still be smaller than the leading Chinese open models. Moonshot’s Kimi K3 currently holds the record as the largest open model released.
Thus, Nvidia's ambition appears to be catching up rather than leading, as they are working to reach a size that China has already surpassed, a fact acknowledged in their briefings.
China has set this benchmark and continues to hold it. Nvidia chose a Chinese model for its benchmark comparison, with Lightning being evaluated against Qwen 3.6 instead of any American open model. At its size, Qwen is currently the model to emulate.
DeepSeek has spent the year reducing prices on open models, making anything less capable impractical. Alibaba and Moonshot have alternated at the top of the open charts during this time, while American open weights have mainly served as a policy discussion rather than an actual market product.
Nemotron 3.5 Lightning makes a small shift in this landscape by being a genuine, open model from a company known for its hardware.
Looking ahead, the critical point will be Nemotron 4, which has an approximate launch timeline. If it arrives in late autumn and claims the top position in open models, Nvidia will secure a place at a table it currently only finances. If it ranks second to Kimi again, the honest interpretation of the strategy would be as demand generation for GPUs.
Both scenarios can coexist, underscoring the rationale behind providing free software while selling hardware.
Other articles
Nvidia is developing an open model with a trillion parameters, which would still be less extensive than China's model.
Nvidia's Nemotron 3.5 Lightning is both free and quick. The upcoming trillion-parameter Nemotron 4 would still lag behind the top Chinese open models.
