Majestic Labs abandons the GPU to overcome Nvidia's memory challenges.
The discussion surrounding AI hardware has become fixated on the term "compute." However, a Tel Aviv-based startup aims to shift this focus to "memory." Majestic Labs, established in 2023 by former engineers from Google and Meta, has introduced a server that it claims can outperform a rack of Nvidia GPUs by targeting a different bottleneck.
According to a report by TechRadar, the argument is that the combination of expensive GPUs with limited high-bandwidth memory has reached an impasse for AI inference. The ability to execute a model is often constrained by access to fast memory rather than raw processing power. As a result, Majestic has moved away from GPUs.
Their server, named Prometheus, replaces graphic processing units with what they refer to as Ignite AI Processing Units. These units combine Arm cores with RISC-V vector and tensor engines. Up to 12 of these units can fit in one server, sharing a single pool of 8TB to 128TB of LPDDR6 memory, which is the more affordable memory typically found in smartphones, rather than the expensive high-bandwidth memory required by GPUs.
A novel approach to overcoming the memory limitation is in the server's wiring. Instead of attaching memory to each GPU, Majestic aggregates it through customized chiplets connected by copper cables up to one meter in length. This results in a unified memory pool that is significantly larger than what a typical GPU setup can access.
The comparison with Nvidia’s hardware is striking. An Nvidia DGX B300 featuring eight Blackwell GPUs comes with 2.3TB of high-bandwidth memory. Majestic claims that Prometheus provides over 50 times this amount of fast memory, while also achieving 1.7 times the interconnect bandwidth. The company asserts that one of its racks is equivalent to 25 Nvidia Vera Rubin racks in terms of fast memory and consumes far less power.
On the software side, Prometheus is designed to adhere to open standards, supporting platforms such as PyTorch, vLLM, and OpenAI's Triton. Models created for GPUs are intended to run seamlessly on this new system. Majestic reports that it has already secured orders from large enterprises, neoclouds, and hyperscalers.
However, every statistic they present comes with a caveat: these figures are based on Majestic’s own assessments and have not yet been validated by independent testing, nor has any hardware been shipped. The company has around 40 employees located in Tel Aviv and Los Angeles and raised $100 million late last year, a relatively small amount compared to its competitors.
Physical concerns also arise; a 128TB memory pool created from 2GB LPDDR6 chips would require about 64,000 individual chips, suggesting that over a hundred aggregation chiplets would be necessary in a single server. This raises the complexity of maintaining coherence across so many components, leading TechRadar to note that potential customers may hold off on switching until independent benchmarks are available.
Majestic is part of an expanding group of startups looking to challenge Nvidia from various angles, including optical chips, edge silicon for inference, and open networking hardware. The unifying theme is that Nvidia's stronghold is being targeted from multiple fronts, including its software ecosystem.
The focus on overcoming the memory bottleneck is the most recent challenge to Nvidia’s dominance, yet it remains the least substantiated. Majestic has depicted a scenario where its rack outperforms a room full of GPUs in terms of memory and efficiency. Next year, as hardware becomes available and others conduct benchmarks, the validity of this vision will be tested.
Other articles
Majestic Labs abandons the GPU to overcome Nvidia's memory challenges.
Majestic Labs' Prometheus server eliminates the GPU for Arm cores and offers up to 128TB of affordable LPDDR6, asserting that it overcomes Nvidia's memory limitations.
