AMD's Helios integrates 72 GPUs and 31 terabytes of HBM4 within a single rack. This is AMD's response to Nvidia's NVL72.

AMD's Helios integrates 72 GPUs and 31 terabytes of HBM4 within a single rack. This is AMD's response to Nvidia's NVL72.

      **TL;DR**: AMD's Helios includes 72 MI455X GPUs, 31TB of HBM4, and delivers 2.9 exaflops of inference in a single rack, built on open standards. Engineering samples expected in the second half of 2026, with mass production starting in Q2 2027.

      AMD's Helios is a rack containing 72 Instinct MI455X GPUs, 31 terabytes of HBM4 memory, and 2.9 exaflops of FP4 inference compute power. This marks AMD's inaugural rack-scale AI system and serves as a direct competitor to Nvidia's Vera Rubin NVL72. The design comprises 18 compute trays, with each tray accommodating four MI455X accelerators based on the new CDNA 5 architecture, along with one sixth-generation EPYC “Venice” CPU. Engineering samples are set to ship in the latter half of 2026, while mass production is slated for Q2 2027.

      The architecture emphasizes open standards, utilizing UALink for the interconnect between GPUs within the rack, adhering to Ultra Ethernet Consortium specifications for networking between racks, and incorporating the OCP Open Rack Wide form factor. In contrast, Nvidia’s NVL72 relies on proprietary NVLink. AMD believes that data center managers seeking flexibility and not wanting to be tied to a single vendor's interconnect will invest in this option. Networking is managed by AMD Pensando AI NICs with programmable hardware and UEC-ready RDMA capabilities.

      The specifications aim to compete on memory in addition to computing power. Each MI455X GPU features HBM4 with 19.6 TB/s bandwidth, resulting in a total rack bandwidth of 260 TB/s for scale-up and 43 TB/s for scale-out. This memory capacity is crucial for training advanced models and long-context inference, where the limitation has transitioned from sheer computational power to the amount of data the system can store and transfer. The AI-driven memory shortage has caused significant increases in HBM prices, and 31TB of HBM4 in a single rack entails a substantial material cost that only large-scale cloud providers and government computing budgets can accommodate.

      Supermicro unveiled Helios hardware at Computex in June. AMD has invested billions into AI infrastructure in the UK, as showcased during London Tech Week, and Helios will form the foundation of these initiatives. The ROCm software stack provides support for frameworks like PyTorch, TensorFlow, and JAX, implying that developers should, theoretically, be able to transition from Nvidia’s CUDA ecosystem without needing to rewrite code. The key question is whether AMD can overcome the software disparity that has historically placed it behind Nvidia in AI compute capabilities. While the hardware specifications are competitive, the broader ecosystem remains the real challenge.

Other articles

This Mac application can transform any window into a CRT-inspired surreal experience or evoke the nostalgia of a Game Boy. This Mac application can transform any window into a CRT-inspired surreal experience or evoke the nostalgia of a Game Boy. Glaze 1.9 employs real-time GPU shaders to enhance every Mac window, video, and game, offering more than 50 different appearances for a one-time cost of $9.99. The early internet created a museum for fading tech sounds, and it surprisingly remains active. The early internet created a museum for fading tech sounds, and it surprisingly remains active. This charmingly retro online museum keeps alive the sounds of dial-up modems, floppy drives, Windows startup chimes, and other auditory memories that faded away as our devices became sleeker, quicker, and significantly less vocal. Google Frozen chip: Gemini integrated into the silicon. Google Frozen chip: Gemini integrated into the silicon. A report indicates that the Google Frozen chip will integrate Gemini directly into silicon, achieving up to 10 times the efficiency by 2028. This has implications for AI. Singles are utilizing ChatGPT and Claude to craft their dating messages, and the apps have no intention of putting a stop to it. Singles are utilizing ChatGPT and Claude to craft their dating messages, and the apps have no intention of putting a stop to it. 26% of adults in the U.S. have sought assistance from AI for dating purposes. Match, Hinge, and Bumble indicate that they are not filtering out messages generated by AI. This trend is referred to as chatfishing. AMD's Helios incorporates 72 GPUs and 31 terabytes of HBM4 within a single rack. This is AMD's response to Nvidia's NVL72. AMD's Helios incorporates 72 GPUs and 31 terabytes of HBM4 within a single rack. This is AMD's response to Nvidia's NVL72. The AMD Helios contains 72 MI455X GPUs, 31TB of HBM4, and delivers 2.9 exaflops of inference computing capabilities within a single rack. Engineering samples are expected to ship in the second half of 2026, with mass production set to begin in the second quarter of 2027. Camera anxiety is transforming the way young adults socialize at parties. Camera anxiety is transforming the way young adults socialize at parties. Students express that they can no longer fully unwind during a night out, caught in the middle of enthusiastic promoters with cameras and glasses that resemble anything but a recording device.

AMD's Helios integrates 72 GPUs and 31 terabytes of HBM4 within a single rack. This is AMD's response to Nvidia's NVL72.

The AMD Helios contains 72 MI455X GPUs, 31TB of HBM4, and delivers 2.9 exaflops of inference computing within a single rack. Engineering samples are set to be shipped in the second half of 2026, with mass production beginning in the second quarter of 2027.