Artificial Intelligence

AMD Launches Helios Rack-Scale AI System With 72 GPUs

by Nitin - 14 hours ago - 4 min read

AMD has officially launched Helios, its first complete rack-scale AI infrastructure platform and its clearest attempt yet to challenge Nvidia in the market for large AI data centres.

Unveiled at AMD’s Advancing AI 2026 event in San Francisco on July 23, Helios has entered production, with customer shipments expected to begin near the end of the third quarter of 2026. Rather than selling individual accelerators, AMD is now combining its GPUs, server CPUs, networking hardware and software into one integrated system designed for frontier-model training and large-scale inference.

A complete AI rack built around 72 GPUs

A full Helios configuration contains 72 AMD Instinct MI455X accelerators and 18 sixth-generation EPYC “Venice” CPUs. It also uses AMD Pensando networking and the company’s ROCm software platform to connect and manage the complete rack.

AMD says one rack can deliver up to 2.9 exaFLOPS of FP4 performance and 1.4 exaFLOPS at FP8 precision. More importantly for increasingly large AI models, Helios includes 31TB of HBM4 memory, alongside 260TB per second of scale-up bandwidth and 43TB per second of scale-out networking bandwidth.

SpecificationAMD HeliosNvidia Vera Rubin NVL72
AI accelerators72 MI455X GPUs72 Rubin GPUs
Host processors18 EPYC Venice CPUs36 Vera CPUs
GPU memory31TB HBM420.7TB HBM4
Peak low-precision inference2.9 EF FP43.6 EF NVFP4
Scale-up bandwidth260TB/s260TB/s
Scale-out bandwidth43TB/s28.8TB/s

The two systems use different precision formats and benchmarking methods, so their headline compute figures are not directly comparable. Nvidia also describes its published Vera Rubin figures as preliminary.

AMD is competing on token economics

AMD claims Helios can produce up to 30% more inference tokens per dollar than the leading competing system. The company is placing greater emphasis on inference economics because running trained models for millions of users is becoming one of the largest ongoing costs for AI companies.

However, the 30% figure comes from AMD’s own performance modelling and theoretical system comparisons rather than independent production testing. Actual results will depend on model architecture, software optimisation, power availability, utilisation rates and the final configuration supplied by AMD’s hardware partners.

Helios is also technically a reference design rather than a single rack sold directly by AMD. Manufacturers including HPE, Lenovo and Supermicro are expected to build commercial systems based on the design, while infrastructure partners will handle manufacturing and deployment.

OpenAI and Microsoft strengthen the launch

The announcement carries more weight because several major AI companies are already preparing to use the platform.

OpenAI expects to bring Helios infrastructure online during the fourth quarter of 2026 and increase deployments throughout 2027. Microsoft will introduce Helios into Azure for frontier-model inference, Azure AI services and customer applications. Meta has started testing Helios racks, while Anthropic plans to deploy up to two gigawatts of AMD Instinct MI455X capacity through its expanded partnership with AMD.

AMD is also working with Cerebras on a disaggregated inference system. The planned setup will combine Helios infrastructure for high-throughput processing with Cerebras hardware for workloads requiring extremely low response latency.

Open standards form the centre of AMD’s strategy

Helios is based on Meta’s double-wide Open Rack Wide design and supports standards including UALink and Ultra Ethernet. AMD argues that this approach gives cloud providers and data-centre operators more control over networking, rack design and component selection than a closed, single-vendor architecture.

That openness is an important part of AMD’s pitch, but the company must still prove that its ROCm software stack, manufacturing partners and deployment support can perform consistently at gigawatt scale.

AMD estimates that the broader computing market could reach approximately $2 trillion by 2030, including $1.4 trillion for AI accelerators and $220 billion for server CPUs. These remain AMD’s projections, but they explain why the company is moving beyond individual chips and competing directly for complete AI data-centre deployments.

Helios does not immediately remove Nvidia’s infrastructure advantage. It does, however, give hyperscalers a credible second rack-scale option backed by large memory capacity, open standards and commitments from some of the world’s biggest AI buyers.