by Nitin - 14 hours ago - 4 min read
AMD has officially launched Helios, its first complete rack-scale AI infrastructure platform and its clearest attempt yet to challenge Nvidia in the market for large AI data centres.
Unveiled at AMD’s Advancing AI 2026 event in San Francisco on July 23, Helios has entered production, with customer shipments expected to begin near the end of the third quarter of 2026. Rather than selling individual accelerators, AMD is now combining its GPUs, server CPUs, networking hardware and software into one integrated system designed for frontier-model training and large-scale inference.
A full Helios configuration contains 72 AMD Instinct MI455X accelerators and 18 sixth-generation EPYC “Venice” CPUs. It also uses AMD Pensando networking and the company’s ROCm software platform to connect and manage the complete rack.
AMD says one rack can deliver up to 2.9 exaFLOPS of FP4 performance and 1.4 exaFLOPS at FP8 precision. More importantly for increasingly large AI models, Helios includes 31TB of HBM4 memory, alongside 260TB per second of scale-up bandwidth and 43TB per second of scale-out networking bandwidth.
| Specification | AMD Helios | Nvidia Vera Rubin NVL72 |
|---|---|---|
| AI accelerators | 72 MI455X GPUs | 72 Rubin GPUs |
| Host processors | 18 EPYC Venice CPUs | 36 Vera CPUs |
| GPU memory | 31TB HBM4 | 20.7TB HBM4 |
| Peak low-precision inference | 2.9 EF FP4 | 3.6 EF NVFP4 |
| Scale-up bandwidth | 260TB/s | 260TB/s |
| Scale-out bandwidth | 43TB/s | 28.8TB/s |
The two systems use different precision formats and benchmarking methods, so their headline compute figures are not directly comparable. Nvidia also describes its published Vera Rubin figures as preliminary.
AMD claims Helios can produce up to 30% more inference tokens per dollar than the leading competing system. The company is placing greater emphasis on inference economics because running trained models for millions of users is becoming one of the largest ongoing costs for AI companies.
However, the 30% figure comes from AMD’s own performance modelling and theoretical system comparisons rather than independent production testing. Actual results will depend on model architecture, software optimisation, power availability, utilisation rates and the final configuration supplied by AMD’s hardware partners.
Helios is also technically a reference design rather than a single rack sold directly by AMD. Manufacturers including HPE, Lenovo and Supermicro are expected to build commercial systems based on the design, while infrastructure partners will handle manufacturing and deployment.
The announcement carries more weight because several major AI companies are already preparing to use the platform.
OpenAI expects to bring Helios infrastructure online during the fourth quarter of 2026 and increase deployments throughout 2027. Microsoft will introduce Helios into Azure for frontier-model inference, Azure AI services and customer applications. Meta has started testing Helios racks, while Anthropic plans to deploy up to two gigawatts of AMD Instinct MI455X capacity through its expanded partnership with AMD.
AMD is also working with Cerebras on a disaggregated inference system. The planned setup will combine Helios infrastructure for high-throughput processing with Cerebras hardware for workloads requiring extremely low response latency.
Helios is based on Meta’s double-wide Open Rack Wide design and supports standards including UALink and Ultra Ethernet. AMD argues that this approach gives cloud providers and data-centre operators more control over networking, rack design and component selection than a closed, single-vendor architecture.
That openness is an important part of AMD’s pitch, but the company must still prove that its ROCm software stack, manufacturing partners and deployment support can perform consistently at gigawatt scale.
AMD estimates that the broader computing market could reach approximately $2 trillion by 2030, including $1.4 trillion for AI accelerators and $220 billion for server CPUs. These remain AMD’s projections, but they explain why the company is moving beyond individual chips and competing directly for complete AI data-centre deployments.
Helios does not immediately remove Nvidia’s infrastructure advantage. It does, however, give hyperscalers a credible second rack-scale option backed by large memory capacity, open standards and commitments from some of the world’s biggest AI buyers.