A wider AMD role inside Azure

Microsoft and AMD have announced an expansion of their long-running infrastructure partnership that moves beyond individual processors or GPU virtual machines. The planned Azure deployment spans AMD’s forthcoming Helios rack-scale platform, sixth-generation EPYC processors codenamed “Venice”, Pensando data-processing units and networking technology, as well as the ROCm software stack.

The announcement, made on July 20, is significant because it places AMD hardware across several layers of Microsoft’s cloud design: AI accelerators, host CPUs, backend networking and customer-facing virtual machines. Microsoft says the systems will serve frontier-model inference for its own services, Azure AI offerings and customers. AMD expects to begin shipping Helios to customers, including Microsoft, in the second half of 2026.

That timing is important. The agreement describes a forthcoming deployment rather than a currently available Azure service, and neither company has disclosed the initial scale, regions, pricing or general-availability dates for the Helios-based instances. The commercial impact will therefore depend on execution, supply availability, software readiness and how rapidly Azure customers adopt the new capacity.

Helios shifts the focus from chips to racks

AMD Helios is designed as an integrated rack-scale reference architecture rather than a conventional server platform assembled from separately chosen components. The design combines Instinct MI455X accelerators, EPYC Venice CPUs, Pensando networking hardware and ROCm software. AMD also positions it around open infrastructure specifications, including Open Compute Project Open Rack Wide, Ultra Accelerator Link and Ultra Ethernet Consortium technologies.

In AMD’s published Helios configuration, a rack integrates 72 MI455X GPUs. That architecture is intended to address a central challenge in large AI deployments: achieving sufficient bandwidth not only within a GPU server, but across accelerators, compute trays and the wider cluster. For inference workloads involving very large models, the movement of model weights, prompts and generated outputs can be as consequential to system utilisation and response time as the nominal compute capability of the accelerators.

Microsoft’s initial emphasis is on inference rather than presenting Helios chiefly as a training platform. That reflects the growing need for cloud infrastructure that can support persistent, production-facing AI services, where predictable throughput, efficient serving and operational manageability matter alongside peak benchmark performance. Microsoft says the platform is intended for reasoning, search and agentic workloads, as well as broader Azure AI services.

This does not mean Helios is limited to inference. AMD describes the platform as suitable for both large-scale training and inference. However, Microsoft’s stated use case illustrates how hyperscalers are increasingly tailoring hardware fleets for different phases of the AI lifecycle rather than relying on a single general-purpose accelerator cluster.

New EPYC instances address data and engineering workloads

The agreement also includes two new Azure virtual-machine series based on sixth-generation AMD EPYC Venice processors. Azure HDv2 is aimed at AI data preparation, search, reinforcement learning and agent coordination. Microsoft says it will offer nearly 500 physical CPU cores, up to 4 TB of memory, 32 TB of local NVMe storage and 400 Gb Azure Boost networking.

Azure HXv2 is targeted at electronic-design automation, simulation and other technical-computing applications. It is planned to use 176 EPYC CPU cores, clock speeds above 5 GHz, large cache capacity and up to 4 TB of memory, with 800 Gb InfiniBand for distributed workloads. These specifications indicate that the release is not only about AI accelerators: CPU capacity remains a major determinant of how efficiently data is prepared, jobs are orchestrated and complex engineering workloads are run.

The split between HDv2 and HXv2 also shows Microsoft’s preference for workload-specific infrastructure. One instance family is designed around data-intensive AI pipelines, while the other targets applications where high single-thread performance, memory capacity and cache behaviour are essential. For chip designers in particular, faster simulation and verification can affect development schedules for the processors and accelerators that underpin the wider AI market.

Networking is a strategic part of the announcement

A notable element of the partnership is the inclusion of AMD Pensando DPUs and the planned integration of AMD silicon with Azure Boost. DPUs can offload networking, storage and security-related processing from host CPUs, while Azure Boost is Microsoft’s infrastructure system for moving virtualisation, networking and storage functions onto dedicated hardware and software.

Microsoft and AMD have not provided detailed performance figures for their combined implementation. Still, the approach matters because cloud AI clusters depend on more than GPU-to-GPU links. They must connect compute nodes to storage, customers and control systems while maintaining isolation and service reliability in multi-tenant environments. Offloading parts of that work can free CPU resources and potentially improve operational efficiency, but the outcome will depend on the final Azure configuration.

Competition, choice and the software test

For Azure, the AMD commitment adds depth to a deliberately heterogeneous AI infrastructure portfolio. Microsoft continues to operate major NVIDIA-based AI systems and has developed its own cloud silicon, while also expanding AMD-based offerings. The result is not a replacement of one supplier by another, but a broader set of options that Microsoft can match to different workloads, supply conditions and customer requirements.

For AMD, securing a planned production-scale deployment at a leading hyperscaler is strategically valuable. It validates the company’s effort to sell an integrated AI platform that includes networking and software, not simply standalone GPUs. Yet widespread customer use will rest heavily on ROCm maturity, framework compatibility, model-serving tooling and the ease with which developers can move workloads between accelerator platforms.

Microsoft’s existing Foundry managed-compute strategy highlights why that software layer matters. Its objective is to let customers deploy models through common management, identity, networking and endpoint experiences while the underlying accelerator infrastructure varies. If Helios capacity can be exposed through similarly straightforward cloud services, AMD hardware may become more accessible to enterprises that do not want to operate specialist AI clusters themselves.

The partnership therefore matters less as a single hardware win than as evidence of a changing cloud architecture. AI infrastructure is becoming a rack- and cluster-level design problem, where accelerators, CPUs, networking, memory, software and cloud operations must work together. Microsoft’s deployment plans give AMD a larger role in that stack, but the practical measure of success will be availability, performance in customer workloads and the economics of running production AI at scale.

Sources