A capacity plan rather than an immediate shipment

AWS and NVIDIA have set out a new expansion plan for cloud AI infrastructure, with AWS intending to deploy two million additional NVIDIA GPUs across its global estate in 2027 and 2028. The commitment covers NVIDIA Blackwell Ultra, Rubin and Rubin Ultra-generation products, and comes on top of AWS’s earlier plan, announced in March 2026, to add more than one million NVIDIA GPUs beginning this year.

The wording matters. This is a multi-year deployment objective, not a claim that two million processors have already been delivered or installed. Turning an order or supply arrangement into usable cloud capacity requires data-centre space, electrical infrastructure, cooling, servers, networking, software validation and regional deployment. For customers, the practical questions will be which instance types become available, in which AWS Regions, on what timetable and at what price.

Still, the scale is material. It signals that AWS expects demand for accelerated computing to extend well beyond the initial wave of generative AI training. The companies frame the build-out around model training and inference, but also data engineering, scientific workloads, enterprise automation, vector search and robotics. That breadth is important because it ties GPU demand to a larger collection of production workloads rather than to a single category of frontier models.

The partnership is moving beyond GPU supply

The announcement is not solely about adding processors. AWS and NVIDIA are deepening integration across the underlying systems needed to operate very large clusters. AWS plans to combine NVIDIA platforms with its Nitro System and Elastic Fabric Adapter networking, while also working with NVIDIA on Spectrum networking. These components matter because a large AI workload depends on communication between many accelerators as much as on the performance of an individual GPU.

The companies also intend to bring NVIDIA Vera CPU-based infrastructure to AWS. That reflects the heterogeneous character of modern AI systems: applications often need conventional CPU resources for orchestration, data preparation and services surrounding models, alongside GPU capacity for highly parallel computation.

Another notable element is the expansion of NVIDIA NVLink Fusion work with AWS’s Annapurna Labs. AWS had previously announced support for the interconnect technology in a future Trainium design. The new plan adds work around NVIDIA’s custom high-bandwidth memory technology. The strategic implication is that AWS is not treating its own AI chips and NVIDIA hardware as wholly separate product lines. Instead, it is seeking ways to combine them at rack scale, preserving a choice of compute architectures while improving how they communicate.

That approach could give AWS flexibility in serving customers with different performance, software and cost requirements. It also acknowledges NVIDIA’s continuing importance in the AI software ecosystem, where many applications and tools have been built around CUDA and related libraries.

Near-term services provide context

AWS already has some of the technologies cited in the new announcement in service. Its EC2 G7 instances, using NVIDIA RTX PRO 4500 Blackwell Server Edition GPUs, became generally available in June 2026 in US East (Ohio) and US West (Oregon). AWS says these instances are aimed at inference, graphics and GPU-accelerated analytics workloads, rather than only the largest training clusters.

The new agreement also identifies data processing as a priority. AWS and NVIDIA are working on GPU acceleration for Amazon EMR using the cuDF library, as well as vector indexing for Amazon OpenSearch. The claimed gains are specific to the relevant benchmark configurations, but the overall direction is clear: cloud providers want GPUs to be consumed by data teams and application developers, not reserved exclusively for specialist AI research groups.

This matters commercially. Training a new foundation model is a high-profile use of compute, but many enterprises will obtain more predictable value from running inference, preparing data, accelerating analytics or operating search systems. Broadening these workloads can improve utilisation of expensive infrastructure and make GPU-backed cloud services more relevant to customers that are not developing large models themselves.

Government and physical AI broaden the addressable market

AWS and NVIDIA also plan AI factories for United States government workloads, including 100,000 GPUs on AWS secure infrastructure for federal and national-security use. The companies say the systems are intended for workloads classified at Impact Level 6 and above. Such deployments will be subject to security, procurement and operational requirements that differ from those of commercial cloud regions, so the headline number should be viewed as part of a longer implementation programme.

The partnership additionally reaches into physical AI through Amazon Robotics’ use of NVIDIA technology for simulation, synthetic data, training and validation. Robotics represents a more demanding workload mix than text generation alone. It can require simulated environments, vision processing, route planning and repeated feedback between virtual systems and real-world operations. AWS’s existing cloud scale and NVIDIA’s software platforms make that a logical area for closer collaboration, although the ultimate pace of adoption will depend on deployment economics and reliability in industrial settings.

A competitive balancing act for AWS

The expansion reinforces NVIDIA’s role as a central supplier to AWS, but it does not replace Amazon’s custom-silicon strategy. AWS continues to develop Graviton CPUs and Trainium AI accelerators, and has highlighted customer commitments involving Trainium capacity. Its stated strategy is to offer a broad set of accelerators rather than making customers choose a single architecture.

For NVIDIA, the agreement extends demand visibility into 2028 and anchors successive generations of its data-centre hardware within one of the largest cloud platforms. For AWS, it helps reduce the risk that customers looking specifically for NVIDIA hardware will take their workloads to competing clouds or build private infrastructure instead.

The limiting factors will be execution rather than ambition. The deployment requires sustained availability of chips and surrounding components, rapid construction of power-intensive data-centre capacity, and networking that keeps large clusters efficiently utilised. AWS and NVIDIA have made an unusually expansive commitment, but its economic significance will ultimately be measured in installed capacity, customer access and the workloads that run on it—not by the announced GPU total alone.

Sources