Amazon Web Services plans to deploy an additional 2 million Nvidia GPUs across its global infrastructure in 2027 and 2028, deepening a partnership that now reaches beyond accelerators into CPUs, networking, robotics software and dedicated “AI factory” infrastructure.

The companies announced the expansion on August 26. It builds on AWS’s March plan to begin adding more than 1 million Nvidia GPUs in 2026. Taken together, the public plans point to more than 3 million additional Nvidia GPUs over several years, although neither company disclosed the contract value, a region-by-region schedule or when specific instance types will become available to ordinary cloud customers.

The size of the commitment is a clear signal that hyperscalers expect demand for training and running AI systems to remain high well beyond the current product cycle. It is not, however, the same as 2 million chips being installed today. The delivery window spans two future years and includes architectures at different stages of their rollout.

What AWS says it will deploy

According to the joint AWS and Nvidia announcement, the additional fleet will include Nvidia Blackwell Ultra, Rubin and Rubin Ultra GPUs. AWS also plans to offer infrastructure based on Nvidia’s Vera CPU, both as part of integrated systems and as a standalone compute option.

The partnership covers more of the rack than the headline number suggests. The companies are working on Nvidia Spectrum networking for large training clusters, integration with the AWS Nitro System and Elastic Fabric Adapter, and support for Nvidia’s Nemotron open models and physical-AI software. AWS also plans more Blackwell capacity, including RTX Pro 4500-based EC2 G7 instances for graphics and inference workloads.

A government component is part of the plan. AWS and Nvidia say they will build secure AI factories for U.S. federal and national-security workloads, including 100,000 GPUs on AWS infrastructure designed for Impact Level 6 requirements. That is a stated future deployment; the announcement does not say the environment is already operational.

Independent reporting supports the central figures. Bloomberg reported that the two million high-end GPUs are additional to the one million AWS previously said it would install, while TechCrunch confirmed the 2027–2028 window and the mix of Blackwell Ultra, Rubin and Rubin Ultra products.

Capacity is becoming a multi-year procurement decision

The announcement shows how far cloud capacity planning has moved from buying individual servers. Modern AI systems depend on tightly connected racks, high-bandwidth networking, large power allocations, cooling and software that can schedule work across thousands of accelerators. Securing GPUs is necessary, but it is only one part of bringing usable capacity online.

For developers, the additional supply could eventually reduce waits for large clusters and widen access to newer Nvidia architectures. It does not guarantee lower cloud prices. AWS has not announced pricing, reservation terms or the share of capacity that will be committed to large customers and government projects. Newer systems may improve performance per watt or per token, but those gains can be absorbed by larger models and more inference demand.

The long timeline also exposes execution risk. Data centers need power, grid connections and cooling before racks can be commissioned. Rubin and Rubin Ultra are future-generation platforms, so final delivery depends on Nvidia’s manufacturing ramp, memory supply, networking availability and AWS’s own construction schedule. A planned GPU count should not be confused with deployed servers, available instances or completed customer workloads.

What it means for Nvidia’s position

AWS develops its own AI chips, but the expanded agreement shows that custom silicon has not removed its need for Nvidia at the high end of the market. Customers often arrive with software built around Nvidia’s CUDA ecosystem, and large training jobs benefit from a mature combination of accelerators, networking and optimized libraries.

The deeper integration could strengthen that advantage. Bringing Nvidia CPUs and networking into AWS means the relationship is moving closer to a full-system design rather than a simple GPU purchase. For AWS, that can accelerate deployment of a known stack. For customers, it may improve performance and availability. For the market, it also concentrates more critical AI infrastructure around one supplier’s roadmap.

Cloud buyers should therefore evaluate portability at the software layer. Containerized workloads, standard model formats, reproducible performance tests and a clear fallback strategy across instance families can reduce—but not eliminate—dependence on one architecture. Teams signing long reservations should compare total job cost and completion time, not only the hourly GPU price.

The number is a plan, not the outcome

Two million additional GPUs is a remarkable procurement signal, and more than 3 million planned additions would give AWS an enormous Nvidia footprint. The useful milestones now are operational: which regions receive capacity, when customers can reserve it, how reliably large clusters perform and what a completed training or inference job costs.

Until those details arrive, the announcement is best read as evidence of sustained confidence in AI compute demand—and of the capital, energy and supply-chain commitments required to serve it. The infrastructure race is no longer measured only by the next chip. It is measured by who can turn millions of promised components into dependable computing services.