AWS and NVIDIA commit to two million more GPUs and push the partnership down the stack to CPUs and robotics
On 26 August 2026, AWS and NVIDIA announced two million additional GPUs to be deployed by 2028, NVIDIA Vera CPUs, NVHBM memory, and 100,000-GPU federal AI factories. This is no longer a silicon supply deal: AWS is now co-engineering interconnect and memory with NVIDIA. Here is what it changes for a cloud engineer.
18 March 2026. At GTC 2026, AWS announces it will add more than one million NVIDIA GPUs starting in 2026. 26 August 2026. The two companies double down: two million additional GPUs to be deployed in 2027-2028. Jensen Huang sums up the move in one sentence: “demand is running ahead of every forecast.” Net result: a partnership that once meant shipping cards now spans CPUs, memory, networking and robotics.
Two million GPUs, with demand outrunning every forecast
The headline number is enormous. AWS plans to deploy 2 million additional NVIDIA GPUs — Blackwell Ultra, Rubin and Rubin Ultra — across its global infrastructure in 2027-2028. That capacity stacks on top of the million GPUs already announced at GTC 2026, a figure real-world demand exceeded before it was even fully deployed.
For an operator, the question is no longer “will we get GPUs?” but “how many, and for what workload profile?”. AWS is also widening the accelerated-instance menu: it is expanding Blackwell capacity, including RTX PRO 4500 Blackwell Server Edition GPUs for Amazon EC2 G7 instances, making it the first major cloud to offer instances accelerated by the RTX PRO 4500. G7 delivers 4.6× the inference performance and 2.1× the graphics performance of the previous G6 generation.
The real constraint is no longer quantity. It is the speed at which a workload moves from pilot to production. When a frontier lab, an enterprise or a government needs to scale up, waiting for a GPU tranche is a genuine opportunity cost.
Vera brings the CPU into the deal
Historically, the AWS-NVIDIA relationship was about accelerators. The 26 August announcement changes the terms by introducing the NVIDIA Vera CPU into the AWS portfolio.
Vera is the processor NVIDIA is building for the next generation of AI compute. Putting it on AWS is not meant to displace Graviton or the Intel/AMD instances: it adds another option for agentic workloads that need high-performance CPU compute alongside accelerators. Matt Garman, CEO of AWS, states the intent plainly — offer “the broadest choice of compute”, from AWS’s own silicon to the latest accelerators and CPUs from partners.
That is a strong strategic signal. A cloud provider that agrees to co-sell a partner’s CPU — alongside its own silicon — is conceding that the AI race is won on how the stack is composed, not on any single component.
NVLink Fusion and NVHBM: memory, made to order
The most technical part of the announcement concerns interconnect and memory. At re:Invent 2025, AWS announced support for NVLink Fusion, NVIDIA’s high-speed interconnect, in its next-generation Trainium chips. The August expansion goes further: Annapurna Labs, Amazon’s silicon arm, will work directly with NVIDIA’s custom high-bandwidth memory — NVHBM — in partnership with memory suppliers.
In practice, this gives Trainium chips access to faster, more power-efficient memory, and lets Trainium and NVIDIA GPUs sit in a common rack-scale architecture. For a cloud engineer the consequence is direct: the boundary between in-house silicon (Trainium) and partner silicon (NVIDIA) is becoming porous. A single cluster will be able to mix the two without changing the operating model.
Federal AI at 100,000 GPUs
The announcement carries a sovereign dimension. AWS and NVIDIA plan to build AI factories for the U.S. government, with a target of 100,000 GPUs on AWS’s secure infrastructure, for workloads classified at Impact Level 6 (IL6) and above.
This is the first public deal of this scale to place a commercial cloud at the centre of defence and national-security AI. The message is twofold: federal agencies will not build their own training capacity alone, and classification-level compliance is now a commercial differentiator in its own right.
Nemotron, cuDF, cuVS: the open software stack lands
Beyond hardware, the deal cements the presence of NVIDIA’s open Nemotron models on Amazon Bedrock (managed, serverless) and Amazon SageMaker (self-hosted fine-tuning). The stated goal is the “broadest model choice”.
Two numbers capture the software acceleration. On Amazon EMR, GPU-accelerated data processing via cuDF on G7 instances reaches 3.7× the speed with 30% better price/performance compared to CPU-based configurations. On Amazon OpenSearch Service, GPU-accelerated vector indexing via cuVS offloads index construction onto dedicated GPUs: 9× faster at a quarter of the cost.
These gains are not cosmetic. RAG pipelines and semantic search stall precisely on vector indexing at billions of records. Making it 9× faster shifts the economic balance of those architectures.
Physical AI enters the cloud
The least expected part is robotics. Amazon Robotics is adopting NVIDIA’s physical AI platform — Jetson, Omniverse, Isaac — to accelerate warehouse automation and next-generation robots. The collaboration spans simulation, synthetic data generation, training, route optimisation, functional safety and real-to-sim validation, all running on GPU-accelerated EC2 instances.
It is a strong sign of convergence between the cloud and physical AI. The same provider that hosts your models is also training the robots that move your packages. “Robotics” becomes a cloud workload like any other, with the same demands for massive simulation and continuous validation.
What it changes for a cloud engineer
The deal will not change your life tomorrow morning. It changes what you can plan 18 to 24 months out.
First, capacity: two million more GPUs is insurance against scarcity, not instant inventory. If your inference or training roadmap depends on a specific tranche of Blackwell Ultra or Rubin, reserve it when the instances are announced, not at go-live.
Second, silicon choice: with Vera, Trainium and NVIDIA GPUs on the same interconnect through NVLink Fusion, the question is no longer “which platform?” but “which mix at what cost?”. Run cross-benchmarks of Trainium against Blackwell on your own workloads before you commit.
Third, sovereignty: the 100,000-GPU federal programme signals that compliance (IL6+, dedicated zones, encryption) is now a product. If you operate in a regulated sector, that segment will set the standards.
Verdict
If you train or serve models on AWS, this partnership is good medium-term news: more capacity, more silicon choice and a unified interconnect. Start comparing Trainium and Blackwell on your real workloads now — that is the only way to turn the announcement into savings.
If you are a provider or a sovereign customer, read the federal track as a strategic marker. Classified AI will no longer be bought as standalone hardware: it will be consumed as a certified cloud service, with NVIDIA and AWS setting the reference.
References
- AWS and NVIDIA to Deliver 2 Million Additional GPUs and Next-Generation Infrastructure for Agentic and Physical AI — NVIDIA Newsroom, 26 August 2026
- AWS and NVIDIA expand partnership for next-gen AI infrastructure — About Amazon, 26 August 2026
- AWS and NVIDIA deepen strategic collaboration at GTC 2026 — AWS Machine Learning Blog