AWS and Nvidia said on August 26 that AWS will add 2 million more Nvidia GPUs to its global infrastructure across 2027 and 2028. The chips span Nvidia's Blackwell Ultra, Rubin, and Rubin Ultra families, and the deal also deepens a chip-to-chip partnership that lets AWS's own custom silicon borrow Nvidia's memory technology.

  • AWS will deploy 2 million additional Nvidia GPUs in 2027 and 2028, on top of GPUs already running in its data centers today.
  • The order covers three chip generations at once: Blackwell Ultra, Rubin, and Rubin Ultra, plus what the companies call next-generation infrastructure for agentic and physical AI.
  • AWS's next-generation Trainium chips will gain access to Nvidia's new custom high-bandwidth memory (NVHBM), built with memory suppliers specifically for this integration.
  • The expansion follows an earlier 1 million GPU commitment that, according to TechTimes, AWS burned through faster than planned, a real-time signal that AI compute demand keeps outrunning supply.

What AWS and Nvidia actually announced

The headline number is 2 million GPUs, deployed over two years rather than all at once. That phasing matters: 2027 gets one wave, 2028 gets another, and the chip families ship in sequence as Nvidia's own roadmap matures. Blackwell Ultra is closest to production today. Rubin and Rubin Ultra are Nvidia's next architecture generations, still ramping toward volume manufacturing, so spreading the order across two years is as much about matching Nvidia's fabrication schedule as it is about AWS's own capacity planning.

RelatedNVIDIA and SK Group Unveil $500B Korea AI Buildout

Nvidia's language about "next-generation infrastructure for agentic and physical AI" is doing real work here too. It signals that this isn't just more chips for chatbots and image generators. Agentic AI systems that plan and execute multi-step tasks, and physical AI workloads like robotics and autonomous systems, need sustained low-latency compute at a different scale than a single inference request. Building for that use case now is a bet that it becomes a much bigger share of AWS's customer base within the next two years.

GPU commitment scale: prior deal versus new deal Bar chart comparing AWS's prior 1 million GPU commitment against the new 2 million GPU commitment announced August 26, 2026, showing a doubling in scale. GPU COMMITMENT SCALE, 2024 DEAL VS 2026 DEAL 1,000,000 PRIOR COMMITMENT 2,000,000 NEW: 2027 TO 2028 genztech.blog
Fig 1 The new 2 million GPU order is double the scale of AWS's earlier 1 million GPU commitment, and it ships over two years instead of one lump delivery.

Why did AWS's 1 million GPU commitment run out early?

This is the part of the story that's easy to skim past. A company doesn't double its GPU order because it feels generous. It doubles the order because the first one wasn't enough, and TechTimes' reporting that the prior 1 million GPU deal "ran out early" is a blunt way of saying AWS underestimated demand even after making what looked like an enormous commitment at the time.

That's the pattern across the entire hyperscaler field right now. Google, Microsoft, and Meta are all simultaneously signing record GPU orders and pouring billions into their own custom chips, because no single supply plan has kept pace with how fast enterprises are adopting AI workloads. Training the next generation of frontier models eats compute. Running inference for millions of daily users eats more. Agentic workflows that call a model dozens of times per task eat still more. When AWS says its prior number ran out ahead of schedule, that's a demand signal worth taking seriously, not a marketing line.

  1. Earlier dealAWS commits to roughly 1 million Nvidia GPUs baseline capacity plan, since exceeded by demand
  2. Aug 26, 2026AWS and Nvidia announce 2 million additional GPUs Blackwell Ultra, Rubin, Rubin Ultra, plus Trainium/NVHBM integration
  3. 2027First wave of new GPU deployments begins expected to lean on Blackwell Ultra availability
  4. 2028Second wave, including Rubin and Rubin Ultra timed to Nvidia's next architecture ramp

What is the Trainium and NVLink Fusion connection, and why does it matter more than it looks?

Most coverage of this deal treats it as a straightforward GPU purchase order. The more interesting piece is buried in the announcement: AWS's next-generation Trainium chips, its homegrown silicon built specifically to reduce dependence on Nvidia, will get access to Nvidia's new custom high-bandwidth memory, NVHBM, developed with memory suppliers as part of this deal. NVLink Fusion, Nvidia's high-speed interconnect technology, is expanding its integration with AWS's custom chips at the same time.

Sit with that for a second. AWS is buying record volumes of Nvidia GPUs while also asking Nvidia to help make AWS's Nvidia-competing chips faster. That's not a contradiction, it's coopetition, and it's a more honest picture of how the AI hardware market actually works than the simple "hyperscaler versus Nvidia" framing most people default to. AWS needs Trainium to be commercially competitive on cost and power efficiency, because Nvidia GPUs are expensive and supply-constrained. But AWS also needs raw GPU volume today, right now, because Trainium alone can't cover the demand AWS customers are generating. Nvidia, for its part, benefits either way. If Trainium succeeds, Nvidia's memory and interconnect technology still gets a design win. If Trainium falls short, AWS keeps buying GPUs at scale. There's no version of this where Nvidia loses.

Who actually benefits from this deal

AWS customers building large models or running agentic products get more compute headroom starting in 2027, which matters for anyone who has hit capacity limits on Nvidia instances inside AWS over the past year. Waitlists for the newest GPU instance types on AWS have been a recurring complaint from AI startups, and this order is a direct response to that pressure.

Nvidia gets a locked-in, multi-year order that spans three chip generations before those generations have even fully shipped. That's about as strong a revenue signal as a chipmaker can get: a customer committing to buy hardware that doesn't exist in volume yet. AWS gets to tell its own customers, and its own investors, that compute scarcity isn't going to be the bottleneck on AWS in 2027 and 2028, at least not for lack of trying.

RelatedNvidia's Groq 3 LPU Bets Inference Isn't a GPU Problem

What it means for the market

For Nvidia (NVDA), this is a multi-year revenue commitment locked in well ahead of the actual GPU generations shipping, the kind of forward visibility that chip companies rarely get and investors tend to reward. For Amazon (AMZN), the signal runs the other direction: committing to 2 million more GPUs is Amazon telling the market it expects AI compute demand to keep climbing through 2028, which is also Amazon's own justification for the capital expenditure this order requires.

The read for anyone watching AI infrastructure spending is straightforward. Sustained, expanding hyperscaler capex commitments are one of the clearest tells available that AI infrastructure spending isn't slowing down, at least not yet. This is market analysis, not investment advice, and nothing here should be read as a recommendation to buy or sell either stock. But when the company buying the chips says its last big order ran out early, that's a data point worth more than most quarterly earnings-call soundbites about AI momentum.

What to watch · 2027-2028
  • Blackwell Ultra deployment pace. Watch whether AWS gets meaningful Blackwell Ultra capacity online in 2027 as planned, or whether supply constraints push the timeline into 2028.
  • Trainium's real-world performance with NVHBM. If Trainium chips paired with Nvidia's memory tech close the performance gap with pure Nvidia GPU instances, that's a bigger story than the GPU order itself.
  • Whether a third top-up follows. If AWS's 2 million GPU commitment also runs out early, that's confirmation this isn't a one-time catch-up order but a structural underestimate of AI compute demand across the industry.
  • Rubin and Rubin Ultra shipping schedules. Both chip families still need to reach volume production; any slip in Nvidia's roadmap pushes AWS's 2028 wave along with it.

Our take

Two million additional GPUs, on top of a million that already ran out, is not the profile of a market that's cooling off. It's the profile of a market where every major buyer keeps discovering their own demand forecasts were too conservative, which is a genuinely different problem than a bubble popping. Bubbles usually show up as supply outrunning demand, with unsold capacity sitting idle. What's happening here looks like the opposite: demand outrunning even aggressive supply commitments, repeatedly, across multiple companies.

That doesn't mean the spending is risk-free. Two million GPUs is an enormous amount of capital tied up in hardware that depreciates, and if AI adoption growth ever plateaus even modestly, all four major hyperscalers are sitting on capacity built for a demand curve that assumed no plateau. The Trainium integration is the more interesting hedge to watch, because it's AWS quietly building an exit ramp even while it signs the biggest GPU order of its history. A company that fully believed Nvidia GPUs would remain irreplaceable wouldn't bother co-developing custom memory tech for its own competing chip. AWS is buying the insurance and the safety net at the same time, and that combination, more than the 2 million number itself, is the real story here.

Primary sources

Original analysis by GenZTech, based on official Nvidia and AWS announcements.