Put simply, the Rubin Ultra HBM despec is a supply issue. Not a demand issue.
I think as much additional HBM will be produced as the supply headroom freed up by the shift from 12Hi to 8Hi.
Nvidia is said to be weighing a plan to halve the memory that goes into Rubin Ultra, its 2027 accelerator. The idea is to bring a chip that was to carry 384GB down to 192GB.
TrendForce said much the same thing in a report dated July 31. Nvidia is still weighing four possible configurations, and what settles the final spec is not a performance target but how the memory makers allocate their wafers. Because it is Nvidia doing the cutting, the news reads easily as a sign that AI demand is cooling.
HBM is DRAM stacked in layers and placed right beside the GPU. How many layers are stacked is what 12 high and 8 high refer to. Twelve dies on the same footprint hold one and a half times the capacity of eight, and consume proportionally more DRAM.

The Rubin Ultra tray Nvidia demonstrated in March held four compute chiplets in one package with 1TB of memory on top. It has since come down to two chiplets, and now the layer count on the remaining memory is on the table too.
Nvidia demonstrates Rubin Ultra tray, the world's first AI GPU with 1TB of HBM4E (Tom's Hardware, 2026-03-17)Fewer layers means less DRAM inside each chip. That DRAM does not disappear. It goes straight into building other HBM.
DRAM is getting scarcer while Nvidia trims capacity. That runs the opposite way from an explanation built on cooling demand.
Retrieved 2026-08-04 · 1TB from Tom's Hardware coverage of the 2026-03-17 demonstration · the 2027 figure from Apacer's chief executive, Tom's Hardware 2026-07-29 · both are press citations
DRAM chip supply to module makers could drop by more than 70% year on year in 2027, says Apacer CEO (Tom's Hardware, 2026-07-29)The head of the Taiwanese module maker Apacer said the DRAM reaching module makers next year could fall to around 30% of this year's level. Roughly 60% of DRAM capacity already goes to server applications. Phone and PC makers are the first to be pushed back in the queue for what is left. The side that is short of parts is the side writing the spec.
Nvidia halved the memory on each chip. The largest customer is cutting its order, and that is where an HBM price break begins.
What was cut is one chip's share, not the total. Build more HBM from the freed DRAM and the bits sold stay the same.
The same decision is read by one side as a smaller order book and by the other as rationing a part that is in short supply.
Citing DigiTimes, the original post says 2027 DRAM and HBM capacity has been fully booked ahead of schedule. Buyers are said to be receiving only 60 to 70 percent of the volumes they first requested. Cloud providers and the large AI buyers are filled first, and everyone else waits. If the lines really are full, cutting the layer count looks less like a decision to buy less and more like a decision to build more accelerators out of the same parts.
Which of the two it was will be answered by the HBM bits sold in 2027. If capacity per chip falls and total bits shipped still rise, what was short was the part. If capacity and bits fall together, what shrank was the order. Until then the readable signal is price. If the lines really are booked out, memory does not get cheaper after the spec comes down.
Retrieved 2026-08-04 · the despec plan and the 2027 booking picture are relayed by the original post from a brokerage note and DigiTimes · the 1TB figure is from coverage of the 2026-03-17 demonstration