Nvidia details the Rubin platform ahead of its 2026 launch
Nvidia has confirmed final specifications for Rubin, the GPU architecture succeeding Blackwell in its data-center lineup. The briefing focused on three changes: a move to HBM4 memory, a next-generation NVLink interconnect, and a redesigned rack-scale system built around the new Vera CPU pairing.
The headline numbers
- HBM4 memory bandwidth roughly 2x Blackwell's HBM3e, addressing the memory-bandwidth bottleneck that dominates large-model inference cost more than raw compute does
- NVLink 6 doubling per-GPU interconnect bandwidth, relevant specifically for the all-to-all communication pattern in mixture-of-experts model inference
- A new rack-scale reference design (Vera Rubin NVL144) targeting significantly higher performance-per-watt than the current Blackwell NVL72 racks
For anyone renting inference capacity rather than buying hardware, the number that matters is performance-per-dollar on the specific workload you run — MoE inference is far more memory-bandwidth-bound than the dense-model benchmarks usually quoted, so Rubin's real-world advantage will vary a lot by model architecture.
Availability
Nvidia is targeting broad cloud-provider availability in the following year, with the usual pattern of large hyperscalers getting early allocation ahead of general access. For teams planning infrastructure budgets now, the practical takeaway is timing a hardware refresh around this transition rather than mid-cycle, since price-per-token on current-generation capacity typically drops once a new architecture starts shipping in volume.
Source: www.nvidia.com