PantheonGet Early Access

NVIDIA GB300 NVL72 vs GB200 NVL72

TL;DR

The GB300 NVL72 is NVIDIA’s current rack-scale flagship on Grace Blackwell Ultra — 72 GPUs at 288 GB of HBM3e each, roughly 20.7 TB of HBM3e per rack. The GB200 NVL72 is the prior Grace Blackwell rack at 192 GB per GPU (≈13.8 TB per rack). Both wire 72 GPUs into one fifth-generation NVLink domain at 1.8 TB/s; the GB300 is a memory step-up, and the market has rotated to it for the largest models. We supply the GB300 NVL72.

On this page

The short answer

Both are rack-scale NVL72 systems — 72 GPUs and 36 Grace CPUs wired into a single fifth-generation NVLink domain — but a generation apart. The GB200 NVL72 is built on the original Grace Blackwell (GB200); the GB300 NVL72 is the newer Grace Blackwell Ultra (GB300).

The defining difference is memory. Each GB300 GPU carries 288 GB of HBM3e versus 192 GB on the GB200 — so a full GB300 rack holds roughly 20.7 TB of HBM3e against about 13.8 TB on the GB200. NVLink, GPU count, and the ConnectX-8 scale-out plane are the same across both. Because the largest models are increasingly memory-bound, the market has rotated to the higher-memory GB300, and that is the rack-scale system we supply.

Spec comparison

How the two rack-scale generations line up. The headline move is memory — 288 GB vs 192 GB per GPU (≈20.7 TB vs ≈13.8 TB per rack) — on top of an otherwise-shared NVL72 platform: 72 GPUs, one fifth-generation NVLink domain, and ConnectX-8 scale-out.

SpecGB300 NVL72GB200 NVL72
GPUGB300 · Grace Blackwell UltraGB200 · Grace Blackwell
HBM3e per GPU288 GB192 GB
HBM3e per rack≈20.7 TB≈13.8 TB
Peak FP8 (dense)~5 PFLOPS FP8 · ~15 PFLOPS FP4 per GPU (~1.1 EFLOPS FP4/rack)~5 PFLOPS FP8 · ~10 PFLOPS FP4 per GPU (~0.72 EFLOPS FP4/rack)
GPUs per rack7272
NVLink5th-gen · 1.8 TB/s5th-gen · 1.8 TB/s
Scale-outConnectX-8 · 800 Gb/sConnectX-8 · 800 Gb/s

What changed: memory

The GB300 is a memory refresh of the rack-scale platform rather than a new fabric. Grace Blackwell Ultra raises on-package HBM3e from 192 GB to 288 GB per GPU — a 50% step — so the pooled memory of a single NVL72 rack grows from roughly 13.8 TB to ≈20.7 TB of HBM3e.

Why it matters: a rack-scale NVL72 behaves like one very large accelerator, and its ceiling is set by how much model and long-context KV cache stay resident across the NVLink domain. More HBM3e per GPU lets a larger model shard across the same 72 GPUs with less spillover, which is exactly the constraint the frontier of AI serving keeps running into. The NVLink generation (fifth-gen, 1.8 TB/s per GPU) and the ConnectX-8 scale-out plane (800 Gb/s per GPU) are common to both — the GB300’s advantage is memory, not fabric.

On compute the two are nearly matched — both deliver ~5 PFLOPS of dense FP8 per GPU — but Blackwell Ultra lifts dense FP4 from ~10 to ~15 PFLOPS per GPU, so a full 72-GPU rack reaches roughly ~1.1 EFLOPS of dense FP4 against about ~0.72 EFLOPS on the GB200 — the throughput lever for large-scale inference. Figures are dense; sparsity doubles them.

Which to deploy

For a new rack-scale buildout, the GB300 NVL72 is the current generation and the one to deploy — more HBM3e per GPU, more pooled memory per rack, and the platform NVIDIA and its integrators are shipping now. The GB200 NVL72 was the prior rack-scale flagship; the market has rotated to the GB300’s higher memory, and that is what we carry — factory-integrated from Supermicro, Lenovo, HPE, and Pegatron.

For the full rack composition — 18 compute nodes of 4 GPUs, 36 Grace CPUs, the in-rack CDU, and scale-out fabric — see the GB300 NVL72 reference architecture. Pricing across every configuration is quoted per configuration on request — tell us the cluster you are standing up and we will return pricing and availability. Browse the GPU catalog to compare integrations.

Frequently asked questions

What’s the difference between GB300 and GB200 NVL72?

Both are rack-scale NVL72 systems — 72 GPUs and 36 Grace CPUs in one fifth-generation NVLink domain at 1.8 TB/s per GPU — but a generation apart. The GB200 NVL72 uses the original Grace Blackwell (GB200) at 192 GB of HBM3e per GPU; the GB300 NVL72 uses Grace Blackwell Ultra (GB300) at 288 GB per GPU. The GB300 is a memory step-up on an otherwise-shared platform (same NVLink, same 72-GPU count, same ConnectX-8 scale-out).

How much more memory does the GB300 have?

Each GB300 GPU carries 288 GB of HBM3e versus 192 GB on the GB200 — a 50% increase. Across a full NVL72 rack that is roughly 20.7 TB of HBM3e on the GB300 against about 13.8 TB on the GB200, addressable across a single NVLink domain.

Can I still buy GB200 NVL72?

The market has rotated to the GB300 NVL72 — the higher-memory Grace Blackwell Ultra generation — which is the rack-scale system we supply, factory-integrated from Supermicro, Lenovo, HPE, and Pegatron. If your requirement specifically calls for the prior-generation GB200, tell us and we will advise; for new rack-scale buildouts we point customers to the GB300.

How much does a GB300 NVL72 cost?

The GB300 NVL72 is quoted per configuration on request — the figure depends on integrator, cooling, scale-out fabric, and support scope. Tell us the cluster you are standing up and we will return pricing and availability.

Related

Last updated