NVIDIA GB200 NVL72 Specs & Datasheet (72-GPU Rack)
TL;DR
The NVIDIA GB200 NVL72 is a liquid-cooled rack that fuses 72 Blackwell GPUs and 36 Grace CPUs — 36 GB200 Grace Blackwell superchips across 18 compute trays — into a single fifth-generation NVLink domain. Each GPU carries 192 GB of HBM3e, so a full rack holds roughly 13.8 TB of pooled high-bandwidth memory behaving as one coherent accelerator. It ships factory-integrated from Supermicro, Dell, and HPE; configuration, site power, and pricing are confirmed at quote.
On this page
What the GB200 NVL72 is
The GB200 NVL72 is NVIDIA’s Grace Blackwell rack-scale system: a single, factory-integrated, liquid-cooled rack that connects 72 B200 GPUs and 36 Grace CPUs into one coherent accelerator over a fifth-generation NVLink fabric.
The building block is the GB200 Grace Blackwell superchip — one Grace CPU paired with two Blackwell GPUs on a single module. Thirty-six of those superchips, two per compute tray across 18 trays, give the rack its 72 GPUs and 36 CPUs. Unlike an HGX B200 node that you rack yourself in a standard chassis, the NVL72 arrives as a complete rack: compute trays, NVLink switch trays, power shelves, networking, and liquid cooling integrated at the factory.
The result is roughly 13.8 TB of HBM3e addressable across a single NVLink domain. It is the Blackwell predecessor to the Blackwell Ultra GB300 NVL72.
Full spec sheet
The GB200 NVL72 rack specification, aggregated from the current integrator builds. The per-GPU and per-rack GPU, memory, and NVLink figures below are fixed characteristics of the platform; site power draw, rack weight, and final configuration vary by integrator and facility, and are confirmed at quote.
| Spec | GB200 NVL72 |
|---|---|
| GPU architecture | NVIDIA B200 · Blackwell |
| GPUs per rack | 72 (18 compute trays × 4) |
| Grace CPUs | 36 |
| Superchips | 36× GB200 (1 Grace CPU + 2 Blackwell GPUs each) |
| HBM3e per GPU | 192 GB |
| HBM3e per rack | ≈13.8 TB |
| Memory bandwidth | ~8 TB/s per GPU |
| Peak FP8 (dense) | ~5 PFLOPS FP8 · ~10 PFLOPS FP4 per GPU (~360 PFLOPS / ~0.72 EFLOPS FP4 per rack, derived) |
| NVLink | 5th-gen · 1.8 TB/s per GPU · 72-GPU domain |
| NVLink aggregate | ~130 TB/s all-to-all in-rack |
| Scale-out | ConnectX-8 · 800 Gb/s per GPU |
| GPU TDP | ~1,200 W per GPU |
| Cooling | Liquid · in-rack CDU |
| Form factor | Single NVL72 rack, fully integrated |
| Integrators | Supermicro · Dell · HPE |
Memory & bandwidth
Each Blackwell GPU in the GB200 NVL72 carries 192 GB of HBM3e — the on-package memory available to the model. Across 72 GPUs that is roughly 13.8 TB of HBM3e per rack (72 × 192 GB = 13,824 GB), pooled and addressable over NVLink so a frontier-scale model and its long-context KV cache can span the entire rack.
Per-GPU memory bandwidth is about ~8 TB/s. What makes that number matter is the fabric: because all 72 GPUs sit in one NVLink domain, the rack behaves like a single very large accelerator rather than 72 networked cards, and the practical ceiling on model size is the pooled 13.8 TB rather than any one GPU’s 192 GB.
On compute, each GPU delivers ~5 PFLOPS of dense FP8 and ~10 PFLOPS of dense FP4, so a full rack reaches roughly ~360 PFLOPS of dense FP8 and ~0.72 EFLOPS of dense FP4. All figures here are dense; NVIDIA’s published rack headline of 1,440 PFLOPS FP4 is the "with sparsity" number, i.e. 2× the dense figure.
Rack topology
The NVL72 rack is built from three repeating elements plus power and cooling:
- 18 compute trays. Each holds two GB200 superchips — 2 Grace CPUs and 4 Blackwell GPUs per tray — for 72 GPUs and 36 Grace CPUs across the rack.
- 9 NVLink switch trays. These carry the fifth-generation NVLink switch silicon that gives every GPU all-to-all connectivity to every other GPU in the rack, for roughly ~130 TB/s of aggregate in-rack NVLink bandwidth.
- Power shelves. Rack-level shelves convert facility input and feed the compute and switch trays over a common busbar, rather than per-server PSUs.
The cabling is what distinguishes NVL72 from a cluster of nodes: an NVLink spine connects the compute trays to the switch trays inside the rack, so GPU-to-GPU traffic never leaves the chassis for the network fabric. That is the topology that makes 72 GPUs act as one NVLink domain — and it is why the rack ships factory-integrated rather than assembled on site.
Power & cooling
The GB200 NVL72 is a high-density, liquid-cooled rack served by an in-rack coolant distribution unit (CDU) with direct-to-chip liquid cooling across the Grace CPUs, Blackwell GPUs, and NVLink switches. At roughly 1,200 W per GPU for 72 GPUs — before CPUs, switches, and conversion losses — liquid cooling is not optional; it is what makes the density thermally viable at all.
NVIDIA and its integrators commonly describe the NVL72 as an approximately 120 kW rack, and that is the right order of magnitude for facility planning. Treat it as a planning figure rather than a spec: the exact site power draw and rack weight depend on the integrator build, the configuration, and the facility, and are confirmed at quote. Plan for facility water and high-density power distribution.
Scale-out networking
Beyond the in-rack NVLink domain, the GB200 NVL72 scales out to multi-rack clusters over NVIDIA ConnectX-8 SuperNICs at 800 Gb/s per GPU, on Quantum-X800 InfiniBand or Spectrum-X Ethernet fabrics. NVLink handles the dense all-to-all traffic inside each rack; ConnectX stitches many racks into a larger training cluster.
GB200 NVL72 vs GB300 NVL72
The GB300 NVL72 is the Blackwell Ultra successor on the same rack-scale platform. The defining difference is memory: 288 GB of HBM3e per GPU versus 192 GB, so a GB300 rack holds roughly 20.7 TB against the GB200’s 13.8 TB. Blackwell Ultra also lifts dense FP4 from ~10 to ~15 PFLOPS per GPU (~1.1 EFLOPS vs ~0.72 EFLOPS per rack), while dense FP8 is unchanged at ~5 PFLOPS.
Everything else is shared: 72 GPUs, 36 Grace CPUs, one fifth-generation NVLink domain at 1.8 TB/s per GPU, and ConnectX-8 scale-out. For the full side-by-side see GB300 NVL72 vs GB200 NVL72.
Available integrators
The GB200 NVL72 ships as a factory-integrated rack from several OEMs, each with its own chassis, cooling detail, and support model. The GPU, memory, and NVLink specifications are common across builds; chassis, cooling, warranty, and lead time differ by integrator. Each build below links to its full specification.
Supermicro
Supermicro SRS-GB200-NVL72
NVIDIA GB200 NVL72 rack-scale GPU system — 72 Blackwell GPUs and 36 Grace CPUs in a single liquid-cooled 19-inch Supermicro NVL72 rack.
Dell
Dell PowerEdge XE9712
NVIDIA GB200 NVL72 rack-scale GPU system — 72 Blackwell GPUs and 36 Grace CPUs in a single liquid-cooled Dell PowerEdge XE9712 (IR7000) rack.
HPE
HPE
NVIDIA GB200 NVL72 rack-scale GPU system — 72 Blackwell GPUs and 36 Grace CPUs in a single liquid-cooled NVL72 rack integrated by HPE.
Procurement
Every GB200 NVL72 build is quoted per configuration on request — the figure depends on integrator, cooling, scale-out fabric, and support scope. Lead times vary by integrator and are confirmed at quote. Tell us the cluster you are standing up and we will return pricing, availability, and a facility-fit review.
For a new rack-scale buildout the current generation is the GB300 NVL72, and that is where we point most buyers; we carry the GB200 NVL72 for requirements that specifically call for it. Browse the GPU catalog to compare integrations.
Frequently asked questions
How much memory does a GB200 NVL72 have?
Each Blackwell GPU carries 192 GB of HBM3e, so a full GB200 NVL72 rack holds roughly 13.8 TB of HBM3e (72 × 192 GB = 13,824 GB) — pooled and addressable across all 72 GPUs over a single fifth-generation NVLink domain at 1.8 TB/s per GPU.
How many GPUs are in a GB200 NVL72?
A GB200 NVL72 rack contains 72 NVIDIA Blackwell GPUs and 36 Grace CPUs, built from 36 GB200 Grace Blackwell superchips (one Grace CPU plus two Blackwell GPUs each) arranged two per tray across 18 compute trays.
What is the GB200 NVL72’s power consumption?
The NVL72 is commonly described as an approximately 120 kW rack, which is the right order of magnitude for facility planning — roughly 1,200 W per GPU across 72 GPUs, before CPUs, NVLink switches, and conversion losses. The exact site power draw depends on the integrator build and facility and is confirmed at quote.
What is the NVL72 rack topology?
Eighteen compute trays (two GB200 superchips each, so 4 GPUs and 2 Grace CPUs per tray) and nine NVLink switch trays, connected by an in-rack NVLink spine and fed by rack-level power shelves over a common busbar. The switch trays give every GPU all-to-all connectivity to every other GPU, for roughly 130 TB/s of aggregate in-rack NVLink bandwidth.
What is the difference between GB200 and GB300 NVL72?
The GB300 NVL72 is the Blackwell Ultra generation: 288 GB of HBM3e per GPU versus the GB200’s 192 GB (≈20.7 TB vs ≈13.8 TB per rack), and ~15 vs ~10 PFLOPS of dense FP4 per GPU. GPU count, Grace CPU count, the fifth-generation NVLink domain, and ConnectX-8 scale-out are the same on both.
How much does a GB200 NVL72 cost?
The GB200 NVL72 is quoted per configuration on request — the figure depends on integrator, cooling, scale-out fabric, and support scope. Tell us the cluster you are standing up and we will return pricing and availability.
Related
Last updated