PantheonGet Early Access

NVIDIA GB300 NVL72 Specs & Datasheet (72-GPU Rack)

TL;DR

The NVIDIA GB300 NVL72 is a liquid-cooled rack that fuses 72 Grace Blackwell Ultra GPUs and 36 Grace CPUs — arranged two GPUs per superchip across 18 compute trays — into a single fifth-generation NVLink domain. Each GPU carries 288 GB of HBM3e, so a full rack holds roughly 20.7 TB of pooled high-bandwidth memory behaving as one coherent accelerator. It ships factory-integrated from Supermicro, Lenovo, HPE, and Pegatron; configuration, site power, and pricing are confirmed at quote.

On this page

What the GB300 NVL72 is

The GB300 NVL72 is NVIDIA’s Grace Blackwell Ultra rack-scale system: a single, factory-integrated, liquid-cooled rack that connects 72 GB300 GPUs and 36 Grace CPUs into one coherent accelerator over a fifth-generation NVLink fabric.

The building block is the GB300 Grace Blackwell Ultra superchip — one Grace CPU paired with two Blackwell Ultra GPUs — arranged two superchips per compute tray across 18 trays, which gives the rack its 72 GPUs and 36 CPUs. Unlike an HGX B300 node that you rack yourself in a standard chassis, the NVL72 arrives as a complete rack: compute trays, NVLink switch trays, power shelves, networking, and liquid cooling integrated at the factory.

The result is roughly 20.7 TB of HBM3e addressable across a single NVLink domain. It is the Blackwell Ultra successor to the GB200 NVL72, and the current generation for a new rack-scale buildout. For the design-level view see the GB300 NVL72 reference architecture.

Full spec sheet

The GB300 NVL72 rack specification, aggregated from the current integrator builds. The per-GPU and per-rack GPU, memory, and NVLink figures below are fixed characteristics of the platform; site power draw, rack weight, and final configuration vary by integrator and facility, and are confirmed at quote.

SpecGB300 NVL72
GPU architectureNVIDIA GB300 · Grace Blackwell Ultra
GPUs per rack72 (18 compute trays × 4)
Grace CPUs36
Superchips36× GB300 (1 Grace CPU + 2 Blackwell Ultra GPUs each)
HBM3e per GPU288 GB
HBM3e per rack≈20.7 TB
Memory bandwidth~8 TB/s per GPU
Peak FP8 (dense)~5 PFLOPS FP8 · ~15 PFLOPS FP4 per GPU (~360 PFLOPS / ~1.1 EFLOPS FP4 per rack, derived)
NVLink5th-gen · 1.8 TB/s per GPU · 72-GPU domain
NVLink aggregate~130 TB/s all-to-all in-rack
Scale-outConnectX-8 · 800 Gb/s per GPU
GPU TDP~1,400 W per GPU
CoolingLiquid · in-rack CDU
Form factorSingle NVL72 rack, fully integrated
IntegratorsSupermicro · Lenovo · HPE · Pegatron

Memory & bandwidth

Each Blackwell Ultra GPU in the GB300 NVL72 carries 288 GB of HBM3e — the Blackwell Ultra memory step-up over the 192 GB GB200. Across 72 GPUs that is roughly 20.7 TB of HBM3e per rack (72 × 288 GB = 20,736 GB), pooled and addressable over NVLink so a frontier-scale model and its long-context KV cache can span the entire rack.

Per-GPU memory bandwidth is about ~8 TB/s. What makes that number matter is the fabric: because all 72 GPUs sit in one NVLink domain, the rack behaves like a single very large accelerator rather than 72 networked cards, and the practical ceiling on model size is the pooled 20.7 TB rather than any one GPU’s 288 GB.

On compute, each GPU delivers ~5 PFLOPS of dense FP8 and ~15 PFLOPS of dense NVFP4, so a full rack reaches roughly ~360 PFLOPS of dense FP8 and ~1.1 EFLOPS of dense FP4 — the low-precision path NVIDIA positions for large-scale reasoning inference. All figures here are dense; NVIDIA’s "with sparsity" rack headline is roughly 2× the dense figure.

Rack topology

The NVL72 rack is built from three repeating elements plus power and cooling:

  • 18 compute trays. Each holds two GB300 superchips — 2 Grace CPUs and 4 Blackwell Ultra GPUs per tray — for 72 GPUs and 36 Grace CPUs across the rack.
  • 9 NVLink switch trays. These carry the fifth-generation NVLink switch silicon that gives every GPU all-to-all connectivity to every other GPU in the rack, for roughly ~130 TB/s of aggregate in-rack NVLink bandwidth.
  • Power shelves. Rack-level shelves convert facility input and feed the compute and switch trays over a common busbar, rather than per-server PSUs.

The cabling is what distinguishes NVL72 from a cluster of nodes: an NVLink spine connects the compute trays to the switch trays inside the rack, so GPU-to-GPU traffic never leaves the chassis for the network fabric. That all-to-all NVLink topology is what makes 72 GPUs act as one NVLink domain — and it is why the rack ships factory-integrated rather than assembled on site.

Power & cooling

The GB300 NVL72 is a high-density, liquid-cooled rack served by an in-rack coolant distribution unit (CDU) with direct-to-chip liquid cooling across the Grace CPUs, Blackwell Ultra GPUs, and NVLink switches. At roughly 1,400 W per GPU for 72 GPUs — before CPUs, switches, and conversion losses — liquid cooling is not optional; it is what makes the density thermally viable at all.

Blackwell Ultra raises per-GPU TDP over the GB200 NVL72 (~1,200 W), so plan for a rack meaningfully denser than that platform’s ~120 kW; some integrator builds specify an in-rack CDU rated around 250 kW of cooling capacity. Treat any single figure as a planning number: the exact site power draw and rack weight depend on the integrator build, the configuration, and the facility, and are confirmed at quote. Plan for facility water and high-density power distribution.

Scale-out networking

Beyond the in-rack NVLink domain, the GB300 NVL72 scales out to multi-rack clusters over NVIDIA ConnectX-8 SuperNICs at 800 Gb/s per GPU, on Quantum-X800 InfiniBand or Spectrum-X Ethernet fabrics. NVLink handles the dense all-to-all traffic inside each rack; ConnectX stitches many racks into a larger SuperPOD-class training cluster.

GB300 NVL72 vs GB200 NVL72

The GB300 NVL72 is the Blackwell Ultra generation on the same rack-scale platform. The defining difference is memory: 288 GB of HBM3e per GPU versus 192 GB, so a GB300 rack holds roughly 20.7 TB against the GB200’s 13.8 TB. Blackwell Ultra also lifts dense FP4 from ~10 to ~15 PFLOPS per GPU (~1.1 EFLOPS vs ~0.72 EFLOPS per rack), while dense FP8 is unchanged at ~5 PFLOPS; per-GPU TDP rises from ~1,200 W to ~1,400 W.

Everything else is shared: 72 GPUs, 36 Grace CPUs, one fifth-generation NVLink domain at 1.8 TB/s per GPU, and ConnectX-8 scale-out. For the full side-by-side see GB300 NVL72 vs GB200 NVL72.

Available integrators

The GB300 NVL72 ships as a factory-integrated rack from several OEMs, each with its own chassis, cooling detail, and support model. The GPU, memory, and NVLink specifications are common across builds; chassis, cooling, warranty, and lead time differ by integrator. Each build below links to its full specification.

Supermicro

Supermicro MGX

NVIDIA GB300 NVL72 rack-scale GPU system — 72 Blackwell Ultra GPUs and 36 Grace CPUs in a single, fully-integrated, liquid-cooled 19-inch Supermicro MGX rack.

Liquid-cooled
View specs

HPE

HPE 48U MGX

NVIDIA GB300 NVL72 rack-scale GPU system — 72 Blackwell Ultra GPUs and 36 Grace CPUs factory-integrated by HPE in a 48U liquid-cooled MGX-compliant rack.

Liquid-cooled
View specs

Lenovo

Lenovo 48U MGX

NVIDIA GB300 NVL72 rack-scale GPU system — 72 Blackwell Ultra GPUs and 36 Grace CPUs on the Lenovo 48U MGX reference rack, factory-integrated with direct water cooling.

Liquid-cooled
View specs

Pegatron

Pegatron

NVIDIA GB300 NVL72 rack-scale GPU system — 72 Blackwell Ultra GPUs and 36 Grace CPUs in a single, fully-integrated, liquid-cooled Pegatron NVL72 rack.

Liquid-cooled
View specs

Procurement

Every GB300 NVL72 build is quoted per configuration on request — the figure depends on integrator, cooling, scale-out fabric, and support scope. Lead times vary by integrator and are confirmed at quote. Tell us the cluster you are standing up and we will return pricing, availability, and a facility-fit review. Browse the GPU catalog to compare integrations.

Frequently asked questions

How much memory does a GB300 NVL72 have?

Each Blackwell Ultra GPU carries 288 GB of HBM3e, so a full GB300 NVL72 rack holds roughly 20.7 TB of HBM3e (72 × 288 GB = 20,736 GB) — pooled and addressable across all 72 GPUs over a single fifth-generation NVLink domain at 1.8 TB/s per GPU.

How many GPUs are in a GB300 NVL72?

A GB300 NVL72 rack contains 72 NVIDIA Blackwell Ultra GPUs and 36 Grace CPUs, built from 36 GB300 Grace Blackwell Ultra superchips (one Grace CPU plus two Blackwell Ultra GPUs each) arranged two per tray across 18 compute trays.

What is the GB300 NVL72’s power consumption?

The GB300 NVL72 is a high-density liquid-cooled rack at roughly 1,400 W per GPU across 72 GPUs — denser than the GB200 NVL72’s ~120 kW, with some integrator builds specifying an in-rack CDU rated around 250 kW of cooling capacity. The exact site power draw depends on the integrator build and facility and is confirmed at quote.

What is the NVL72 rack topology?

Eighteen compute trays (two GB300 superchips each, so 4 GPUs and 2 Grace CPUs per tray) and nine NVLink switch trays, connected by an in-rack NVLink spine and fed by rack-level power shelves over a common busbar. The switch trays give every GPU all-to-all connectivity to every other GPU, for roughly 130 TB/s of aggregate in-rack NVLink bandwidth.

What is the difference between GB300 and GB200 NVL72?

The GB300 NVL72 is the Blackwell Ultra generation: 288 GB of HBM3e per GPU versus the GB200’s 192 GB (≈20.7 TB vs ≈13.8 TB per rack), ~15 vs ~10 PFLOPS of dense FP4 per GPU, and ~1,400 vs ~1,200 W per GPU. GPU count, Grace CPU count, the fifth-generation NVLink domain, and ConnectX-8 scale-out are the same on both.

How much does a GB300 NVL72 cost?

The GB300 NVL72 is quoted per configuration on request — the figure depends on integrator, cooling, scale-out fabric, and support scope. Tell us the cluster you are standing up and we will return pricing and availability.

Related

Last updated