PantheonGet Early Access

NVIDIA HGX H200 Specs & Datasheet (8-GPU Node)

TL;DR

The NVIDIA HGX H200 is an 8-GPU Hopper node: eight H200 SXM GPUs with 141 GB of HBM3e each — 1,128 GB per node — connected over a fourth-generation NVLink baseboard at 900 GB/s per GPU, with roughly 4.8 TB/s of memory bandwidth per GPU for memory-bound inference and large-model training. It ships factory-integrated from several OEMs (HPE, Dell, Supermicro, Gigabyte) in air-cooled form factors; configuration, power, and pricing are confirmed at quote.

On this page

What the HGX H200 is

The HGX H200 is NVIDIA’s 8-GPU Hopper server platform — the memory-upgraded successor to the HGX H100, built for large-model training and, especially, memory-bound inference. Eight H200 SXM GPUs sit on a single fourth-generation NVLink baseboard, so the node behaves as one tightly-coupled 8-GPU accelerator rather than eight discrete cards.

The H200 shares the Hopper compute architecture of the H100 but swaps in HBM3e memory: 141 GB per GPU instead of the H100’s 80 GB of HBM3, at higher bandwidth. That makes the HGX H200 a node you rack in a standard air-cooled data center — a 5U–6U chassis that drops into a conventional rack — and the mainstream Hopper platform for inference workloads whose bottleneck is memory capacity and bandwidth rather than raw compute. It is the Hopper predecessor to the Blackwell HGX B200 8-GPU node.

Full spec sheet

The HGX H200 8-GPU node specification, aggregated from the current integrator builds (HPE, Dell, Supermicro, Gigabyte). The per-GPU and per-node GPU, memory, and NVLink figures below are fixed characteristics of the platform; chassis, CPU, cooling, site power draw, and final configuration vary by integrator and are confirmed at quote.

SpecHGX H200 (8-GPU node)
GPU architectureNVIDIA H200 · Hopper
GPUs per node8× HGX H200 SXM
HBM3e per GPU141 GB
HBM3e per node1,128 GB
Memory bandwidth~4.8 TB/s per GPU
Peak FP8 (dense)1,979 TFLOPS (~2 PFLOPS) per GPU · no FP4 (~15.8 PFLOPS per node, derived)
NVLink4th-gen · 900 GB/s per GPU
Scale-outConnectX-7 NDR400 · 400 Gb/s (~3.2 Tb/s/node)
CoolingAir (5U–6U), by integrator · DLC option on the HPE build
IntegratorsHPE · Dell · Supermicro · Gigabyte

Memory & bandwidth (VRAM)

Each H200 GPU in the HGX H200 carries 141 GB of HBM3e — the on-package GPU memory (VRAM) available to the model. Across the 8-GPU baseboard that is 1,128 GB of HBM3e per node, pooled and addressable over NVLink so a large model and its long-context KV cache can span all eight GPUs.

Memory bandwidth is roughly ~4.8 TB/s per GPU, and it is the headline reason to choose H200 over H100: the H200 keeps the same Hopper compute but nearly doubles memory capacity (141 GB vs 80 GB) and lifts bandwidth by moving from HBM3 to HBM3e. For memory-bound inference — where throughput is gated by how fast weights and KV cache stream out of VRAM — that extra capacity and bandwidth is what lets a single HGX H200 node serve larger models, longer contexts, and more concurrent requests than the H100 it replaces.

Compute is unchanged from the H100 — about 1,979 TFLOPS (~2 PFLOPS) of dense FP8 per GPU (~15.8 PFLOPS across the node) — since the H200 is a memory upgrade of the same Hopper die; the gain is entirely capacity and bandwidth. Sparsity doubles the FP8 figure.

Inside the node, the eight H200 GPUs are wired over a fourth-generation NVLink baseboard at 900 GB/s per GPU, giving all-to-all connectivity across the 8-GPU domain with minimal communication overhead — the same NVLink generation and bandwidth as the HGX H100.

To scale beyond a single node, the HGX H200 uses NVIDIA ConnectX-7 NDR400 adapters at 400 Gb/s each with GPUDirect — eight per node, for roughly ~3.2 Tb/s of scale-out bandwidth per node on InfiniBand fabrics. NVLink handles the dense traffic inside the node; ConnectX-7 NDR400 stitches many nodes into a larger training or inference cluster.

Available integrators

The HGX H200 ships as a factory-integrated, air-cooled 8-GPU node from several OEMs — HPE (Cray XD670), Dell (PowerEdge XE9680), Supermicro, and Gigabyte — each with its own chassis, host CPU, and support model. All are new and factory-integrated: the GPU, memory, and NVLink specifications are common across builds, while chassis, host CPU, GPU:NIC ratio, warranty, and lead time differ by integrator. Each build below links to its full specification.

Supermicro

Supermicro 6U 8-GPU

NVIDIA HGX H200 8-GPU Hopper server — 1,128 GB HBM3e per node on the air-cooled Supermicro 6U SXM platform, offered in networking-optimized (1:1 GPU:NIC) and compute-optimized (1:2) builds.

Air-cooled
View specs

Dell

Dell PowerEdge XE9680

NVIDIA HGX H200 8-GPU Hopper server — 1,128 GB HBM3e per node on the air-cooled Dell PowerEdge XE9680 with 4th-gen NVLink and ConnectX-7 NDR400 InfiniBand.

Air-cooled
View specs

Gigabyte

Gigabyte 8-GPU

NVIDIA HGX H200 8-GPU Hopper server — 1,128 GB HBM3e per node on the air-cooled Gigabyte platform with dual AMD EPYC Genoa CPUs and ConnectX-7 NDR400 InfiniBand.

Air-cooled
View specs

HPE

HPE Cray XD670

NVIDIA HGX H200 8-GPU Hopper server — 1,128 GB HBM3e per node on the HPE Cray XD670 (5U) with 4th-gen NVLink and ConnectX-7 NDR400 InfiniBand.

Air-cooled
View specs

Procurement

Every HGX H200 build is quoted per configuration on request — the figure depends on integrator, host CPU, GPU:NIC ratio, networking, and quantity. Lead times run from ready-to-ship to a few weeks by integrator, confirmed at quote. Tell us the build you are standing up and we will return pricing, availability, and a facility-fit review. Browse the GPU catalog to compare integrations, see how the H200 compares to the HGX H100, or step up to the Blackwell HGX B200.

Frequently asked questions

How much memory (VRAM) does an HGX H200 have?

Each H200 GPU in the HGX H200 carries 141 GB of HBM3e, so a full 8-GPU node holds 1,128 GB of HBM3e — pooled and addressable across the GPUs over fourth-generation NVLink, with roughly ~4.8 TB/s of memory bandwidth per GPU.

How many GPUs are in an HGX H200 node?

An HGX H200 node has 8 NVIDIA H200 (Hopper) SXM GPUs on a single fourth-generation NVLink baseboard, connected all-to-all at 900 GB/s per GPU.

What is the H200’s memory bandwidth?

Each H200 GPU delivers roughly ~4.8 TB/s of memory bandwidth from its 141 GB of HBM3e — the upgrade from the H100’s HBM3 that makes the H200 the stronger choice for memory-bound inference at the same Hopper compute.

How much does an HGX H200 server cost?

The HGX H200 is quoted per configuration on request — the figure depends on integrator, host CPU, GPU:NIC ratio, and networking. Tell us the build you need and we will return pricing and availability.

Related

Share this page

Last updated