PantheonGet Early Access

NVIDIA HGX H200 vs H100

TL;DR

The H200 and H100 are the same Hopper GPU — the H200 is a memory and bandwidth upgrade of the H100, not a new architecture. It carries 141 GB of HBM3e per GPU at ~4.8 TB/s versus 80 GB of HBM3 at ~3.35 TB/s on the H100; the tensor compute is identical. That makes the H200 the pick for memory-bound inference and large-context serving, while the H100 is often the better value where 80 GB per GPU is enough. Both are new, 8-GPU Hopper nodes.

On this page

The short answer

These are two configurations of the same Hopper GPU. The H200 is a memory upgrade of the H100 — not a new architecture. The compute is identical; what changes is on-package memory and bandwidth.

The H100 carries 80 GB of HBM3 per GPU at roughly 3.35 TB/s; the H200 carries 141 GB of HBM3e at roughly 4.8 TB/s — about 76% more memory and a third more bandwidth on the same Hopper silicon. Because the tensor math is the same, the H200’s advantage shows up on memory-bound work: large-model inference and long-context serving, where capacity and bandwidth are the ceiling. Where 80 GB per GPU is enough, the H100 is often the better value. Both are new, factory-integrated, 8-GPU HGX nodes.

Spec comparison

Per-GPU and per-node figures for the two 8-GPU Hopper platforms. The two share the same Hopper compute and the same fourth-generation NVLink at 900 GB/s per GPU — the differences are memory capacity (141 GB vs 80 GB per GPU) and bandwidth (~4.8 vs ~3.35 TB/s per GPU).

SpecHGX H200HGX H100
ArchitectureHopperHopper
HBM per GPU141 GB HBM3e80 GB HBM3
Per 8-GPU node1,128 GB640 GB
Memory bandwidth~4.8 TB/s per GPU~3.35 TB/s per GPU
Peak FP8 (dense)1,979 TFLOPS (~2 PFLOPS) — same as H1001,979 TFLOPS (~2 PFLOPS) — same as H200
NVLink4th-gen · 900 GB/s4th-gen · 900 GB/s
ComputeSame Hopper GPUSame Hopper GPU

What changed: memory & bandwidth

The H200 keeps the H100’s Hopper compute and fourth-generation NVLink (900 GB/s per GPU) and upgrades the memory subsystem. Per GPU, HBM grows from 80 GB of HBM3 to 141 GB of HBM3e, and bandwidth rises from roughly 3.35 TB/s to roughly 4.8 TB/s. Across an 8-GPU node that is 1,128 GB of HBM3e on the H200 versus 640 GB of HBM3 on the H100.

More capacity means a larger model — and its long-context KV cache — stays resident on fewer GPUs, cutting cross-GPU traffic and simplifying serving. More bandwidth means the same Hopper cores are fed faster. Both are exactly what memory-bound inference is limited by, which is why the H200 pulls ahead there despite having identical tensor throughput.

That tensor throughput is identical on both — about 1,979 TFLOPS (~2 PFLOPS) of dense FP8 per GPU — because they share the Hopper die; the spec table lists it flat by design. (NVIDIA’s "with sparsity" figure is double, but it is the same on each.)

Which to buy

Match the node to whether the workload is memory-bound:

  • Memory-bound inference and large-context serving → the HGX H200. 141 GB of HBM3e per GPU (1,128 GB per node) and ~4.8 TB/s of bandwidth handle very large models and long-context KV caches with headroom the H100 can’t match.
  • Value, or workloads that fit in 80 GB per GPU → the HGX H100. Same Hopper compute; where capacity isn’t the ceiling, it delivers the useful throughput at a lower cost, and it comes in both air-cooled and liquid-cooled builds.

Cross-shopping the newer generation? See HGX B200 vs H200 for the Blackwell step-up. Pricing across every configuration is quoted per configuration on request — browse the GPU catalog and tell us the build you need.

Frequently asked questions

What’s the difference between H200 and H100?

The H200 is a memory and bandwidth upgrade of the same Hopper GPU — not a new architecture. It carries 141 GB of HBM3e per GPU at ~4.8 TB/s versus 80 GB of HBM3 at ~3.35 TB/s on the H100 (1,128 GB vs 640 GB across an 8-GPU node). Tensor compute and fourth-generation NVLink (900 GB/s per GPU) are identical between the two.

Is the H200 faster than the H100?

The tensor compute is the same — both are the Hopper GPU. The H200 is faster on memory-bound work because it has more memory (141 GB vs 80 GB) and higher bandwidth (~4.8 vs ~3.35 TB/s per GPU): large-model inference and long-context serving see the gain. On workloads that already fit in 80 GB per GPU and aren’t bandwidth-limited, the two perform similarly.

Should I still buy H100?

Yes, for many workloads. Where 80 GB of HBM3 per GPU (640 GB per node) is enough and the job isn’t memory-bandwidth-limited, the H100 delivers the same Hopper compute at a lower cost — often the better value. It also comes in both air-cooled and liquid-cooled builds. Step up to the H200 when capacity or bandwidth is your ceiling.

How much does an HGX H200 or H100 server cost?

GPU systems are quoted per configuration on request — the figure depends on model, integrator, cooling, and networking. Tell us the build you need and we will return pricing and availability.

Related

Share this page

Last updated