NVIDIA Blackwell vs Hopper
TL;DR
Hopper (H100/H200) is NVIDIA’s prior data-center GPU generation; Blackwell (B200, and Blackwell Ultra B300/GB300) is newer — more HBM3e memory, higher bandwidth, fifth-generation NVLink, and FP4/FP6 low-precision on a second-generation Transformer Engine. Blackwell Ultra pushes memory further again and powers the rack-scale GB300 NVL72. Which to buy is workload-driven: Blackwell for frontier-scale training and highest-throughput inference, while the Hopper H200 remains strong for memory-bound inference and value.
On this page
The short answer
These are two GPU architectures, a generation apart. Hopper is the prior generation — the H100 (80 GB) and the memory-upgraded H200 (141 GB). Blackwell is newer: the B200 (180 GB), and the Blackwell Ultra refresh (B300 / GB300, 288 GB).
Blackwell wins on raw capability — more HBM3e per GPU, higher memory bandwidth, a faster fifth-generation NVLink fabric, and added FP4/FP6 low-precision on a second-generation Transformer Engine — and only Blackwell reaches rack-scale as a single NVLink domain (GB200 / GB300 NVL72). Hopper is not obsolete: the H200’s 141 GB of HBM3e still makes it a strong pick for memory-bound inference, and it typically wins on value and availability. The right generation is workload-driven, not simply "newest."
Generational spec comparison
How the two generations line up, with Blackwell split into the original Blackwell (B200) and the Blackwell Ultra refresh (B300 / GB300). The headline moves are memory (80/141 GB → 180 GB → 288 GB per GPU), a doubling of per-GPU NVLink bandwidth (900 GB/s → 1.8 TB/s), the added FP4/FP6 formats, and Blackwell’s reach into rack-scale NVL72.
| Hopper | Blackwell | Blackwell Ultra | |
|---|---|---|---|
| Example GPUs | H100 / H200 | B200 | B300 / GB300 |
| HBM per GPU | 80 GB HBM3 (H100) · 141 GB HBM3e (H200) | 180 GB HBM3e | 288 GB HBM3e |
| NVLink | 4th-gen · 900 GB/s | 5th-gen · 1.8 TB/s | 5th-gen · 1.8 TB/s |
| Low-precision | FP8 / FP16 | + FP4 / FP6 | + FP4 / FP6 |
| Peak FP8 (dense) | ~2 PFLOPS (1,979 TFLOPS) · no FP4 | ~4.5 PFLOPS · ~9 PFLOPS FP4 | ~4.5 PFLOPS · ~13.5 PFLOPS FP4 |
| Rack-scale | — | GB200 NVL72 | GB300 NVL72 |
What changed: memory
The clearest generational step is memory capacity and bandwidth. Hopper started at 80 GB of HBM3 per GPU on the H100 and grew to 141 GB of HBM3e on the H200. Blackwell’s B200 carries 180 GB of HBM3e — about 28% more than the H200 — and the Blackwell Ultra B300 / GB300 pushes to 288 GB per GPU.
More on-package memory means a larger model (and its long-context KV cache) stays resident on fewer GPUs, which cuts cross-GPU traffic and simplifies serving. Blackwell also raises memory bandwidth, so the GPU is fed faster — the difference that matters most for memory-bound inference. Exact per-GPU bandwidth figures depend on the SKU and are confirmed at quote.
NVLink & fabric
Hopper uses fourth-generation NVLink at 900 GB/s per GPU. Blackwell moves to fifth-generation NVLink at 1.8 TB/s per GPU — double the per-GPU bandwidth — across both the B200 and the Blackwell Ultra B300 / GB300.
The bigger shift is scope. On Hopper, NVLink connects the GPUs inside an 8-GPU HGX node. Blackwell extends the NVLink domain to rack scale: the GB200 and GB300 NVL72 wire 72 GPUs into a single all-to-all NVLink fabric (roughly 130 TB/s in-rack on the GB300 NVL72), so an entire rack behaves like one large accelerator instead of many networked nodes.
Precision & the Transformer Engine
Hopper introduced the Transformer Engine with FP8 (alongside FP16) to accelerate transformer training and inference. Blackwell adds a second-generation Transformer Engine that introduces even lower-precision FP4 and FP6 formats.
Lower precision means more tokens per second per GPU wherever the model and toolchain can use it — a large lever for high-throughput inference. Hopper remains fully capable at FP8/FP16; Blackwell simply extends the range downward for the workloads that benefit.
In throughput terms, an HGX B200 GPU delivers roughly ~4.5 PFLOPS of dense FP8 and ~9 PFLOPS of dense FP4 (Blackwell Ultra HGX B300: ~4.5 / ~13.5 PFLOPS), versus about ~2 PFLOPS of dense FP8 (1,979 TFLOPS) on Hopper — with FP4, absent on Hopper, the lever for high-throughput inference. The rack-scale GB200/GB300 superchips run a higher clock bin (~5 / ~10 and ~5 / ~15 PFLOPS dense per GPU). All figures are dense; with 2:4 sparsity NVIDIA quotes double.
Blackwell vs Blackwell Ultra
Within Blackwell there are two tiers. The original Blackwell (B200, and the GB200 Grace-Blackwell superchip) carries 180 GB of HBM3e per GPU. The Blackwell Ultra refresh (B300, and the GB300 Grace Blackwell Ultra superchip) steps memory up to 288 GB per GPU while keeping the same fifth-generation NVLink at 1.8 TB/s and the same FP4/FP6 precision.
In practice: both are Blackwell, but Blackwell Ultra is the higher-memory option, and the rack-scale flagship shifts from GB200 NVL72 to GB300 NVL72 at 288 GB/GPU (≈20.7 TB of HBM3e per rack).
Which to buy
Match the generation to the workload rather than defaulting to the newest silicon:
- Frontier-scale training or the highest-throughput inference → Blackwell. Reach for the HGX B200, the Blackwell Ultra HGX B300, or — at rack scale — the GB300 NVL72.
- Memory-bound inference, value, or a faster facilities fit → the Hopper HGX H200, which delivers most of the useful serving throughput at lower cost and drops into air-cooled facilities.
For a head-to-head on the two most-cross-shopped 8-GPU nodes, see NVIDIA HGX B200 vs H200. Pricing across every configuration is quoted per configuration on request — browse the GPU catalog and tell us the build you need.
Frequently asked questions
Is Blackwell better than Hopper?
Blackwell is the newer, more capable generation — more HBM3e memory per GPU, higher bandwidth, fifth-generation NVLink at 1.8 TB/s (versus 900 GB/s on Hopper), added FP4/FP6 precision, and rack-scale NVLink domains (GB200/GB300 NVL72). Whether it is "better" for you depends on the workload: Blackwell shines for frontier-scale training and highest-throughput inference, while the Hopper H200 is still excellent — and often the better value — for memory-bound inference.
What is the difference between Blackwell and Blackwell Ultra?
Both are the Blackwell architecture. The original Blackwell (B200, GB200) carries 180 GB of HBM3e per GPU; the Blackwell Ultra refresh (B300, GB300) steps that up to 288 GB per GPU while keeping the same fifth-generation NVLink at 1.8 TB/s and FP4/FP6 precision. The rack-scale flagship moves from GB200 NVL72 to GB300 NVL72 accordingly.
Should I still buy Hopper (H200)?
Yes, for many workloads. The H200’s 141 GB of HBM3e per GPU handles very large models and long-context KV caches well, making it one of the strongest platforms for memory-bound inference and large-model serving. It typically wins on value and availability, and air-cooled H200 nodes drop into standard facilities without a liquid-cooling retrofit.
How much do Blackwell vs Hopper servers cost?
GPU systems are quoted per configuration on request — the figure depends on model, integrator, cooling, and networking. Tell us the build you need and we will return pricing and availability.
Related
Last updated