
The question lands in my inbox almost every week, phrased a dozen different ways: which GPU should we buy? Usually it arrives attached to a spec sheet someone screenshotted, a number they saw on a keynote slide, and a quiet assumption that the newest part is the right part. It rarely is. The right accelerator is the cheapest one that clears your actual workload with headroom — and figuring out which one that is has almost nothing to do with the top-line FLOPS number everyone quotes.
We source this hardware, both new off the OEM line and used off expired contracts, so I watch the whole ladder move at once: Hopper repricing as Blackwell ships, Blackwell Ultra landing at the frontier, rack-scale systems selling out before they're built. This is the field note I wish I could just hand people at the start of that conversation. It's a map of the ladder and a way to find your rung. Public data only, no invented numbers, and no prices — pricing is per-configuration and confirmed at quote, for reasons I'll get to at the end.
The ladder, rung by rung
Think of NVIDIA's current data-center lineup as a ladder with two families on it — Hopper below, Blackwell above — and a rack-scale system sitting at the very top. You don't climb it for its own sake. You climb until you hit the rung that fits your model, then you stop.
Hopper: H100 and H200
H100 is the workhorse that built the current era, and in 2026 it's the value rung. Same Hopper compute it always had, now trading well off its scarcity-era peak as buyers rotate to Blackwell. For steady inference, fine-tuning, and production load that doesn't need the frontier, a used or refurbished H100 still pencils out — the card didn't get worse at its job, a newer one just showed up. We wrote about that repricing in the H100 cliff that isn't; the short version is that cheap H100 is a vintage signal, not a demand signal.
H200 is the more interesting Hopper rung, and the one people skip past too fast. It's the same Hopper architecture as the H100, but NVIDIA swapped in HBM3e: 141 GB of memory per GPU instead of 80 GB, at higher bandwidth. That's it — same compute, far more memory, moving faster. It sounds boring until you're serving a large model with a long context window and discover your bottleneck was never compute. It was how fast weights and KV cache stream out of VRAM. For memory-bound inference, an H200 node routinely serves bigger models, longer contexts, and more concurrent requests than the H100 it replaces. If you're deciding between the two, the whole argument is laid out in HGX H200 vs H100, and the full node spec sheet is in the HGX H200 datasheet.
Blackwell: B200
B200 is the current mainstream training rung. Blackwell is a genuine generational step over Hopper — more memory per GPU, much higher NVLink bandwidth inside the node, and the low-precision formats that modern training and high-throughput inference lean on. If you're training or fine-tuning at real scale and you want current-generation hardware you can rack in a conventional data center, the HGX B200 8-GPU node is the default answer. The honest generational comparison — what actually changed and what didn't — is in Blackwell vs Hopper, and the case for stepping up from Hopper's memory king to Blackwell is in HGX B200 vs H200.
Blackwell Ultra: B300 and the GB300 NVL72
Then there's the top of the ladder. B300 (Blackwell Ultra) is the frontier node — more memory and more of everything than B200, aimed at the largest training runs and the most demanding inference. The step up from B200 is spelled out in HGX B300 vs B200.
Above even that is the GB300 NVL72 — not a node but a rack: dozens of Blackwell Ultra GPUs lashed to Grace CPUs over a single NVLink domain so the whole rack behaves like one enormous accelerator. It's a different kind of purchase from an 8-GPU node, with different facility, networking, and integration realities, and it's genuinely worth understanding before you assume you need one. The generational comparison is GB300 NVL72 vs GB200 NVL72, and the decision that actually matters for most buyers — rack-scale system versus a cluster of HGX nodes — is in GB300 NVL72 rack vs HGX nodes.
How to actually choose
Here's the part the spec sheets won't do for you. Walk these in order and stop at the first honest answer.
-
What's the workload — training or inference? If you're training frontier-scale models, you belong on Blackwell (B200) or Blackwell Ultra (B300 / GB300 NVL72), full stop; there's no used shortcut to the frontier. If you're running inference, fine-tuning, or steady production load, you have real choices lower on the ladder and should take them.
-
Is inference memory-bound? This is the question that saves the most money. If your bottleneck is model size, context length, or KV-cache pressure rather than raw compute, the answer is more and faster memory, not more FLOPS. That's the entire argument for H200 over H100, and often the argument against jumping to Blackwell before you need to. Measure where your inference actually stalls before you buy up.
-
What's your budget and availability window? The frontier is sold out. HBM supply and advanced packaging capacity are the binding upstream constraint, and the newest rack-scale systems are spoken for well ahead. That queue has a mechanic behind it, and it explains why a quoted lead time is what it is — see what GPU allocation actually means. If you need capacity this quarter, Hopper — including a well-priced used H100 or a new H200 node — may be the only rung that's actually in the room, regardless of what you'd prefer on paper.
-
What can your facility power and cool? This is where more buildouts die than on any spec. The Hopper HGX H200 nodes are air-cooled and drop into conventional racks. Rack-scale Blackwell Ultra pushes power density into liquid-cooling territory and brings real facility requirements with it. Don't buy a rung your building can't feed. If your read on the whole lineup is confident and your facility is the open question, sort that before you sort the silicon.
If you want the full spec grids for any of these while you're deciding, the /learn comparison pieces carry the memory, NVLink, and scale-out numbers side by side, and the whole catalog is at GPUs & AI compute for sale. For a straight read on new-versus-used pricing structure — why it moves and how to not overpay — see NVIDIA GPU server pricing. If the question is where these actually get bought rather than which one, how data-center GPUs are sold maps the channel, and new surplus vs refurbished vs used decodes the condition words on any listing.
Why there's no price on this page
Because the honest answer to "how much?" is it depends, and anyone who quotes you a single number before knowing your configuration is selling, not advising. Integrator, host CPU, GPU:NIC ratio, cooling, networking, quantity, and lead time all move the figure, and the market itself reprices faster than any blog post could keep up. We publish public spec data, not our own quotes. When you know your rung, we return pricing, availability, and a facility-fit review against your real build.
So: pick your workload, find where it stalls, check what your building can cool, and buy the cheapest rung that clears it. If you want a second set of eyes on that call for your actual project — browse the catalog or request a quote and tell us what you're standing up.
Pantheon Research is our series on the infrastructure behind AI: power, turbines, cooling, compute, and the procurement reality that decides who actually ships. Field notes from the deal flow, not the keynote.
Milo
Expert in manufacturing technology and industrial solutions, sharing insights on the latest trends and best practices.

