PantheonGet Early Access

Learn

GPUs & AI compute

How data-center GPUs actually compare — Blackwell vs Hopper, rack-scale vs node, what the specs mean for a buildout, and the network fabric that ties the cluster together.

NVIDIA accelerator hardware

NVIDIA HGX B200 vs H200

The HGX B200 is NVIDIA’s newer Blackwell platform: more memory (180 GB vs 141 GB per GPU), higher bandwidth, faster NVLink, and lower-precision FP4/FP8 tensor throughput built for frontier-scale training and the highest-throughput inference. The HGX H200 is the top of the Hopper generation — still excellent for memory-bound inference and large-model serving, and usually the better value and availability play. Both are new, 8-GPU nodes and available now; the right pick is workload-driven.

Read →
NVIDIA data-center GPU

NVIDIA HGX H200 vs H100

The H200 and H100 are the same Hopper GPU — the H200 is a memory and bandwidth upgrade of the H100, not a new architecture. It carries 141 GB of HBM3e per GPU at ~4.8 TB/s versus 80 GB of HBM3 at ~3.35 TB/s on the H100; the tensor compute is identical. That makes the H200 the pick for memory-bound inference and large-context serving, while the H100 is often the better value where 80 GB per GPU is enough. Both are new, 8-GPU Hopper nodes.

Read →
Server motherboard

NVIDIA Blackwell vs Hopper

Hopper (H100/H200) is NVIDIA’s prior data-center GPU generation; Blackwell (B200, and Blackwell Ultra B300/GB300) is newer — more HBM3e memory, higher bandwidth, fifth-generation NVLink, and FP4/FP6 low-precision on a second-generation Transformer Engine. Blackwell Ultra pushes memory further again and powers the rack-scale GB300 NVL72. Which to buy is workload-driven: Blackwell for frontier-scale training and highest-throughput inference, while the Hopper H200 remains strong for memory-bound inference and value.

Read →
Data-center racks

NVIDIA GB300 NVL72 Reference Architecture & Rack Specs

The NVIDIA GB300 NVL72 is a single liquid-cooled rack that fuses 72 GB300 (Grace Blackwell Ultra) GPUs and 36 Grace CPUs across 18 compute nodes into one fifth-generation NVLink domain — roughly 20.7 TB of HBM3e behaving as one coherent accelerator, built for frontier-scale AI training and inference. It ships factory-integrated from several OEMs (Supermicro, Lenovo, HPE, Pegatron); configuration, power, weight, and pricing are confirmed at quote.

Read →
Data-center cold aisle

NVIDIA GB300 NVL72 vs GB200 NVL72

The GB300 NVL72 is NVIDIA’s current rack-scale flagship on Grace Blackwell Ultra — 72 GPUs at 288 GB of HBM3e each, roughly 20.7 TB of HBM3e per rack. The GB200 NVL72 is the prior Grace Blackwell rack at 192 GB per GPU (≈13.8 TB per rack). Both wire 72 GPUs into one fifth-generation NVLink domain at 1.8 TB/s; the GB300 is a memory step-up, and the market has rotated to it for the largest models. We supply the GB300 NVL72.

Read →
Data-center hall

Air-Cooled vs Liquid-Cooled HGX H100 Servers

Both our HGX H100 builds carry the same eight H100 80 GB SXM GPUs (640 GB HBM3 per node) — the choice is cooling and density. The air-cooled 8U server drops into a standard air-cooled facility with no liquid loop, so it deploys fast with no retrofit. The liquid-cooled 4U server packs the same eight GPUs into half the rack units for higher density but needs a coolant distribution unit (CDU) and a facility liquid loop. The GPU is identical; the decision is your facility and density target.

Read →
Server board detail

NVIDIA HGX B200 Specs & Datasheet (8-GPU Node)

The NVIDIA HGX B200 is an 8-GPU Blackwell node: eight B200 GPUs with 180 GB of HBM3e each — 1,440 GB per node — connected over a fifth-generation NVLink baseboard, with roughly 62 TB/s of aggregate memory bandwidth for large-model training and inference. It ships factory-integrated from several OEMs (Lenovo, Dell, Supermicro, Gigabyte) in air- or liquid-cooled form factors; configuration, power, and pricing are confirmed at quote.

Read →
Hands holding a smartphone showing the NVIDIA logo

NVIDIA HGX H100 Specs & Datasheet (8-GPU Node)

The NVIDIA HGX H100 is an 8-GPU Hopper node: eight H100 SXM GPUs with 80 GB of HBM3 each — 640 GB per node — connected over a fourth-generation NVLink baseboard at 900 GB/s per GPU, with 3.35 TB/s of memory bandwidth per GPU for large-model training and inference. Each GPU is a 700 W SXM5 module delivering about 989.5 TFLOPS of dense FP16 and 1,979 TFLOPS of dense FP8. It ships factory-integrated from several OEMs (Supermicro, Dell, Gigabyte, HPE) in air- or liquid-cooled form factors; configuration, power, and pricing are confirmed at quote.

Read →
Hands holding a smartphone showing the NVIDIA logo

NVIDIA A100 Specs & Datasheet (80GB SXM · HGX & DGX A100)

The NVIDIA A100 SXM 80GB is an Ampere-generation data-center GPU: 80 GB of HBM2e at 2.04 TB/s per GPU, about 312 TFLOPS of dense FP16 Tensor compute, 400 W, connected over third-generation NVLink at 600 GB/s across an 8-GPU domain. It ships in two 8-GPU forms — the HGX A100 baseboard (integrator-built) and the NVIDIA DGX A100 640GB system, which is exactly eight of these GPUs (640 GB HBM2e per node). A100 is a prior-generation part (Hopper H100/H200 and Blackwell are the current silicon); it is still widely deployed and available, with configuration and pricing confirmed at quote.

Read →
An AMD Ryzen processor seated in a motherboard socket

AMD Instinct MI325X Specs & Datasheet (8-GPU Platform)

The AMD Instinct MI325X is an AMD CDNA 3 data-center accelerator: 256 GB of HBM3e per GPU at 6 TB/s of bandwidth, deployed as an 8-GPU OAM platform holding 2,048 GB (~2 TB) of pooled memory. That is the largest single-GPU memory in its class — 33% more than the 192 GB MI300X on the same compute core. It ships factory-integrated on an 8-GPU platform; configuration, power, and pricing are confirmed at quote.

Read →
An NVIDIA NVLink bridge connector

NVIDIA GB300 NVL72 Specs & Datasheet (72-GPU Rack)

The NVIDIA GB300 NVL72 is a liquid-cooled rack that fuses 72 Grace Blackwell Ultra GPUs and 36 Grace CPUs — arranged two GPUs per superchip across 18 compute trays — into a single fifth-generation NVLink domain. Each GPU carries 288 GB of HBM3e, so a full rack holds roughly 20.7 TB of pooled high-bandwidth memory behaving as one coherent accelerator. It ships factory-integrated from Supermicro, Lenovo, HPE, and Pegatron; configuration, site power, and pricing are confirmed at quote.

Read →
A smartphone displaying the NVIDIA logo

NVIDIA HGX B300 Specs & Datasheet (8-GPU Node)

The NVIDIA HGX B300 is an 8-GPU Blackwell Ultra node: eight B300 GPUs with 288 GB of HBM3e each — about 2.3 TB per node — connected over a fifth-generation NVLink baseboard at 1.8 TB/s per GPU, with roughly 8 TB/s of memory bandwidth per GPU for the largest-memory training and inference. It is the Blackwell Ultra successor to the HGX B200 and ships factory-integrated from several OEMs (Supermicro, Dell, HPE, Lenovo) in air- or liquid-cooled builds; configuration, power, and pricing are confirmed at quote.

Read →
GPU secondary market

NVIDIA GPU Server Pricing: What Drives the Cost & How to Quote

NVIDIA GPU systems are quoted per configuration on request — there is no list price because the figure depends on the model and generation, the integrator, cooling, networking, memory and configuration, and quantity. A single 8-GPU node and a full NVL72 rack are different orders of magnitude. This page explains what moves the number and how to get an exact quote.

Read →
Server memory modules

NVIDIA HGX B300 vs B200

The HGX B300 is NVIDIA’s Blackwell Ultra 8-GPU node: 288 GB of HBM3e per GPU (2,304 GB per node) and ConnectX-8 800 Gb/s scale-out. The HGX B200 is the original Blackwell node at 180 GB per GPU (1,440 GB per node) with ConnectX-7 400 Gb/s. Both share the same fifth-generation NVLink at 1.8 TB/s and the same 8 GPUs per node — the B300 is a memory and scale-out step-up. Reach for the B300 when models or headroom demand more memory; the B200 is available now and strong for most training and inference.

Read →
AMD accelerator hardware

NVIDIA HGX B200 vs AMD Instinct MI300X

The AMD Instinct MI300X is a capable data-center accelerator — 192 GB of HBM3 per GPU (slightly more than the B200’s 180 GB) at ~5.3 TB/s. The NVIDIA HGX B200 leads on interconnect (fifth-generation NVLink at 1.8 TB/s per GPU vs Infinity Fabric) and, decisively for most teams, on the maturity of the CUDA software ecosystem versus AMD’s maturing ROCm. Memory capacity is close; the practical choice usually turns on interconnect and software. We source NVIDIA HGX — the B200 and H200 — so that is where we point customers.

Read →
Data-center site

GB300 NVL72 Rack vs HGX 8-GPU Nodes

The choice is form factor and granularity. A GB300 NVL72 is a factory-integrated rack — 72 GPUs wired into one fifth-generation NVLink domain, liquid-cooled — bought as a whole unit; it is the pick for the largest single-domain training, maximum density, and new build-outs. HGX 8-GPU nodes (B200 or B300, air or liquid) deploy incrementally into standard racks and suit adding capacity node-by-node, standard facilities, and mixed workloads. Both are NVIDIA Blackwell; the right one depends on scale, facilities, and how you want to grow.

Read →
High-bandwidth memory

NVIDIA HGX H200 Specs & Datasheet (8-GPU Node)

The NVIDIA HGX H200 is an 8-GPU Hopper node: eight H200 SXM GPUs with 141 GB of HBM3e each — 1,128 GB per node — connected over a fourth-generation NVLink baseboard at 900 GB/s per GPU, with roughly 4.8 TB/s of memory bandwidth per GPU for memory-bound inference and large-model training. It ships factory-integrated from several OEMs (HPE, Dell, Supermicro, Gigabyte) in air-cooled form factors; configuration, power, and pricing are confirmed at quote.

Read →
Power infrastructure detail

NVIDIA GB200 NVL72 Specs & Datasheet (72-GPU Rack)

The NVIDIA GB200 NVL72 is a liquid-cooled rack that fuses 72 Blackwell GPUs and 36 Grace CPUs — 36 GB200 Grace Blackwell superchips across 18 compute trays — into a single fifth-generation NVLink domain. Each GPU carries 192 GB of HBM3e, so a full rack holds roughly 13.8 TB of pooled high-bandwidth memory behaving as one coherent accelerator. It ships factory-integrated from Supermicro, Dell, and HPE; configuration, site power, and pricing are confirmed at quote.

Read →
AMD silicon

AMD Instinct MI355X Specs & Datasheet (8-GPU Server)

The AMD Instinct MI355X is AMD’s CDNA 4 flagship data-center accelerator: 288 GB of HBM3E per GPU at 8 TB/s, deployed as an 8-GPU node holding 2,304 GB of pooled memory. That matches the memory of NVIDIA’s Blackwell Ultra parts and is 60% more per GPU than an HGX B200. It ships factory-integrated on a 4U direct-liquid-cooled Supermicro platform; configuration, power, and pricing are confirmed at quote.

Read →
Close-up of an AMD Radeon graphics card with its glowing logo

AMD Instinct MI300X Specs & Datasheet (8-GPU Server)

The AMD Instinct MI300X is AMD’s CDNA 3 data-center accelerator: 192 GB of HBM3 per GPU at 5.3 TB/s of bandwidth, deployed as an 8-GPU OAM node holding 1,536 GB of pooled memory. At 750 W per accelerator it is the most deployable member of the Instinct line — air-cooled 8-GPU builds ship from Supermicro and Dell without any facility-water requirement. Configuration, power, and pricing are confirmed at quote.

Read →
Close-up of SK hynix memory chips on a module

What is HBM and Why Does It Matter for AI?

HBM is high-bandwidth memory — DRAM stacked vertically and mounted next to the GPU die on the same package, giving far more bandwidth than the GDDR used on graphics cards. It matters because large-model inference is usually memory-bound, not compute-bound: capacity decides whether a model and its KV cache fit at all, and bandwidth decides how fast tokens come out. For most buyers, "how many GPUs do I need" is really "how much HBM do I need, and how fast".

Read →
Inside a computer showing a motherboard with an installed processor, memory, and graphics card

GPU vs Server vs Rack: What You Actually Buy

Data-center AI compute is sold in layers. The accelerator is a module, not a product you rack on its own; the unit of purchase is an integrated 8-GPU node built by an OEM around an NVIDIA HGX baseboard, complete with host CPUs, memory, NICs, power and cooling. Nodes go into racks, racks into a cluster tied together by a scale-out fabric. Rack-scale systems such as the GB300 NVL72 change the layering itself: an entire rack becomes one coherent accelerator domain rather than a set of networked nodes.

Read →
A smartphone displaying the NVIDIA logo

HGX vs DGX: What is the Difference?

HGX and DGX are the same NVIDIA GPU silicon sold two ways. HGX is a reference platform — an 8-GPU baseboard with NVSwitch that NVIDIA supplies to server manufacturers, who build it into their own servers with their own CPUs, NICs, chassis, cooling and support. DGX is NVIDIA’s own complete system built on that same baseboard, with a fixed configuration, NVIDIA’s software stack and NVIDIA support. HGX gives configurability, integrator choice and more paths to supply; DGX gives one vendor, one validated stack, and a turnkey deployment.

Read →
Close-up of a printed circuit board with surface-mount components

How to Size a GPU Cluster

Size a cluster in one direction: workload first, then memory footprint, then GPU count, then node count, then the fabric, and only then power, cooling and floor space. The step most people miss is the last one — nameplate GPU TDP is not facility load. Host platforms, networking, storage and cooling overhead typically push the number at the meter to roughly two-and-a-half to three times the sum of the GPU TDPs. That is why AI buildouts turn into power procurement projects.

Read →
A data center aisle lined with server racks and cabling

InfiniBand vs Ethernet for AI Clusters

InfiniBand and Ethernet are both scale-out fabrics — the network between GPU nodes. They are not alternatives to NVLink, which is the scale-up interconnect inside a node or rack. InfiniBand is lossless by design, RDMA-native and long-established in HPC and large AI training. Ethernet with RoCEv2 — including AI-tuned variants and the Ultra Ethernet effort — has closed much of the gap and brings a broader vendor ecosystem and familiar operations. Both are deployed at serious scale in 2026; the right answer depends on cluster size, workload and the skills of the team running it.

Read →
A data center server with rows of drives and status indicator lights

Buying vs Leasing vs Cloud GPUs

The hinge is sustained utilization. Owned hardware costs the same whether it runs at 20% or 95%, so ownership only pencils when demand is steady and high; cloud bills for what you consume, which is why it is genuinely the right answer for spiky, uncertain or short-lived workloads. Leasing sits in the middle, converting a large purchase into a predictable payment. Lead time, depreciation, residual value, control and data residency then decide the edge cases. Many organizations land on a hybrid: own the steady baseline, burst to cloud.

Read →
A data center server with rows of drives and status indicator lights

NVIDIA SN4600C Specs & Datasheet (MSN4600-CS2RC)

The NVIDIA Spectrum-3 SN4600C is a 2U open Ethernet switch with 64 QSFP28 ports of 100GbE — 6.4 Tb/s of switching capacity and 8.4 Bpps of packet processing — built for spine and super-spine roles in large virtualized data centers and as the 100GbE storage and management fabric around GPU clusters. It ships under several MSN4600 ordering part numbers that differ only in airflow direction and preloaded network OS; MSN4600-CS2RC is the C2P, Cumulus Linux variant.

Read →
Blue-lit blade servers mounted in a data center enclosure

MSN4600-CS2RC vs MSN4600-CS2FC: C2P or P2C Airflow?

MSN4600-CS2RC and MSN4600-CS2FC are the same NVIDIA Spectrum-3 SN4600C switch — 64 QSFP28 ports of 100GbE, Cumulus Linux, dual AC power supplies — differing only in the direction air moves through the chassis. CS2RC is C2P (air enters at the ports); CS2FC is P2C (air enters at the power-supply side). Airflow is fixed at manufacture, so the choice is determined by which side of the switch faces your cold aisle.

Read →
A data center aisle lined with server racks and cabling

Cumulus Linux vs Onyx: NVIDIA Switch Operating Systems

Cumulus Linux and Onyx are the two network operating systems NVIDIA has shipped preloaded on Spectrum Ethernet switches. Cumulus Linux is a Linux-native NOS managed with standard Linux tooling; Onyx (formerly MLNX-OS) is the older Mellanox-lineage switch OS. Current SN4600C ordering part numbers ship Cumulus Linux — marked by a trailing C in the part number — and the Onyx part numbers for that platform have been withdrawn.

Read →

Can you use third-party optics in NVIDIA switches?

Yes, in almost every case. An optical transceiver is built to a Multi-Source Agreement (MSA) standard, so a third-party module that matches the port — same form factor, same speed, same reach, and coded for the platform — moves traffic identically to the OEM-branded part in the same slot. Two things decide whether a given module works: the coding in its EEPROM (what the switch reads when it powers the port up) and matching the optic's reach to the actual fiber run. The "third-party optics void your warranty" line is, in the United States, mostly a myth. What genuinely breaks a link is a reach or fiber-type mismatch — not the brand on the label.

Read →

What is the GPU memory wall?

The "memory wall" is a thirty-year-old observation — compute speed improves faster than memory speed and capacity, so eventually memory, not the processor, sets the ceiling. In AI infrastructure that ceiling has arrived: a modern GPU can do far more math per second than its fixed on-package HBM can hold or feed, so large-model inference is usually memory-bound, not compute-bound. The practical symptom is teams buying more GPUs than their compute needs just to get memory capacity, then watching that compute sit idle. The structural response is to stop pinning every byte to one GPU or server and instead pool memory — a disaggregated pool shared across servers over a fast fabric.

Read →

Why is my KV cache too large for VRAM?

The KV cache is the per-request attention state a transformer keeps so it does not recompute the whole prompt for every new token. It grows linearly with context length and linearly with the number of concurrent requests, and it lives in the same HBM as the model weights. At short context it is a rounding error; at long context and real concurrency it becomes the dominant consumer and is usually what triggers the out-of-memory error. Quantization and offload stretch the budget, but past a point the cache needs somewhere real to land — CPU DRAM, or a disaggregated memory pool the whole cluster can share.

Read →

What is stranded DRAM, and why is half my memory idle?

Every server is bought with enough DRAM for its worst-case moment, so on average a large fraction of that memory is idle — "stranded" on a CPU that is not using it while a neighbor a rack away is memory-starved. Hyperscale measurements put the waste around a quarter to a half of all DRAM. Because memory is one of the biggest and now fastest-rising lines in a server's cost, stranded DRAM is real capital sitting dark. The fix is the same as for the GPU memory wall: stop welding memory to individual servers and pool it, so idle capacity in one place serves demand in another.

Read →