AMD Instinct MI325X Specs & Datasheet (8-GPU Platform)
TL;DR
The AMD Instinct MI325X is an AMD CDNA 3 data-center accelerator: 256 GB of HBM3e per GPU at 6 TB/s of bandwidth, deployed as an 8-GPU OAM platform holding 2,048 GB (~2 TB) of pooled memory. That is the largest single-GPU memory in its class — 33% more than the 192 GB MI300X on the same compute core. It ships factory-integrated on an 8-GPU platform; configuration, power, and pricing are confirmed at quote.
On this page
What the MI325X is
The AMD Instinct MI325X is an AMD data-center accelerator built on the CDNA 3 architecture. It is the memory-forward member of the Instinct line: 256 GB of HBM3e per GPU at roughly 6 TB/s of bandwidth — the largest single-GPU memory capacity in its class.
Like NVIDIA’s HGX platforms, the MI325X is sold as a platform, not a card. Eight accelerators sit on an OAM baseboard connected by AMD Infinity Fabric, giving 2,048 GB (~2 TB) of pooled HBM3e in one chassis. The MI325X shares the same CDNA 3 compute core as the MI300X; what it adds is more memory (256 GB vs 192 GB) and more bandwidth (6 TB/s vs 5.3 TB/s). The CDNA 4 MI355X is its successor.
Full spec sheet
The MI325X per-GPU and 8-GPU platform specification. The accelerator, memory, and fabric figures are fixed characteristics of the platform; chassis, host CPU, cooling, site power draw, and final configuration vary by build and are confirmed at quote.
| Spec | AMD Instinct MI325X |
|---|---|
| Architecture | AMD Instinct MI325X · CDNA 3 |
| Form factor | OAM · 8-GPU platform |
| HBM3e per GPU | 256 GB |
| HBM3e per 8-GPU platform | 2,048 GB (~2 TB) |
| Memory bandwidth | 6.0 TB/s per GPU |
| Dense FP16 / BF16 | ~1,307.4 TFLOPS (~1.31 PFLOPS) per GPU |
| Dense FP8 | ~2,614.9 TFLOPS (~2.61 PFLOPS) per GPU |
| Dense FP16 (8-GPU platform) | ~10.5 PFLOPS |
| Dense FP8 (8-GPU platform) | ~20.9 PFLOPS |
| GPU interconnect | AMD Infinity Fabric · 896 GB/s |
| GPU TDP | 1,000 W per GPU |
| Integrators | HPE |
Memory & bandwidth
Memory is the MI325X’s defining characteristic. Each accelerator carries 256 GB of HBM3e at about 6 TB/s — so an 8-GPU platform pools 2,048 GB (~2 TB) of high-bandwidth memory addressable over Infinity Fabric.
That capacity is the reason to look at the part. It is the largest single-GPU memory in its class, and 33% more per GPU than the 192 GB MI300X on the identical CDNA 3 compute core — with bandwidth stepping up from 5.3 TB/s to 6.0 TB/s. For memory-bound workloads — large models held resident, long-context inference with a big KV cache — capacity per GPU is frequently the binding constraint, and this is where the MI325X is most competitive: fewer GPUs to hold a given model, and more headroom per GPU.
Compute
On compute, each MI325X delivers roughly ~1,307.4 TFLOPS (~1.31 PFLOPS) of dense FP16/BF16 and ~2,614.9 TFLOPS (~2.61 PFLOPS) of dense FP8. Across an 8-GPU platform that is about ~10.5 PFLOPS dense FP16 and ~20.9 PFLOPS dense FP8.
Those are the same CDNA 3 throughput figures as the MI300X — the MI325X’s generational gain is memory and bandwidth, not raw FLOPS. All numbers here are dense; AMD’s datasheet headline figures double these with 2:4 sparsity.
Infinity Fabric & scale-out
Inside the platform, the eight accelerators communicate over AMD Infinity Fabric at 896 GB/s. This is the on-baseboard fabric that lets the eight GPUs share their 2,048 GB of pooled HBM3e as one platform, rather than eight isolated cards.
The practical consequence: the MI325X is deployed as a platform-scale unit — where ~2 TB of pooled memory in one chassis does the work of a memory-bound job without sharding a model across separate hosts. Scale-out networking beyond the platform (InfiniBand or RoCE) is specified per build and confirmed at quote.
Power & cooling
Each MI325X has a 1,000 W TDP. An 8-GPU baseboard therefore draws on the order of 8 kW from the accelerators alone, before host CPUs, system memory, and conversion losses.
That density is what drives the cooling and facility design of a platform. Chassis, cooling method, and exact per-node site power draw depend on the integrator build and are confirmed at quote.
MI325X vs MI300X and MI355X
Where the MI325X sits in the AMD Instinct line. It shares the MI300X’s CDNA 3 compute core and adds memory and bandwidth; the MI355X is the CDNA 4 successor.
| Spec | MI300X | MI325X | MI355X |
|---|---|---|---|
| Architecture | CDNA 3 | CDNA 3 | CDNA 4 |
| HBM per GPU | 192 GB HBM3 | 256 GB HBM3e | 288 GB HBM3E |
| Memory bandwidth | 5.3 TB/s | 6.0 TB/s | 8 TB/s |
| Dense FP16 | ~1.31 PFLOPS | ~1.31 PFLOPS | ~2.5 PFLOPS |
| Dense FP8 | ~2.61 PFLOPS | ~2.61 PFLOPS | ~5 PFLOPS |
| GPU interconnect | Infinity Fabric | Infinity Fabric · 896 GB/s | Infinity Fabric · 896 GB/s |
Available integrators
The MI325X ships as a factory-integrated 8-GPU platform. The accelerator, memory, and fabric specifications are fixed; chassis, host CPU, cooling, warranty, and lead time depend on the build.
HPE
HPE ProLiant Compute XD685
AMD Instinct MI325X 8-GPU server — 2,048 GB HBM3E per node on the HPE ProLiant Compute XD685 (5U), 256 GB per GPU for the largest memory-bound models.
What drives the price of an MI325X platform
An 8-GPU MI325X platform is quoted per configuration on request — there is no meaningful list price, because two platforms with the same eight accelerators can differ substantially in cost. The variables that actually move the number:
- Accelerator count. The 8-GPU baseboard is the standard unit of sale; the accelerators dominate the bill of materials.
- Cooling and chassis. The 1,000 W per-GPU TDP drives the cooling and facility design of the build.
- Host CPU and memory. The host CPU tier and system DDR5 capacity feeding the accelerators.
- Scale-out networking. How much InfiniBand or RoCE fabric per GPU beyond the platform, and the switching to match.
- Quantity and support scope. Single platform versus a multi-rack cluster, and the warranty and on-site service term.
- Lead time. What you need delivered, and when.
For the wider picture of how accelerator servers are actually quoted, see NVIDIA GPU server pricing. Tell us the build you are standing up and we will return pricing, availability, and a facility-fit review. Browse the GPU catalog to compare integrations.
Frequently asked questions
What are the AMD Instinct MI325X specs?
The MI325X is an AMD CDNA 3 accelerator with 256 GB of HBM3e per GPU at 6.0 TB/s of bandwidth, ~1,307.4 TFLOPS (~1.31 PFLOPS) of dense FP16/BF16 and ~2,614.9 TFLOPS (~2.61 PFLOPS) of dense FP8, and a 1,000 W TDP. It ships as an 8-GPU OAM platform connected by AMD Infinity Fabric at 896 GB/s, pooling 2,048 GB (~2 TB) of HBM3e.
How much memory does the MI325X have?
Each MI325X carries 256 GB of HBM3e at roughly 6.0 TB/s of bandwidth — the largest single-GPU memory in its class. An 8-GPU platform therefore pools 2,048 GB (~2 TB) of high-bandwidth memory over AMD Infinity Fabric.
How does the MI325X compare to the MI300X?
The MI325X and MI300X share the same CDNA 3 compute core, so their dense FP16 (~1.31 PFLOPS) and FP8 (~2.61 PFLOPS) per-GPU throughput match. The MI325X’s gains are memory and bandwidth: 256 GB of HBM3e at 6.0 TB/s versus the MI300X’s 192 GB at 5.3 TB/s — 33% more memory per GPU on the identical compute engine.
What is the MI325X’s TDP and power consumption?
Each MI325X has a 1,000 W TDP, so an 8-GPU baseboard draws on the order of 8 kW from the accelerators alone before host CPUs, system memory, and conversion losses. That density drives the cooling and facility design; exact per-node draw is confirmed at quote.
How much does an 8-GPU MI325X platform cost?
It is quoted per configuration on request. The figure depends on accelerator count, cooling and chassis, host CPU tier, system memory, how much scale-out fabric per GPU, quantity, support scope, and lead time. Tell us the build you need and we will return pricing and availability.
Related
Last updated