AMD Instinct MI325X vs NVIDIA GB10 Comparison

AMD
RADEON

AMD Instinct MI325X

CORE STATE Aqua Vanjaram
VRAM 256 GB
CLOCK SPEED 2100 MHz
TDP 1000 W
BUS WIDTH 8192 bit
ARCHITECTURE CDNA 3.0
nm
PROCESS 5 nm
LAUNCH DATE 2024
VS
NVIDIA
GEFORCE

GB10

CORE STATE GB20B
VRAM 128 GB
CLOCK SPEED 2418 MHz
TDP 140 W
BUS WIDTH 256 bit
ARCHITECTURE Blackwell 2.0
nm
PROCESS 5 nm
LAUNCH DATE 2025

PERFORMANCE BENCHMARKS

geekbench_opencl
N/A
120,137
geekbench_vulkan
N/A
114,648

Analysis: AMD Instinct MI325X vs NVIDIA GB10

Head-to-Head Benchmarks

The recorded data shows no direct head-to-head benchmark runs between the AMD Instinct MI325X and the NVIDIA GB10. The AMD Instinct MI325X has no benchmark entries in the database, with an average benchmark score of zero and a percentile ranking of 50 among all GPUs. The NVIDIA GB10, by contrast, has two recorded benchmark results: a Geekbench OpenCL score of 120,137 and a Geekbench Vulkan score of 114,648, producing an average benchmark score of 117,393 and a percentile ranking of 95.

The NVIDIA GB10's average score places it slightly ahead of the NVIDIA RTX 4000 SFF Ada Generation, which scores 117,088, a delta of 0.3 percent. It also leads the AMD Radeon PRO W7700, which scores 118,976, by a margin of 1.3 percent in the opposite direction, meaning the GB10 trails that card by that percentage. Against the NVIDIA Tesla V100 SXM2 16 GB, which scores 114,395, the GB10 is ahead by 2.6 percent. The NVIDIA RTX A5500 Mobile scores 113,944, and the GB10 leads it by 3 percent. These comparisons indicate that the GB10, despite its low power envelope, sits in a competitive performance tier among workstation and server-class accelerators.

The AMD Instinct MI325X cannot be placed into this ranking through measured scores because the database contains no benchmark results for it. Its percentile of 50 is the default midpoint for an unmeasured product, not an indication of actual performance. The absence of scores means the only quantitative comparison available is through the specification differences, which show a substantial gap in raw compute resources. Any verdict on relative performance must therefore rely on the architectural data rather than direct measurements.

Architecture Differences

The AMD Instinct MI325X uses the Aqua Vanjaram chip with CDNA 3.0 architecture, built on a 5 nm process at TSMC. The NVIDIA GB10 uses the GB20B chip with Blackwell 2.0 architecture, also on a 5 nm process at TSMC. Both share the same manufacturing node and foundry, but the silicon implementations diverge sharply.

The MI325X contains 153,000 million transistors on a die size of 1,017 mm², yielding a transistor density of 150.4 million per mm². The GB10's transistor count is listed as unknown, but its die size is 382 mm². The MI325X die is more than 2.6 times larger, and the transistor count, where known, is far higher. The MI325X uses a base clock of 1,000 MHz and a boost clock of 2,100 MHz, while the GB10 runs at a higher base of 1,665 MHz and a boost of 2,418 MHz. The GB10's higher clocks reflect its smaller, lower-power design.

Memory architecture differs fundamentally. The MI325X uses 256 GB of HBM3e on an 8,192-bit bus, delivering 6.14 TB/s of bandwidth. The GB10 uses 128 GB of LPDDR5X on a 256-bit bus, delivering 273.2 GB/s. The MI325X has 32 times the bus width and more than 22 times the memory bandwidth, a decisive advantage for memory-bound workloads. The GB10's memory clock is 1,067 MHz with 8.5 Gbps effective, while the MI325X memory runs at 1,500 MHz with 6 Gbps effective.

Compute resources also differ widely. The MI325X has 19,456 shading units and 1,216 texture mapping units, with zero ROPs. The GB10 has 6,144 shading units, 384 TMUs, and 48 ROPs. The MI325X reports a pixel rate of 0 MPixel/s due to its lack of ROPs, while the GB10 delivers 116.1 GPixel/s. The texture rate for the MI325X is 2,553.6 GTexel/s, compared to 928.5 GTexel/s for the GB10. FP32 throughput is 81.72 TFLOPS for the MI325X and 29.71 TFLOPS for the GB10, a ratio of approximately 2.75 to 1. Both list FP16 at 1:1 with their FP32 figures. The GB10 includes 48 RT cores and 384 tensor cores, while the MI325X lists no RT cores or tensor cores in the data.

Power requirements are starkly different. The MI325X has a TDP of 1,000 W with a suggested PSU of 1,400 W, while the GB10 has a TDP of 140 W with a suggested PSU of 300 W. The GB10 is a 150 mm by 51 mm by 150 mm IGP module, while the MI325X is an OAM module with no listed dimensions. The GB10 has a single HDMI output, while the MI325X has no display outputs. Both use PCIe 5.0 x16 and neither uses external power connectors.

FAQ

Q: What is the memory capacity difference between the two accelerators?

A: The AMD Instinct MI325X has 256 GB of HBM3e memory, while the NVIDIA GB10 has 128 GB of LPDDR5X memory. The MI325X offers double the capacity.

Q: How do the memory bandwidth figures compare?

A: The MI325X delivers 6.14 TB/s over an 8,192-bit bus. The GB10 delivers 273.2 GB/s over a 256-bit bus. The MI325X bandwidth is more than 22 times higher.

Q: Which chip has a higher boost clock?

A: The NVIDIA GB10 has a boost clock of 2,418 MHz, compared to the AMD Instinct MI325X boost clock of 2,100 MHz. The GB10 also has a higher base clock at 1,665 MHz versus 1,000 MHz.

Q: What is the TDP of each product?

A: The AMD Instinct MI325X has a TDP of 1,000 W with a suggested PSU of 1,400 W. The NVIDIA GB10 has a TDP of 140 W with a suggested PSU of 300 W.

Q: Does the NVIDIA GB10 support display output?

A: Yes, the GB10 has one HDMI output. The AMD Instinct MI325X has no display outputs.

Q: What process node and foundry do both chips use?

A: Both the AMD Instinct MI325X and the NVIDIA GB10 are built on a 5 nm process at TSMC.

The Verdict

The database records no benchmark scores for the AMD Instinct MI325X, so its measured performance cannot be stated. The NVIDIA GB10 has an average benchmark score of 117,393 and sits at the 95th percentile of all GPUs, placing it above the NVIDIA RTX 4000 SFF Ada Generation, the AMD Radeon PRO W7700, the NVIDIA Tesla V100 SXM2 16 GB, and the NVIDIA RTX A5500 Mobile in the recorded delta comparisons. For use cases where verified performance data is required, the GB10 is the only option with recorded results.

On specifications, the MI325X dominates in raw compute and memory resources. It has 19,456 shading units versus 6,144, 1,216 TMUs versus 384, 81.72 TFLOPS FP32 versus 29.71 TFLOPS, and 256 GB of HBM3e with 6.14 TB/s bandwidth versus 128 GB of LPDDR5X with 273.2 GB/s. The MI325X is designed for maximum throughput in large-scale compute workloads, as its 1,000 W TDP and OAM form factor indicate. The GB10, with a 140 W TDP and IGP form factor, is a compact, low-power accelerator with a display output and a 3,999 USD launch MSRP.

The choice depends on the workload environment. For data-center scale compute with massive memory bandwidth requirements, the MI325X provides the necessary resources on paper. For a smaller, lower-power system where a single HDMI output and a 95th percentile benchmark score are priorities, the GB10 is the measured performer. The MI325X's lack of recorded benchmarks means its actual performance in the database is unverified, while the GB10's scores are concrete.

Specification Differences

| Field | AMD Instinct MI325X | NVIDIA GB10 |

|---|---|---|

| Chip | Aqua Vanjaram | GB20B |

| Architecture | CDNA 3.0 | Blackwell 2.0 |

| Process Node | 5 nm | 5 nm |

| Foundry | TSMC | TSMC |

| Transistors | 153,000 million | unknown |

| Die Size | 1017 mm² | 382 mm² |

| Base Clock | 1000 MHz | 1665 MHz |

| Boost Clock | 2100 MHz | 2418 MHz |

| Memory Size | 256 GB | 128 GB |

| Memory Type | HBM3e | LPDDR5X |

| Memory Bus Width | 8192 bit | 256 bit |

| Memory Bandwidth | 6.14 TB/s | 273.2 GB/s |

| Shading Units | 19456 | 6144 |

| TMUs | 1216 | 384 |

| ROPs | 0 | 48 |

| RT Cores | null | 48 |

| Tensor Cores | null | 384 |

| Pixel Rate | 0 MPixel/s | 116.1 GPixel/s |

| Texture Rate | 2,553.6 GTexel/s | 928.5 GTexel/s |

| FP32 | 81.72 TFLOPS | 29.71 TFLOPS |

| FP16 | 81.72 TFLOPS (1:1) | 29.71 TFLOPS (1:1) |

| TDP | 1000 W | 140 W |

| Slot Width | OAM Module | IGP |

| Suggested PSU | 1400 W | 300 W |

| Display Outputs | No outputs | 1x HDMI |

| Dimensions | null | 150 mm 5.9 inches, 51 mm 2 inches, 150 mm 5.9 inches |

| Production Status | null | Active |

| Release Date | 2024-10-09 | 2025-10-14 |

| Predecessor | Radeon Instinct | Server Hopper |

| Successor | null | Server Rubin |

Where Each One Wins

The AMD Instinct MI325X wins on every raw compute and memory metric in the specification comparison. It has more than three times the shading units, more than three times the TMUs, more than 2.7 times the FP32 throughput, double the memory capacity, and more than 22 times the memory bandwidth. Its 1,017 mm² die and 153,000 million transistors represent a much larger silicon investment aimed at sustained high-throughput compute. It also has no display outputs, confirming a compute-only design. For workloads such as large model training or dense matrix operations where memory bandwidth is the limiting factor, the MI325X is the stronger choice on paper.

The NVIDIA GB10 wins on measured benchmark performance, power efficiency, and physical integration. It has a recorded average benchmark score of 117,393 and a 95th percentile ranking, while the MI325X has no recorded scores. The GB10 has a 140 W TDP versus 1,000 W, a suggested PSU of 300 W versus 1,400 W, and a compact IGP form factor with dimensions of 150 mm by 51 mm by 150 mm. It includes an HDMI output, 48 RT cores, and 384 tensor cores, features that the MI325X lacks in the data. The GB10 also has a faster boost clock at 2,418 MHz versus 2,100 MHz and a production status of Active. For systems requiring a low-power, compact accelerator with verified performance and display capability, the GB10 is the clear selection.

DETAILED SPECIFICATIONS

SPECIFICATION
Instinct MI325X
GB10
Core Specs
Shading Units
19,456
6,144 -68.4%
Shaders
19,456
6,144 -68.4%
TMUs
1,216
384 -68.4%
ROPs
0
48 +∞%
Compute Units
304
SM Count
48
Clocks
Base Clock
1000 MHz
1665 MHz
Boost Clock
2100 MHz
2418 MHz
Memory Clock
1500 MHz 6 Gbps effective
1067 MHz 8.5 Gbps effective
Memory
Memory Size
256 GB
128 GB
VRAM (MB)
262,144
131,072 -50.0%
Memory Type
HBM3e
LPDDR5X
Memory Bus
8192 bit
256 bit
Bandwidth
6.14 TB/s
273.2 GB/s
Cache
L1 Cache
16 KB (per CU)
128 KB (per SM)
L2 Cache
16 MB
50 MB
L3 Cache
256 MB
Performance
Pixel Rate
0 MPixel/s
116.1 GPixel/s
Texture Rate
2,553.6 GTexel/s
928.5 GTexel/s
FP32 (TFLOPS)
81.72 TFLOPS
29.71 TFLOPS
FP64 (TFLOPS)
40.86 TFLOPS (1:2)
464.3 GFLOPS (1:64)
FP16 (TFLOPS)
81.72 TFLOPS (1:1)
29.71 TFLOPS (1:1)
AI/RT
RT Cores
48
Tensor Cores
384
Matrix Cores
1,216
Power
TDP
1000 W
140 W
TDP (W)
1,000
140 -86.0%
Suggested PSU
1400 W
300 W
Power Connectors
None
None
Architecture
Architecture
CDNA 3.0
Blackwell 2.0
GPU Name
Aqua Vanjaram
GB20B
Generation
Instinct (MIx)
Server Blackwell (Bxx)
Process Size
5 nm
5 nm
Transistors
153,000 million
unknown
Die Size
1017 mm²
382 mm²
Foundry
TSMC
TSMC
Density
150.4M / mm²
AMD MCM
MCM
2
API Support
OpenCL
3.0
3.0
CUDA
12.1
Physical
Slot Width
OAM Module
IGP
Length
150 mm 5.9 inches
Height
51 mm 2 inches
Outputs
No outputs
1x HDMI
Bus Interface
PCIe 5.0 x16
PCIe 5.0 x16
Other
Launch Price
3,999 USD
Production
Active
Predecessor
Radeon Instinct
Server Hopper
Successor
Server Rubin
View Instinct MI325X Details View GB10 Details