AMD Instinct MI350P vs NVIDIA GB10 Comparison

AMD
RADEON

AMD Instinct MI350P

CORE STATE MI350 128CU
VRAM 144 GB
CLOCK SPEED 2200 MHz
TDP 600 W
BUS WIDTH 8192 bit
ARCHITECTURE CDNA 4.0
nm
PROCESS 3 nm
LAUNCH DATE 2026
VS
NVIDIA
GEFORCE

GB10

CORE STATE GB20B
VRAM 128 GB
CLOCK SPEED 2418 MHz
TDP 140 W
BUS WIDTH 256 bit
ARCHITECTURE Blackwell 2.0
nm
PROCESS 5 nm
LAUNCH DATE 2025

PERFORMANCE BENCHMARKS

geekbench_opencl
N/A
120,137
geekbench_vulkan
N/A
114,648

Analysis: AMD Instinct MI350P vs NVIDIA GB10

The Verdict

The AMD Instinct MI350P and NVIDIA GB10 serve fundamentally different segments of the compute market, and the recorded data makes the separation clear. The MI350P is a high-power, high-bandwidth accelerator aimed at large-scale compute workloads, while the GB10 is a power-efficient, integrated solution with a verified benchmark presence. The GB10 holds a 95th percentile ranking among all GPUs in the database, with an average benchmark score of 117,393 across Geekbench OpenCL and Vulkan tests. The MI350P, by contrast, has no recorded benchmark scores and sits at the 50th percentile, indicating that its performance profile is not yet quantified in the database.

For users requiring substantial memory bandwidth and raw texture throughput, the MI350P is the clear choice based on its specifications. Its 8.19 TB/s memory bandwidth is an order of magnitude higher than the GB10's 273.2 GB/s, and its 1,126.4 GTexel/s texture rate exceeds the GB10's 928.5 GTexel/s. Conversely, the GB10 offers a verified performance baseline, a lower power envelope, and a display output, making it suitable for deployments where a compact, self-contained compute node is needed. The MI350P has no display outputs, reinforcing its role as a dedicated compute accelerator. The GB10 is actively in production and has a defined successor, while the MI350P's production status is not recorded.

Architecture Differences

The two processors come from different architectural lineages. The MI350P uses the CDNA 4.0 architecture on a 3 nm TSMC process, while the GB10 uses Blackwell 2.0 on a 5 nm TSMC process. The MI350P integrates 73,000 million transistors on a 1,190 mm² die, yielding a transistor density of 61.3 million per square millimeter. The GB10 uses a 382 mm² die with an unknown transistor count and no recorded density figure. The MI350P's die is more than three times larger, reflecting its focus on massive parallel compute resources.

Core configurations diverge sharply. The MI350P has 8,192 shading units, 512 texture mapping units, and no ROPs, resulting in a pixel rate of 0 MPixel/s. The GB10 has 6,144 shading units, 384 TMUs, and 48 ROPs, producing a pixel rate of 116.1 GPixel/s. The MI350P lacks dedicated ray tracing and tensor cores in the recorded data, while the GB10 includes 48 ray tracing cores and 384 tensor cores. Both processors report FP32 and FP16 performance at a 1:1 ratio, with the MI350P delivering 36.04 TFLOPS in both precisions and the GB10 delivering 29.71 TFLOPS in both.

Memory architecture is a major differentiator. The MI350P uses 144 GB of HBM3e across an 8,192-bit bus, achieving 8.19 TB/s bandwidth. The GB10 uses 128 GB of LPDDR5X across a 256-bit bus, achieving 273.2 GB/s. The MI350P's memory clock is 2000 MHz with 8 Gbps effective data rate; the GB10's memory clock is 1067 MHz with 8.5 Gbps effective. The MI350P's memory bus width is 32 times larger, which accounts for its massive bandwidth advantage. Power consumption reflects this disparity: the MI350P has a 600 W TDP and requires a 1000 W suggested PSU with a single 16-pin connector, while the GB10 has a 140 W TDP and a 300 W suggested PSU with no power connectors, as it is an integrated graphics processor (IGP).

Physical specifications also differ. The MI350P is a dual-slot card measuring 267 mm in length, 111 mm in height, and 40 mm in width. The GB10 is an IGP measuring 150 mm by 51 mm by 150 mm. The MI350P uses a PCIe 5.0 x16 interface, as does the GB10, but the MI350P has no display outputs while the GB10 includes a single HDMI port. Both processors report N/A for DirectX, OpenGL, and Vulkan API support, indicating they are not intended for conventional graphics workloads.

Head-to-Head Benchmarks

The database contains no direct head-to-head benchmark results between the MI350P and the GB10. However, the GB10 has two recorded Geekbench scores: 120,137 in OpenCL and 114,648 in Vulkan. These yield an average benchmark score of 117,393. The MI350P has no recorded benchmark scores, so direct numerical comparison is impossible from the available data.

The GB10's nearest rivals provide context for its performance. The NVIDIA RTX 4000 SFF Ada Generation scores 117,088, which is 0.3% below the GB10. The AMD Radeon PRO W7700 scores 118,976, which is 1.3% above the GB10. The NVIDIA Tesla V100 SXM2 16 GB scores 114,395, which is 2.6% below. The NVIDIA RTX A5500 Mobile scores 113,944, which is 3.0% below. These deltas show the GB10 performing within a narrow band around its nearest competitors, with the largest gap being 3.0% behind the RTX A5500 Mobile and 1.3% behind the Radeon PRO W7700. The GB10 outperforms two of its four nearest rivals by 0.3% and 2.6%, respectively.

The MI350P's FP32 throughput of 36.04 TFLOPS is 21.3% higher than the GB10's 29.71 TFLOPS. Its texture rate of 1,126.4 GTexel/s is 21.3% higher than the GB10's 928.5 GTexel/s. The MI350P's memory bandwidth of 8.19 TB/s is nearly 30 times the GB10's 273.2 GB/s. These specification advantages are substantial, but without benchmark scores for the MI350P, the actual workload performance cannot be quantified in the database.

FAQ

Q: What is the average benchmark score for the NVIDIA GB10?

A: The GB10 has an average benchmark score of 117,393, derived from a Geekbench OpenCL score of 120,137 and a Geekbench Vulkan score of 114,648. It ranks in the 95th percentile among all GPUs in the database.

Q: How does the GB10 compare to its nearest rivals?

A: The GB10 is 0.3% ahead of the NVIDIA RTX 4000 SFF Ada Generation and 2.6% ahead of the NVIDIA Tesla V100 SXM2 16 GB. It trails the AMD Radeon PRO W7700 by 1.3% and the NVIDIA RTX A5500 Mobile by 3.0%.

Q: Does the AMD Instinct MI350P have any recorded benchmark scores?

A: No. The MI350P has an empty benchmark array in the database, giving it an average benchmark score of 0 and a 50th percentile ranking among all GPUs.

Q: What are the memory capacities and types of each processor?

A: The MI350P has 144 GB of HBM3e memory with an 8,192-bit bus and 8.19 TB/s bandwidth. The GB10 has 128 GB of LPDDR5X memory with a 256-bit bus and 273.2 GB/s bandwidth.

Q: What are the power requirements for each processor?

A: The MI350P has a 600 W TDP with a suggested PSU of 1000 W and a single 16-pin power connector. The GB10 has a 140 W TDP with a suggested PSU of 300 W and no power connectors, as it is an integrated processor.

Q: Which processor includes display outputs?

A: The GB10 includes one HDMI output. The MI350P has no display outputs, indicating it is designed exclusively for compute tasks.

Where Each One Wins

The MI350P wins decisively in memory bandwidth and capacity. Its 8.19 TB/s bandwidth and 144 GB of HBM3e are designed for data-intensive workloads where large datasets must be fed to compute units rapidly. The 8,192-bit memory bus is the enabling factor, allowing the MI350P to move data at a scale the GB10 cannot approach. Its FP32 throughput of 36.04 TFLOPS is also higher, giving it a 21.3% advantage in raw single-precision compute. The MI350P's texture rate of 1,126.4 GTexel/s similarly exceeds the GB10's 928.5 GTexel/s, though the MI350P's lack of ROPs and pixel output means it is not suited for rasterization tasks. The MI350P's 3 nm process node and 73,000 million transistors indicate a design optimized for maximum parallel throughput. Its dual-slot form factor and 600 W TDP confirm that it is intended for dedicated server installations with ample power delivery.

The GB10 wins in verified performance and operational efficiency. Its average benchmark score of 117,393, backed by two Geekbench runs, provides a concrete performance baseline that the MI350P lacks. The GB10's 140 W TDP is less than a quarter of the MI350P's 600 W TDP, and its 300 W suggested PSU reflects a much lower system power requirement. The GB10 includes 48 ray tracing cores and 384 tensor cores, which the MI350P does not list, suggesting the GB10 is equipped for workloads that use these specialized units. The GB10's 48 ROPs and 116.1 GPixel/s pixel rate give it a graphics output capability that the MI350P entirely lacks. The GB10's IGP form factor, 150 mm by 51 mm by 150 mm dimensions, and single HDMI output make it a self-contained unit that can fit in compact systems. Its production status is active, and it has a defined successor in Server Rubin, indicating an active product lifecycle.

The recorded data also shows the GB10's competitive position among its peers. It outperforms the RTX 4000 SFF Ada Generation by 0.3% and the Tesla V100 SXM2 16 GB by 2.6%, while trailing the Radeon PRO W7700 by 1.3% and the RTX A5500 Mobile by 3.0%. These deltas place it in the upper tier of its comparison group. The MI350P, with no benchmarks and no rivals listed, has no such positioning in the database.

For compute workloads that prioritize memory bandwidth and raw FP32 throughput, the MI350P's specifications are superior. For deployments that require a compact, low-power compute node with verified performance and graphics output, the GB10 is the only option with recorded data. The MI350P's lack of display outputs and higher power envelope restrict it to headless accelerator roles, while the GB10's integrated design and HDMI output broaden its applicability. The choice between them depends on whether the workload demands the MI350P's extreme memory subsystem or the GB10's balanced, validated performance profile.

DETAILED SPECIFICATIONS

SPECIFICATION
Instinct MI350P
GB10
Core Specs
Shading Units
8,192
6,144 -25.0%
Shaders
8,192
6,144 -25.0%
TMUs
512
384 -25.0%
ROPs
0
48 +∞%
Compute Units
128
SM Count
48
Clocks
Base Clock
1000 MHz
1665 MHz
Boost Clock
2200 MHz
2418 MHz
Memory Clock
2000 MHz 8 Gbps effective
1067 MHz 8.5 Gbps effective
Memory
Memory Size
144 GB
128 GB
VRAM (MB)
147,456
131,072 -11.1%
Memory Type
HBM3e
LPDDR5X
Memory Bus
8192 bit
256 bit
Bandwidth
8.19 TB/s
273.2 GB/s
Cache
L1 Cache
16 KB (per CU)
128 KB (per SM)
L2 Cache
16 MB
50 MB
L3 Cache
128 MB
Performance
Pixel Rate
0 MPixel/s
116.1 GPixel/s
Texture Rate
1,126.4 GTexel/s
928.5 GTexel/s
FP32 (TFLOPS)
36.04 TFLOPS
29.71 TFLOPS
FP64 (TFLOPS)
18.02 TFLOPS (1:2)
464.3 GFLOPS (1:64)
FP16 (TFLOPS)
36.04 TFLOPS (1:1)
29.71 TFLOPS (1:1)
AI/RT
RT Cores
48
Tensor Cores
384
Matrix Cores
512
Power
TDP
600 W
140 W
TDP (W)
600
140 -76.7%
Suggested PSU
1000 W
300 W
Power Connectors
1x 16-pin
None
Architecture
Architecture
CDNA 4.0
Blackwell 2.0
GPU Name
MI350 128CU
GB20B
Generation
Instinct (MIx)
Server Blackwell (Bxx)
Process Size
3 nm
5 nm
Transistors
73,000 million
unknown
Die Size
1190 mm²
382 mm²
Foundry
TSMC
TSMC
Density
61.3M / mm²
AMD MCM
MCM
2
API Support
OpenCL
3.0
3.0
CUDA
12.1
Physical
Slot Width
Dual-slot
IGP
Length
267 mm 10.5 inches
150 mm 5.9 inches
Height
111 mm 4.4 inches
51 mm 2 inches
Outputs
No outputs
1x HDMI
Bus Interface
PCIe 5.0 x16
PCIe 5.0 x16
Other
Launch Price
3,999 USD
Production
Active
Predecessor
Radeon Instinct
Server Hopper
Successor
Server Rubin
View Instinct MI350P Details View GB10 Details