AMD Instinct MI350P vs NVIDIA H200 NVL Comparison

AMD
RADEON

AMD Instinct MI350P

CORE STATE MI350 128CU
VRAM 144 GB
CLOCK SPEED 2200 MHz
TDP 600 W
BUS WIDTH 8192 bit
ARCHITECTURE CDNA 4.0
nm
PROCESS 3 nm
LAUNCH DATE 2026
VS
NVIDIA
GEFORCE

H200 NVL

CORE STATE GH100
VRAM 141 GB
CLOCK SPEED 1785 MHz
TDP 600 W
BUS WIDTH 6144 bit
ARCHITECTURE Hopper
nm
PROCESS 5 nm
LAUNCH DATE 2024

PERFORMANCE BENCHMARKS

geekbench_opencl
N/A
334,891

Analysis: AMD Instinct MI350P vs NVIDIA H200 NVL

Head-to-Head Benchmarks

The recorded database contains only one benchmark result for this pairing: a Geekbench OpenCL score for the NVIDIA H200 NVL. The AMD Instinct MI350P has no benchmark entries in the database, so a direct numerical comparison between the two is not possible from the available data. The H200 NVL achieves a Geekbench OpenCL score of 334,891, placing it at the 100th percentile among all GPUs in the database.

The H200 NVL's nearest rivals provide context for its performance. The NVIDIA B200 scores 345,482, which is 3.1% higher than the H200 NVL. The AMD Instinct MI300X scores 317,994, which is 5.3% lower than the H200 NVL. The NVIDIA B300 SXM6 AC scores 369,831, 9.4% higher than the H200 NVL. The NVIDIA L40S scores 295,763, which is 13.2% lower than the H200 NVL. These comparisons show the H200 NVL sitting near the top of the database, with only the B200 and B300 SXM6 AC ahead of it among its listed rivals.

The MI350P carries an average benchmark score of 0 and a percentile rank of 50, indicating that no measured results have been recorded for it. The database shows zero head-to-head benchmark entries, zero wins for the MI350P, and zero wins for the H200 NVL in this pairing. Without a measured score for the MI350P, the data cannot establish a performance hierarchy between these two accelerators.

Architecture Differences

The MI350P uses the MI350 128CU chip built on CDNA 4.0 architecture, manufactured on a 3 nm process at TSMC. The H200 NVL uses the GH100 chip built on Hopper architecture, manufactured on a 5 nm process at TSMC. Both are dual-slot cards with no display outputs, and both connect via PCIe 5.0 x16.

The MI350P contains 73,000 million transistors on a die size of 1190 mm², yielding a transistor density of 61.3 million per mm². The H200 NVL contains 80,000 million transistors on a die size of 814 mm², yielding a transistor density of 98.3 million per mm². The H200 NVL packs more transistors into a smaller die, giving it a higher density despite the older process node.

Clock speeds differ substantially. The MI350P has a base clock of 1000 MHz and a boost clock of 2200 MHz. The H200 NVL has a base clock of 1365 MHz and a boost clock of 1785 MHz. The MI350P has a higher boost clock by a significant margin, while the H200 NVL starts from a higher base clock.

Memory configurations diverge in capacity, bandwidth, and bus width. The MI350P offers 144 GB of HBM3e memory on an 8192-bit bus, delivering 8.19 TB/s of bandwidth. The H200 NVL offers 141 GB of HBM3e memory on a 6144-bit bus, delivering 4.89 TB/s of bandwidth. The MI350P leads in memory capacity, bus width, and bandwidth, with the bandwidth advantage being particularly large.

The compute unit counts differ sharply. The MI350P has 8192 shading units, 512 texture mapping units, and no ROPs (0 MPixel/s pixel rate). The H200 NVL has 16,896 shading units, 528 texture mapping units, and 24 ROPs, with a pixel rate of 42.84 GPixel/s. The H200 NVL has more than double the shading units and a higher texture rate at 942.5 GTexel/s versus the MI350P's 1,126.4 GTexel/s, which actually favors the MI350P in texture throughput.

The H200 NVL includes 528 tensor cores, while the MI350P lists no tensor cores in the database. Floating-point performance shows a clear split: the MI350P delivers 36.04 TFLOPS for both FP32 and FP16 (1:1 ratio), while the H200 NVL delivers 60.32 TFLOPS FP32 and 120.6 TFLOPS FP16 (2:1 ratio). The H200 NVL leads in both precision formats, with its FP16 advantage being especially pronounced.

Power delivery differs in connector type. The MI350P uses a single 16-pin connector, while the H200 NVL uses an 8-pin EPS connector. Both have a TDP of 600 W and a suggested PSU of 1000 W. Both cards measure 267 mm in length and 111 mm in height; the MI350P has a width of 40 mm, while the H200 NVL's width is not recorded.

Where Each One Wins

Based on the recorded data, the NVIDIA H200 NVL wins in raw compute throughput. Its FP32 performance of 60.32 TFLOPS is 67% higher than the MI350P's 36.04 TFLOPS. Its FP16 performance of 120.6 TFLOPS is more than three times the MI350P's 36.04 TFLOPS. The H200 NVL also has more shading units (16,896 versus 8,192) and includes tensor cores, which the MI350P lacks in the database.

The AMD Instinct MI350P wins in memory bandwidth and capacity. Its 8.19 TB/s bandwidth is 67% higher than the H200 NVL's 4.89 TB/s. It also offers 144 GB of memory versus 141 GB, and a wider 8192-bit bus versus 6144-bit. The MI350P has a higher boost clock at 2200 MHz versus 1785 MHz, and a higher texture rate at 1,126.4 GTexel/s versus 942.5 GTexel/s.

The H200 NVL has a production status of Active, while the MI350P's production status is not recorded. The H200 NVL was released on 2024-11-17, while the MI350P has a release date of 2026-05-06. The H200 NVL has a recorded successor in Server Blackwell, while the MI350P lists Radeon Instinct as its predecessor and no successor.

The H200 NVL's benchmark score of 334,891 at the 100th percentile confirms its position among the fastest accelerators in the database. The MI350P has no measured score, so its actual performance position cannot be determined from the data.

FAQ

Q: How does the AMD Instinct MI350P compare to the NVIDIA H200 NVL in memory bandwidth?

A: The MI350P delivers 8.19 TB/s of bandwidth from 144 GB of HBM3e memory on an 8192-bit bus. The H200 NVL delivers 4.89 TB/s from 141 GB of HBM3e memory on a 6144-bit bus. The MI350P has 67% higher bandwidth and a 2048-bit wider bus.

Q: Which accelerator has higher FP32 performance?

A: The NVIDIA H200 NVL has 60.32 TFLOPS of FP32 performance, compared to the AMD Instinct MI350P's 36.04 TFLOPS. The H200 NVL is 67% faster in FP32 according to the recorded specifications.

Q: What does the H200 NVL's benchmark score indicate about its performance relative to rivals?

A: The H200 NVL scores 334,891 in Geekbench OpenCL, placing it at the 100th percentile. It is 3.1% behind the NVIDIA B200 (345,482), 5.3% ahead of the AMD Instinct MI300X (317,994), 9.4% behind the NVIDIA B300 SXM6 AC (369,831), and 13.2% ahead of the NVIDIA L40S (295,763).

Q: Does the AMD Instinct MI350P have tensor cores?

A: The database does not list tensor cores for the MI350P. The NVIDIA H200 NVL has 528 tensor cores. The MI350P's FP16 performance is 36.04 TFLOPS at a 1:1 ratio, while the H200 NVL's FP16 performance is 120.6 TFLOPS at a 2:1 ratio.

Q: What are the process node differences between these two accelerators?

A: The AMD Instinct MI350P uses a 3 nm process at TSMC, while the NVIDIA H200 NVL uses a 5 nm process at TSMC. Despite the older node, the H200 NVL has a higher transistor density at 98.3 million per mm² versus 61.3 million per mm² for the MI350P.

Q: When were these accelerators released?

A: The NVIDIA H200 NVL has a release date of 2024-11-17 and is marked as Active in production. The AMD Instinct MI350P has a release date of 2026-05-06, with no production status recorded.

The Verdict

The data points to a clear split in strengths. The NVIDIA H200 NVL is the compute leader, dominating in FP32 (60.32 TFLOPS versus 36.04 TFLOPS), FP16 (120.6 TFLOPS versus 36.04 TFLOPS), shading units (16,896 versus 8,192), and it includes tensor cores that the MI350P lacks. Its measured benchmark score of 334,891 at the 100th percentile confirms its standing among the fastest accelerators in the database.

The AMD Instinct MI350P is the memory leader, with 67% higher bandwidth (8.19 TB/s versus 4.89 TB/s), a wider memory bus (8192-bit versus 6144-bit), and slightly more memory capacity (144 GB versus 141 GB). It also has a higher boost clock (2200 MHz versus 1785 MHz) and higher texture rate (1,126.4 GTexel/s versus 942.5 GTexel/s).

The H200 NVL is the safer selection for general compute workloads given its verified benchmark performance, its Active production status, and its established position against named rivals. The MI350P, with no recorded benchmark scores and a future release date, cannot be positioned against the H200 NVL in measured performance.

For workloads that depend on memory bandwidth, the MI350P's 8.19 TB/s is the strongest figure in this comparison. For workloads that depend on floating-point throughput or tensor operations, the H200 NVL is the clear choice from the recorded data. The H200 NVL's 120.6 TFLOPS FP16 output is more than triple the MI350P's 36.04 TFLOPS, making it the dominant option for mixed-precision compute.

The MI350P's 3 nm process node and higher boost clock indicate architectural differences that could matter in specific contexts, but without benchmark measurements, the database cannot quantify their impact. The H200 NVL's 5.3% lead over the MI300X and its proximity to the B200 (3.1% behind) provide concrete reference points. The MI350P has no such reference points.

Specification Differences

| Specification | AMD Instinct MI350P | NVIDIA H200 NVL |

|---|---|---|

| Chip | MI350 128CU | GH100 |

| Architecture | CDNA 4.0 | Hopper |

| Process Node | 3 nm | 5 nm |

| Transistors | 73,000 million | 80,000 million |

| Die Size | 1190 mm² | 814 mm² |

| Transistor Density | 61.3M / mm² | 98.3M / mm² |

| Base Clock | 1000 MHz | 1365 MHz |

| Boost Clock | 2200 MHz | 1785 MHz |

| Memory Clock | 2000 MHz 8 Gbps effective | 1593 MHz 6.4 Gbps effective |

| Memory Size | 144 GB | 141 GB |

| Memory Type | HBM3e | HBM3e |

| Memory Bus Width | 8192 bit | 6144 bit |

| Memory Bandwidth | 8.19 TB/s | 4.89 TB/s |

| Shading Units | 8192 | 16896 |

| TMUs | 512 | 528 |

| ROPs | 0 | 24 |

| Tensor Cores | Not listed | 528 |

| Pixel Rate | 0 MPixel/s | 42.84 GPixel/s |

| Texture Rate | 1,126.4 GTexel/s | 942.5 GTexel/s |

| FP32 Performance | 36.04 TFLOPS | 60.32 TFLOPS |

| FP16 Performance | 36.04 TFLOPS (1:1) | 120.6 TFLOPS (2:1) |

| TDP | 600 W | 600 W |

| Power Connectors | 1x 16-pin | 8-pin EPS |

| Suggested PSU | 1000 W | 1000 W |

| Slot Width | Dual-slot | Dual-slot |

| Bus Interface | PCIe 5.0 x16 | PCIe 5.0 x16 |

| Display Outputs | No outputs | No outputs |

| Dimensions | 267 mm x 111 mm x 40 mm | 267 mm x 111 mm |

| Release Date | 2026-05-06 | 2024-11-17 |

| Production Status | Not recorded | Active |

| Predecessor | Radeon Instinct | Server Ada |

| Successor | Not recorded | Server Blackwell |

DETAILED SPECIFICATIONS

SPECIFICATION
Instinct MI350P
H200 NVL
Core Specs
Shading Units
8,192
16,896 +106.3%
Shaders
8,192
16,896 +106.3%
TMUs
512
528 +3.1%
ROPs
0
24 +∞%
Compute Units
128
—
SM Count
—
132
Clocks
Base Clock
1000 MHz
1365 MHz
Boost Clock
2200 MHz
1785 MHz
Memory Clock
2000 MHz 8 Gbps effective
1593 MHz 6.4 Gbps effective
Memory
Memory Size
144 GB
141 GB
VRAM (MB)
147,456
144,384 -2.1%
Memory Type
HBM3e
HBM3e
Memory Bus
8192 bit
6144 bit
Bandwidth
8.19 TB/s
4.89 TB/s
Cache
L1 Cache
16 KB (per CU)
256 KB (per SM)
L2 Cache
16 MB
50 MB
L3 Cache
128 MB
—
Performance
Pixel Rate
0 MPixel/s
42.84 GPixel/s
Texture Rate
1,126.4 GTexel/s
942.5 GTexel/s
FP32 (TFLOPS)
36.04 TFLOPS
60.32 TFLOPS
FP64 (TFLOPS)
18.02 TFLOPS (1:2)
30.16 TFLOPS (1:2)
FP16 (TFLOPS)
36.04 TFLOPS (1:1)
120.6 TFLOPS (2:1)
AI/RT
Tensor Cores
—
528
Matrix Cores
512
—
Power
TDP
600 W
600 W
TDP (W)
600
600 0.0%
Suggested PSU
1000 W
1000 W
Power Connectors
1x 16-pin
8-pin EPS
Architecture
Architecture
CDNA 4.0
Hopper
GPU Name
MI350 128CU
GH100
Generation
Instinct (MIx)
Server Hopper (Hxx)
Process Size
3 nm
5 nm
Transistors
73,000 million
80,000 million
Die Size
1190 mm²
814 mm²
Foundry
TSMC
TSMC
Density
61.3M / mm²
98.3M / mm²
AMD MCM
MCM
2
—
API Support
OpenCL
3.0
3.0
CUDA
—
9.0
Physical
Slot Width
Dual-slot
Dual-slot
Length
267 mm 10.5 inches
267 mm 10.5 inches
Height
111 mm 4.4 inches
111 mm 4.4 inches
Outputs
No outputs
No outputs
Bus Interface
PCIe 5.0 x16
PCIe 5.0 x16
Other
Production
—
Active
Predecessor
Radeon Instinct
Server Ada
Successor
—
Server Blackwell
View Instinct MI350P Details View H200 NVL Details