AMD Instinct MI350P vs NVIDIA GeForce RTX 5090 Comparison

AMD
RADEON

AMD Instinct MI350P

CORE STATE MI350 128CU
VRAM 144 GB
CLOCK SPEED 2200 MHz
TDP 600 W
BUS WIDTH 8192 bit
ARCHITECTURE CDNA 4.0
nm
PROCESS 3 nm
LAUNCH DATE 2026
VS
NVIDIA
GEFORCE

GeForce RTX 5090

CORE STATE GB202
VRAM 32 GB
CLOCK SPEED 2407 MHz
TDP 575 W
BUS WIDTH 512 bit
ARCHITECTURE Blackwell 2.0
nm
PROCESS 5 nm
LAUNCH DATE 2025

PERFORMANCE BENCHMARKS

3dmark_3dmark_steel_nomad_dx12
N/A
18,355
geekbench_opencl
N/A
334,370
geekbench_vulkan
N/A
376,728
passmark_directx_10
N/A
226
passmark_directx_11
N/A
341
passmark_directx_12
N/A
185
passmark_directx_9
N/A
395
passmark_g2d
N/A
1,413
passmark_g3d
N/A
39,650
passmark_gpu_compute
N/A
26,756

Analysis: AMD Instinct MI350P vs NVIDIA GeForce RTX 5090

AMD Instinct MI350P and NVIDIA GeForce RTX 5090 target completely different workloads, and the recorded data reflects that split clearly. The MI350P is a data-center accelerator with no display outputs, while the RTX 5090 is a client GPU with full graphics and compute APIs. The database shows the RTX 5090 has an average benchmark score of 79,842, placing it in the 92nd percentile of all GPUs, while the MI350P has no recorded benchmark scores and sits at the 50th percentile.

Head-to-Head Benchmarks

The RTX 5090 dominates the measured benchmarks, and the numbers are decisive. In 3DMark Steel Nomad DX12, the RTX 5090 scores 18,355. Its Geekbench OpenCL score is 334,370, and its Vulkan score is 376,728. PassMark results further illustrate the gap: G3D score of 39,650, GPU compute score of 26,756, DirectX 12 score of 185, DirectX 11 score of 341, DirectX 10 score of 226, DirectX 9 score of 395, and G2D score of 1,413. These are the only recorded performance figures in the database for either product. The MI350P has zero benchmark entries, meaning no direct head-to-head comparison exists in our measurements. The nearest rivals for the RTX 5090 show a tight cluster: the NVIDIA Tesla P100 PCIe 16 GB trails by 0.3% with an average score of 79,605, the Tesla P100 PCIe 12 GB trails by 0.6% with 79,396, the AMD Radeon RX 6850M XT trails by 1.1% with 78,940, and the AMD Radeon Pro Vega 64X leads by 1.4% with 80,959. These deltas are small, indicating the RTX 5090’s measured scores sit in a competitive band, but the MI350P has no comparable data points.

The largest measurable win for the RTX 5090 is in raw compute throughput. Its FP32 and FP16 performance both reach 104.8 TFLOPS, while the MI350P delivers 36.04 TFLOPS in both formats. That is a 2.9x advantage for the RTX 5090 in floating-point workloads, based directly on the recorded figures. Texture rate also favors the RTX 5090: 1,636.8 GTexel/s versus 1,126.4 GTexel/s for the MI350P, a 45% lead. Pixel rate is even more lopsided. The RTX 5090 produces 423.6 GPixel/s, while the MI350P shows 0 MPixel/s, reflecting its lack of rasterization hardware. The MI350P has no ROPs, whereas the RTX 5090 has 176. In every measured category, the RTX 5090 leads, and the MI350P’s absence of benchmark scores means there is no counterbalancing performance evidence in the database.

Where Each One Wins

The RTX 5090 wins every recorded benchmark category because it is the only one with measured data. Its 104.8 TFLOPS FP32 and FP16 performance suits tasks that rely on shader or compute throughput, such as rendering, simulation, and general GPU compute. The 423.6 GPixel/s pixel rate and 176 ROPs indicate strong rasterization capability, which the MI350P lacks entirely with 0 ROPs and 0 MPixel/s. The RTX 5090 also supports DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4, making it functional for client-side graphics APIs. The MI350P reports N/A for DirectX, OpenGL, and Vulkan, confirming it has no graphics API support.

The MI350P wins in memory capacity and bandwidth, which are not benchmark scores but architectural advantages. It carries 144 GB of HBM3e memory on an 8192-bit bus, yielding 8.19 TB/s of bandwidth. The RTX 5090 has 32 GB of GDDR7 on a 512-bit bus, with 1.79 TB/s. The MI350P offers 4.5x the memory capacity and 4.6x the bandwidth. For workloads where data residency and memory throughput matter more than raw shader speed, such as large model inference or scientific computing, the MI350P’s memory subsystem is the clear advantage. Its FP16 performance at 36.04 TFLOPS, while lower than the RTX 5090, is still substantial and comes with a 1:1 ratio to FP32, meaning no throughput penalty for half-precision work.

The RTX 5090’s launch MSRP is 1,999 USD, which appears once in this analysis. The MI350P has no launch MSRP recorded. The RTX 5090’s 575 W TDP and 950 W suggested PSU are lower than the MI350P’s 600 W TDP and 1000 W suggested PSU, indicating the MI350P draws more power per the recorded specifications. The MI350P uses a 1000 MHz base clock and 2200 MHz boost clock, while the RTX 5090 runs at 2017 MHz base and 2407 MHz boost. Higher clocks on the RTX 5090 contribute to its compute advantage, though the MI350P’s lower clock is typical for a memory-bound accelerator.

Architecture Differences

The MI350P uses AMD’s CDNA 4.0 architecture, built on a 3 nm process at TSMC. The RTX 5090 uses NVIDIA’s Blackwell 2.0 architecture, on a 5 nm process, also at TSMC. The process node difference is significant: 3 nm versus 5 nm, yet the transistor counts tell a different story. The MI350P has 73,000 million transistors on a 1190 mm² die, resulting in a density of 61.3 million transistors per mm². The RTX 5090 has 92,200 million transistors on a 750 mm² die, giving a density of 122.9 million per mm². Despite the larger process node, the RTX 5090 packs more transistors into a smaller area, nearly doubling the density. The MI350P’s die is 440 mm² larger, but it contains 19,200 million fewer transistors.

The chip identifiers confirm the design split. The MI350P uses a chip named “MI350 128CU,” while the RTX 5090 uses “GB202.” The MI350P has 8,192 shading units, 512 TMUs, and 0 ROPs. The RTX 5090 has 21,760 shading units, 680 TMUs, and 176 ROPs. The RTX 5090 also includes 170 RT cores and 680 tensor cores; the MI350P has no recorded RT or tensor core counts. The MI350P’s texture rate is 1,126.4 GTexel/s, lower than the RTX 5090’s 1,636.8 GTexel/s, and its pixel rate is 0 MPixel/s compared to 423.6 GPixel/s.

Memory architecture diverges completely. The MI350P uses HBM3e, while the RTX 5090 uses GDDR7. The MI350P’s bus width is 8192 bits versus 512 bits, and its bandwidth is 8.19 TB/s versus 1.79 TB/s. Memory clocks also differ: the MI350P runs at 2000 MHz with 8 Gbps effective, while the RTX 5090 runs at 1750 MHz with 28 Gbps effective. The effective data rate is higher on the RTX 5090, but the MI350P’s wider bus and larger capacity dominate. The MI350P has 144 GB of memory, the RTX 5090 has 32 GB.

Power and physical design differ as well. The MI350P has a 600 W TDP, the RTX 5090 has 575 W. Both use a dual-slot cooler and a single 16-pin power connector. The MI350P suggests a 1000 W PSU, the RTX 5090 suggests 950 W. Dimensions favor the MI350P in length: 267 mm versus 304 mm for the RTX 5090. Height is 111 mm versus 137 mm. Width is identical at 40 mm. The MI350P has no display outputs, while the RTX 5090 has 1x HDMI 2.1b and 3x DisplayPort 2.1b. The RTX 5090 supports DirectX 12 Ultimate, OpenGL 4.6, and Vulkan 1.4; the MI350P reports N/A for all three. The RTX 5090 is listed as Active production status, while the MI350P has no production status recorded. Release dates: the MI350P is dated 2026-05-06, the RTX 5090 is dated 2025-01-29. The RTX 5090’s predecessor is GeForce 40 and successor is GeForce 60; the MI350P’s predecessor is Radeon Instinct with no successor listed.

The Verdict

The data supports a clear split: pick the RTX 5090 for any task that involves graphics, rendering, or standard GPU compute benchmarks, because it is the only one with recorded performance scores. Its 92nd percentile ranking and average benchmark score of 79,842 place it among the top GPUs in the database. The MI350P has no benchmark scores, so any performance claim about it cannot be substantiated from our measurements. Its 50th percentile rank is a placeholder, not a measured result.

Pick the MI350P for workloads that require massive memory capacity and bandwidth. The 144 GB HBM3e pool and 8.19 TB/s bandwidth are unmatched by the RTX 5090’s 32 GB GDDR7 and 1.79 TB/s. If the task involves holding large datasets on the GPU, the MI350P’s memory subsystem is the deciding factor. Its 36.04 TFLOPS FP16 performance, while lower than the RTX 5090’s 104.8 TFLOPS, is still substantial and operates at a 1:1 ratio with FP32, which suits mixed-precision workloads.

The RTX 5090 is the only option with graphics API support. DirectX 12 Ultimate, OpenGL 4.6, and Vulkan 1.4 are present on the RTX 5090, while the MI350P lists N/A for all. The RTX 5090 also has display outputs, enabling direct connection to monitors, which the MI350P cannot do. For any use case requiring visual output or standard graphics workloads, the RTX 5090 is the only viable choice from the recorded data.

The RTX 5090’s higher transistor density (122.9M per mm² versus 61.3M) and higher clock speeds (2407 MHz boost versus 2200 MHz boost) explain its compute lead. The MI350P’s lower density and clocks are offset by its larger memory footprint. The RTX 5090 has 21,760 shading units versus 8,192, and 680 TMUs versus 512, which directly explains its 2.9x FP32 advantage. The MI350P’s 0 ROPs and 0 MPixel/s pixel rate confirm it is not designed for rasterization.

Who should pick which: if the workload involves graphics, gaming, rendering, or any DirectX/Vulkan/OpenGL compute, the RTX 5090 is the answer. If the workload is memory-bound, such as large-scale inference or data processing that fits within 144 GB, the MI350P provides the necessary capacity. The RTX 5090’s launch MSRP is 1,999 USD, which is the only pricing information available.

FAQ

Q: Which GPU has a higher FP32 performance?

A: The RTX 5090 delivers 104.8 TFLOPS FP32, while the MI350P provides 36.04 TFLOPS FP32.

Q: What is the memory capacity difference?

A: The MI350P has 144 GB of HBM3e memory, while the RTX 5090 has 32 GB of GDDR7 memory.

Q: Does either GPU support DirectX?

A: The RTX 5090 supports DirectX 12 Ultimate (12_2), while the MI350P reports N/A for DirectX.

Q: Which GPU has more shading units?

A: The RTX 5090 has 21,760 shading units, compared to 8,192 on the MI350P.

Q: What are the TDP ratings?

A: The MI350P has a 600 W TDP, and the RTX 5090 has a 575 W TDP.

Q: Which GPU has display outputs?

A: The RTX 5090 has 1x HDMI 2.1b and 3x DisplayPort 2.1b, while the MI350P has no display outputs.

Specification Differences

| Feature | AMD Instinct MI350P | NVIDIA GeForce RTX 5090 |

| --- | --- | --- |

| Architecture | CDNA 4.0 | Blackwell 2.0 |

| Process node | 3 nm | 5 nm |

| Transistors | 73,000 million | 92,200 million |

| Die size | 1190 mm² | 750 mm² |

| Transistor density | 61.3M / mm² | 122.9M / mm² |

| Base clock | 1000 MHz | 2017 MHz |

| Boost clock | 2200 MHz | 2407 MHz |

| Memory clock | 2000 MHz 8 Gbps effective | 1750 MHz 28 Gbps effective |

| Memory size | 144 GB | 32 GB |

| Memory type | HBM3e | GDDR7 |

| Memory bus width | 8192 bit | 512 bit |

| Memory bandwidth | 8.19 TB/s | 1.79 TB/s |

| Shading units | 8192 | 21760 |

| TMUs | 512 | 680 |

| ROPs | 0 | 176 |

| RT cores | N/A | 170 |

| Tensor cores | N/A | 680 |

| Pixel rate | 0 MPixel/s | 423.6 GPixel/s |

| Texture rate | 1,126.4 GTexel/s | 1,636.8 GTexel/s |

| FP32 | 36.04 TFLOPS | 104.8 TFLOPS |

| FP16 | 36.04 TFLOPS (1:1) | 104.8 TFLOPS (1:1) |

| TDP | 600 W | 575 W |

| Suggested PSU | 1000 W | 950 W |

| Power connectors | 1x 16-pin | 1x 16-pin |

| Slot width | Dual-slot | Dual-slot |

| Length | 267 mm | 304 mm |

| Height | 111 mm | 137 mm |

| Width | 40 mm | 40 mm |

| Display outputs | No outputs | 1x HDMI 2.1b, 3x DisplayPort 2.1b |

| DirectX | N/A | 12 Ultimate (12_2) |

| OpenGL | N/A | 4.6 |

| Vulkan | N/A | 1.4 |

| Bus interface | PCIe 5.0 x16 | PCIe 5.0 x16 |

| Production status | N/A | Active |

| Release date | 2026-05-06 | 2025-01-29 |

| Predecessor | Radeon Instinct | GeForce 40 |

| Successor | N/A | GeForce 60 |

| Launch MSRP | N/A | 1,999 USD |

| Average benchmark score | 0 | 79,842 |

| Percentile vs all GPUs | 50 | 92 |

DETAILED SPECIFICATIONS

SPECIFICATION
Instinct MI350P
RTX 5090
Core Specs
Shading Units
8,192
21,760 +165.6%
Shaders
8,192
21,760 +165.6%
TMUs
512
680 +32.8%
ROPs
0
176 +∞%
Compute Units
128
—
SM Count
—
170
Clocks
Base Clock
1000 MHz
2017 MHz
Boost Clock
2200 MHz
2407 MHz
Memory Clock
2000 MHz 8 Gbps effective
1750 MHz 28 Gbps effective
Memory
Memory Size
144 GB
32 GB
VRAM (MB)
147,456
32,768 -77.8%
Memory Type
HBM3e
GDDR7
Memory Bus
8192 bit
512 bit
Bandwidth
8.19 TB/s
1.79 TB/s
Cache
L1 Cache
16 KB (per CU)
128 KB (per SM)
L2 Cache
16 MB
96 MB
L3 Cache
128 MB
—
Performance
Pixel Rate
0 MPixel/s
423.6 GPixel/s
Texture Rate
1,126.4 GTexel/s
1,636.8 GTexel/s
FP32 (TFLOPS)
36.04 TFLOPS
104.8 TFLOPS
FP64 (TFLOPS)
18.02 TFLOPS (1:2)
1.637 TFLOPS (1:64)
FP16 (TFLOPS)
36.04 TFLOPS (1:1)
104.8 TFLOPS (1:1)
AI/RT
RT Cores
—
170
Tensor Cores
—
680
Matrix Cores
512
—
Power
TDP
600 W
575 W
TDP (W)
600
575 -4.2%
Suggested PSU
1000 W
950 W
Power Connectors
1x 16-pin
1x 16-pin
Architecture
Architecture
CDNA 4.0
Blackwell 2.0
GPU Name
MI350 128CU
GB202
Generation
Instinct (MIx)
GeForce 50
Process Size
3 nm
5 nm
Transistors
73,000 million
92,200 million
Die Size
1190 mm²
750 mm²
Foundry
TSMC
TSMC
Density
61.3M / mm²
122.9M / mm²
AMD MCM
MCM
2
—
API Support
DirectX
—
12 Ultimate (12_2)
OpenGL
—
4.6
Vulkan
—
1.4
OpenCL
3.0
3.0
CUDA
—
12.0
Shader Model
—
6.9
Physical
Slot Width
Dual-slot
Dual-slot
Length
267 mm 10.5 inches
304 mm 12 inches
Height
111 mm 4.4 inches
137 mm 5.4 inches
Outputs
No outputs
1x HDMI 2.1b3x DisplayPort 2.1b
Bus Interface
PCIe 5.0 x16
PCIe 5.0 x16
Other
Launch Price
—
1,999 USD
Production
—
Active
Predecessor
Radeon Instinct
GeForce 40
Successor
—
GeForce 60
View Instinct MI350P Details View GeForce RTX 5090 Details