AMD Instinct MI350P vs NVIDIA GeForce RTX 5090 Comparison
AMD Instinct MI350P
GeForce RTX 5090
PERFORMANCE BENCHMARKS
Analysis: AMD Instinct MI350P vs NVIDIA GeForce RTX 5090
AMD Instinct MI350P and NVIDIA GeForce RTX 5090 target completely different workloads, and the recorded data reflects that split clearly. The MI350P is a data-center accelerator with no display outputs, while the RTX 5090 is a client GPU with full graphics and compute APIs. The database shows the RTX 5090 has an average benchmark score of 79,842, placing it in the 92nd percentile of all GPUs, while the MI350P has no recorded benchmark scores and sits at the 50th percentile.
Head-to-Head Benchmarks
The RTX 5090 dominates the measured benchmarks, and the numbers are decisive. In 3DMark Steel Nomad DX12, the RTX 5090 scores 18,355. Its Geekbench OpenCL score is 334,370, and its Vulkan score is 376,728. PassMark results further illustrate the gap: G3D score of 39,650, GPU compute score of 26,756, DirectX 12 score of 185, DirectX 11 score of 341, DirectX 10 score of 226, DirectX 9 score of 395, and G2D score of 1,413. These are the only recorded performance figures in the database for either product. The MI350P has zero benchmark entries, meaning no direct head-to-head comparison exists in our measurements. The nearest rivals for the RTX 5090 show a tight cluster: the NVIDIA Tesla P100 PCIe 16 GB trails by 0.3% with an average score of 79,605, the Tesla P100 PCIe 12 GB trails by 0.6% with 79,396, the AMD Radeon RX 6850M XT trails by 1.1% with 78,940, and the AMD Radeon Pro Vega 64X leads by 1.4% with 80,959. These deltas are small, indicating the RTX 5090’s measured scores sit in a competitive band, but the MI350P has no comparable data points.
The largest measurable win for the RTX 5090 is in raw compute throughput. Its FP32 and FP16 performance both reach 104.8 TFLOPS, while the MI350P delivers 36.04 TFLOPS in both formats. That is a 2.9x advantage for the RTX 5090 in floating-point workloads, based directly on the recorded figures. Texture rate also favors the RTX 5090: 1,636.8 GTexel/s versus 1,126.4 GTexel/s for the MI350P, a 45% lead. Pixel rate is even more lopsided. The RTX 5090 produces 423.6 GPixel/s, while the MI350P shows 0 MPixel/s, reflecting its lack of rasterization hardware. The MI350P has no ROPs, whereas the RTX 5090 has 176. In every measured category, the RTX 5090 leads, and the MI350P’s absence of benchmark scores means there is no counterbalancing performance evidence in the database.
Where Each One Wins
The RTX 5090 wins every recorded benchmark category because it is the only one with measured data. Its 104.8 TFLOPS FP32 and FP16 performance suits tasks that rely on shader or compute throughput, such as rendering, simulation, and general GPU compute. The 423.6 GPixel/s pixel rate and 176 ROPs indicate strong rasterization capability, which the MI350P lacks entirely with 0 ROPs and 0 MPixel/s. The RTX 5090 also supports DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4, making it functional for client-side graphics APIs. The MI350P reports N/A for DirectX, OpenGL, and Vulkan, confirming it has no graphics API support.
The MI350P wins in memory capacity and bandwidth, which are not benchmark scores but architectural advantages. It carries 144 GB of HBM3e memory on an 8192-bit bus, yielding 8.19 TB/s of bandwidth. The RTX 5090 has 32 GB of GDDR7 on a 512-bit bus, with 1.79 TB/s. The MI350P offers 4.5x the memory capacity and 4.6x the bandwidth. For workloads where data residency and memory throughput matter more than raw shader speed, such as large model inference or scientific computing, the MI350P’s memory subsystem is the clear advantage. Its FP16 performance at 36.04 TFLOPS, while lower than the RTX 5090, is still substantial and comes with a 1:1 ratio to FP32, meaning no throughput penalty for half-precision work.
The RTX 5090’s launch MSRP is 1,999 USD, which appears once in this analysis. The MI350P has no launch MSRP recorded. The RTX 5090’s 575 W TDP and 950 W suggested PSU are lower than the MI350P’s 600 W TDP and 1000 W suggested PSU, indicating the MI350P draws more power per the recorded specifications. The MI350P uses a 1000 MHz base clock and 2200 MHz boost clock, while the RTX 5090 runs at 2017 MHz base and 2407 MHz boost. Higher clocks on the RTX 5090 contribute to its compute advantage, though the MI350P’s lower clock is typical for a memory-bound accelerator.
Architecture Differences
The MI350P uses AMD’s CDNA 4.0 architecture, built on a 3 nm process at TSMC. The RTX 5090 uses NVIDIA’s Blackwell 2.0 architecture, on a 5 nm process, also at TSMC. The process node difference is significant: 3 nm versus 5 nm, yet the transistor counts tell a different story. The MI350P has 73,000 million transistors on a 1190 mm² die, resulting in a density of 61.3 million transistors per mm². The RTX 5090 has 92,200 million transistors on a 750 mm² die, giving a density of 122.9 million per mm². Despite the larger process node, the RTX 5090 packs more transistors into a smaller area, nearly doubling the density. The MI350P’s die is 440 mm² larger, but it contains 19,200 million fewer transistors.
The chip identifiers confirm the design split. The MI350P uses a chip named “MI350 128CU,” while the RTX 5090 uses “GB202.” The MI350P has 8,192 shading units, 512 TMUs, and 0 ROPs. The RTX 5090 has 21,760 shading units, 680 TMUs, and 176 ROPs. The RTX 5090 also includes 170 RT cores and 680 tensor cores; the MI350P has no recorded RT or tensor core counts. The MI350P’s texture rate is 1,126.4 GTexel/s, lower than the RTX 5090’s 1,636.8 GTexel/s, and its pixel rate is 0 MPixel/s compared to 423.6 GPixel/s.
Memory architecture diverges completely. The MI350P uses HBM3e, while the RTX 5090 uses GDDR7. The MI350P’s bus width is 8192 bits versus 512 bits, and its bandwidth is 8.19 TB/s versus 1.79 TB/s. Memory clocks also differ: the MI350P runs at 2000 MHz with 8 Gbps effective, while the RTX 5090 runs at 1750 MHz with 28 Gbps effective. The effective data rate is higher on the RTX 5090, but the MI350P’s wider bus and larger capacity dominate. The MI350P has 144 GB of memory, the RTX 5090 has 32 GB.
Power and physical design differ as well. The MI350P has a 600 W TDP, the RTX 5090 has 575 W. Both use a dual-slot cooler and a single 16-pin power connector. The MI350P suggests a 1000 W PSU, the RTX 5090 suggests 950 W. Dimensions favor the MI350P in length: 267 mm versus 304 mm for the RTX 5090. Height is 111 mm versus 137 mm. Width is identical at 40 mm. The MI350P has no display outputs, while the RTX 5090 has 1x HDMI 2.1b and 3x DisplayPort 2.1b. The RTX 5090 supports DirectX 12 Ultimate, OpenGL 4.6, and Vulkan 1.4; the MI350P reports N/A for all three. The RTX 5090 is listed as Active production status, while the MI350P has no production status recorded. Release dates: the MI350P is dated 2026-05-06, the RTX 5090 is dated 2025-01-29. The RTX 5090’s predecessor is GeForce 40 and successor is GeForce 60; the MI350P’s predecessor is Radeon Instinct with no successor listed.
The Verdict
The data supports a clear split: pick the RTX 5090 for any task that involves graphics, rendering, or standard GPU compute benchmarks, because it is the only one with recorded performance scores. Its 92nd percentile ranking and average benchmark score of 79,842 place it among the top GPUs in the database. The MI350P has no benchmark scores, so any performance claim about it cannot be substantiated from our measurements. Its 50th percentile rank is a placeholder, not a measured result.
Pick the MI350P for workloads that require massive memory capacity and bandwidth. The 144 GB HBM3e pool and 8.19 TB/s bandwidth are unmatched by the RTX 5090’s 32 GB GDDR7 and 1.79 TB/s. If the task involves holding large datasets on the GPU, the MI350P’s memory subsystem is the deciding factor. Its 36.04 TFLOPS FP16 performance, while lower than the RTX 5090’s 104.8 TFLOPS, is still substantial and operates at a 1:1 ratio with FP32, which suits mixed-precision workloads.
The RTX 5090 is the only option with graphics API support. DirectX 12 Ultimate, OpenGL 4.6, and Vulkan 1.4 are present on the RTX 5090, while the MI350P lists N/A for all. The RTX 5090 also has display outputs, enabling direct connection to monitors, which the MI350P cannot do. For any use case requiring visual output or standard graphics workloads, the RTX 5090 is the only viable choice from the recorded data.
The RTX 5090’s higher transistor density (122.9M per mm² versus 61.3M) and higher clock speeds (2407 MHz boost versus 2200 MHz boost) explain its compute lead. The MI350P’s lower density and clocks are offset by its larger memory footprint. The RTX 5090 has 21,760 shading units versus 8,192, and 680 TMUs versus 512, which directly explains its 2.9x FP32 advantage. The MI350P’s 0 ROPs and 0 MPixel/s pixel rate confirm it is not designed for rasterization.
Who should pick which: if the workload involves graphics, gaming, rendering, or any DirectX/Vulkan/OpenGL compute, the RTX 5090 is the answer. If the workload is memory-bound, such as large-scale inference or data processing that fits within 144 GB, the MI350P provides the necessary capacity. The RTX 5090’s launch MSRP is 1,999 USD, which is the only pricing information available.
FAQ
Q: Which GPU has a higher FP32 performance?
A: The RTX 5090 delivers 104.8 TFLOPS FP32, while the MI350P provides 36.04 TFLOPS FP32.
Q: What is the memory capacity difference?
A: The MI350P has 144 GB of HBM3e memory, while the RTX 5090 has 32 GB of GDDR7 memory.
Q: Does either GPU support DirectX?
A: The RTX 5090 supports DirectX 12 Ultimate (12_2), while the MI350P reports N/A for DirectX.
Q: Which GPU has more shading units?
A: The RTX 5090 has 21,760 shading units, compared to 8,192 on the MI350P.
Q: What are the TDP ratings?
A: The MI350P has a 600 W TDP, and the RTX 5090 has a 575 W TDP.
Q: Which GPU has display outputs?
A: The RTX 5090 has 1x HDMI 2.1b and 3x DisplayPort 2.1b, while the MI350P has no display outputs.
Specification Differences
| Feature | AMD Instinct MI350P | NVIDIA GeForce RTX 5090 |
| --- | --- | --- |
| Architecture | CDNA 4.0 | Blackwell 2.0 |
| Process node | 3 nm | 5 nm |
| Transistors | 73,000 million | 92,200 million |
| Die size | 1190 mm² | 750 mm² |
| Transistor density | 61.3M / mm² | 122.9M / mm² |
| Base clock | 1000 MHz | 2017 MHz |
| Boost clock | 2200 MHz | 2407 MHz |
| Memory clock | 2000 MHz 8 Gbps effective | 1750 MHz 28 Gbps effective |
| Memory size | 144 GB | 32 GB |
| Memory type | HBM3e | GDDR7 |
| Memory bus width | 8192 bit | 512 bit |
| Memory bandwidth | 8.19 TB/s | 1.79 TB/s |
| Shading units | 8192 | 21760 |
| TMUs | 512 | 680 |
| ROPs | 0 | 176 |
| RT cores | N/A | 170 |
| Tensor cores | N/A | 680 |
| Pixel rate | 0 MPixel/s | 423.6 GPixel/s |
| Texture rate | 1,126.4 GTexel/s | 1,636.8 GTexel/s |
| FP32 | 36.04 TFLOPS | 104.8 TFLOPS |
| FP16 | 36.04 TFLOPS (1:1) | 104.8 TFLOPS (1:1) |
| TDP | 600 W | 575 W |
| Suggested PSU | 1000 W | 950 W |
| Power connectors | 1x 16-pin | 1x 16-pin |
| Slot width | Dual-slot | Dual-slot |
| Length | 267 mm | 304 mm |
| Height | 111 mm | 137 mm |
| Width | 40 mm | 40 mm |
| Display outputs | No outputs | 1x HDMI 2.1b, 3x DisplayPort 2.1b |
| DirectX | N/A | 12 Ultimate (12_2) |
| OpenGL | N/A | 4.6 |
| Vulkan | N/A | 1.4 |
| Bus interface | PCIe 5.0 x16 | PCIe 5.0 x16 |
| Production status | N/A | Active |
| Release date | 2026-05-06 | 2025-01-29 |
| Predecessor | Radeon Instinct | GeForce 40 |
| Successor | N/A | GeForce 60 |
| Launch MSRP | N/A | 1,999 USD |
| Average benchmark score | 0 | 79,842 |
| Percentile vs all GPUs | 50 | 92 |