AMD Instinct MI300X vs NVIDIA GeForce RTX 5090 D V2 Comparison
AMD Instinct MI300X
GeForce RTX 5090 D V2
PERFORMANCE BENCHMARKS
Analysis: AMD Instinct MI300X vs NVIDIA GeForce RTX 5090 D V2
Head-to-Head Benchmarks
The recorded data for these two accelerators comes from different benchmark suites, so a direct apples-to-apples comparison requires careful interpretation. The AMD Instinct MI300X posts a Geekbench OpenCL score of 317,994, while the NVIDIA GeForce RTX 5090 D V2 delivers a 3DMark Steel Nomad DX12 score of 16,504. These metrics target different workloads: OpenCL is a general-purpose compute API, while Steel Nomad DX12 is a graphics-oriented test. The database shows the MI300X sits at the 100th percentile among all GPUs, meaning it outperforms every other recorded device in its benchmark category. The RTX 5090 D V2, by contrast, holds the 59th percentile in its own test, placing it slightly above the median but far from the top.
The MI300X’s nearest rivals in the database provide context for its OpenCL result. It trails the NVIDIA H200 NVL, which scores 334,891, by 5%. It also sits 8% behind the NVIDIA B200, which scores 345,482. Against the NVIDIA L40S, which scores 295,763, the MI300X leads by 7.5%. The NVIDIA RTX 6000 Ada Generation scores 287,237, and the MI300X is 10.7% ahead of that part. These deltas show the MI300X is competitive with, though not the absolute leader of, the highest-performing data center compute accelerators in the database.
The RTX 5090 D V2’s rivals in the Steel Nomad test tell a different story. Its score of 16,504 is statistically tied with the NVIDIA T400, which scores 16,508, a delta of 0%. It is 0.5% ahead of the AMD Radeon PRO W7500 at 16,415, 0.6% ahead of the NVIDIA RTX PRO 6000 Blackwell at 16,408, and 0.9% ahead of the AMD Radeon RX 5700 XT at 16,361. These are extremely narrow margins, indicating that in this specific DirectX 12 graphics workload, the RTX 5090 D V2 does not separate itself from much older or lower-tier hardware. The benchmark results indicate that the RTX 5090 D V2’s strength is not in this particular test, despite its modern architecture.
The raw compute specifications reinforce the divergence. The MI300X delivers 81.72 TFLOPS of FP32 throughput and the same 81.72 TFLOPS for FP16, with a 1:1 ratio. The RTX 5090 D V2 delivers 104.8 TFLOPS for both FP32 and FP16, also at a 1:1 ratio. On paper, the NVIDIA part has 28% more floating-point throughput, but its actual benchmark score sits near the bottom of its peer group, suggesting that raw TFLOPS do not translate directly into the measured Steel Nomad result. The MI300X, despite lower theoretical FP32, achieves a top-percentile OpenCL score, indicating that its memory subsystem and compute scheduling are well matched to that workload.
Memory is another decisive differentiator. The MI300X carries 192 GB of HBM3 across an 8192-bit bus, yielding 5.32 TB/s of bandwidth. The RTX 5090 D V2 has 24 GB of GDDR7 on a 384-bit bus, producing 1.34 TB/s. The MI300X has 8 times the memory capacity and roughly 4 times the bandwidth. For large-scale compute tasks that fit within a single device’s memory, this is a substantial advantage. The RTX 5090 D V2’s memory clock runs at 1750 MHz with 28 Gbps effective, while the MI300X runs at 1300 MHz with 5.2 Gbps effective; the bandwidth gap is due primarily to the vastly wider bus on the AMD part.
FAQ
Q: Which GPU has the higher FP32 compute throughput?
A: The NVIDIA GeForce RTX 5090 D V2 delivers 104.8 TFLOPS of FP32, compared to the AMD Instinct MI300X’s 81.72 TFLOPS.
Q: How much memory bandwidth does each accelerator provide?
A: The MI300X provides 5.32 TB/s over an 8192-bit HBM3 interface. The RTX 5090 D V2 provides 1.34 TB/s over a 384-bit GDDR7 interface.
Q: What are the benchmark percentile rankings for each part?
A: The MI300X ranks at the 100th percentile among all GPUs in its Geekbench OpenCL test. The RTX 5090 D V2 ranks at the 59th percentile in its 3DMark Steel Nomad DX12 test.
Q: How does the MI300X compare to the NVIDIA H200 NVL?
A: The MI300X scores 317,994 in OpenCL, which is 5% lower than the H200 NVL’s 334,891, and 8% lower than the NVIDIA B200’s 345,482.
Q: How close is the RTX 5090 D V2 to its nearest rivals in the Steel Nomad test?
A: The RTX 5090 D V2’s score of 16,504 is within 0.9% of the AMD Radeon RX 5700 XT, and within 0.6% of the NVIDIA RTX PRO 6000 Blackwell, showing very tight clustering.
Q: What is the transistor count and die size for each chip?
A: The MI300X uses 153,000 million transistors on a 1017 mm² die. The RTX 5090 D V2 uses 92,200 million transistors on a 750 mm² die.
Where Each One Wins
The AMD Instinct MI300X wins decisively in memory capacity and bandwidth. With 192 GB of HBM3 and 5.32 TB/s, it can hold large model weights and datasets entirely on-device, avoiding PCIe transfers. Its OpenCL score of 317,994 at the 100th percentile confirms that it is a top-tier compute accelerator for general-purpose parallel workloads. The MI300X also has a higher texture rate at 2,553.6 GTexel/s versus 1,636.8 GTexel/s for the RTX 5090 D V2, and it uses 1216 texture mapping units compared to 680. For tasks that stress compute throughput and memory capacity, the MI300X is the stronger part.
The NVIDIA GeForce RTX 5090 D V2 wins in raw FP32 and FP16 throughput, delivering 104.8 TFLOPS in both, which is 28% higher than the MI300X’s 81.72 TFLOPS. It also has a significantly higher pixel rate at 423.6 GPixel/s, while the MI300X has a pixel rate of 0 MPixel/s due to its lack of ROPs. The RTX 5090 D V2 has 176 ROPs, 170 RT cores, and 680 tensor cores, all of which are absent or null in the MI300X’s specification. The RTX 5090 D V2 supports DirectX 12 Ultimate, OpenGL 4.6, and Vulkan 1.4, while the MI300X reports N/A for all graphics APIs. It also provides display outputs (1x HDMI 2.1b, 3x DisplayPort 2.1b), whereas the MI300X has no display outputs at all. For graphics rendering, ray tracing, and client-side workloads, the RTX 5090 D V2 is the only one of the two that is functionally viable.
The RTX 5090 D V2 also wins on clock speeds. Its base clock is 2017 MHz and boost is 2407 MHz, compared to the MI300X’s 1000 MHz base and 2100 MHz boost. This higher clock rate contributes to its FP32 lead. The RTX 5090 D V2 has a smaller die (750 mm² versus 1017 mm²) and fewer transistors (92,200 million versus 153,000 million), but a higher transistor density at 122.9 million per mm², which is slightly below the MI300X’s 150.4 million per mm². The MI300X has more shading units at 19,456 versus 21,760 for the RTX 5090 D V2, but the NVIDIA part still achieves higher FP32 due to its clocks.
Specification Differences
The two accelerators differ in nearly every measured specification. The MI300X uses HBM3 memory with 192 GB capacity, an 8192-bit bus, and 5.32 TB/s bandwidth. The RTX 5090 D V2 uses GDDR7 with 24 GB capacity, a 384-bit bus, and 1.34 TB/s bandwidth. The MI300X has 0 ROPs and a 0 MPixel/s pixel rate, while the RTX 5090 D V2 has 176 ROPs and a 423.6 GPixel/s pixel rate. The MI300X has 1216 TMUs and a 2,553.6 GTexel/s texture rate; the RTX 5090 D V2 has 680 TMUs and a 1,636.8 GTexel/s texture rate.
The MI300X has no RT cores and no tensor cores listed, while the RTX 5090 D V2 has 170 RT cores and 680 tensor cores. The MI300X has 19,456 shading units versus 21,760 for the RTX 5090 D V2. The MI300X’s TDP is 750 W with a suggested PSU of 1150 W, while the RTX 5090 D V2’s TDP is 575 W with a suggested PSU of 950 W. The MI300X uses an OAM module slot width with no power connectors, while the RTX 5090 D V2 is dual-slot with a 1x 16-pin connector. The MI300X has no display outputs; the RTX 5090 D V2 has 1x HDMI 2.1b and 3x DisplayPort 2.1b. The MI300X reports N/A for DirectX, OpenGL, and Vulkan; the RTX 5090 D V2 supports DirectX 12 Ultimate, OpenGL 4.6, and Vulkan 1.4.
The MI300X’s memory clock is 1300 MHz with 5.2 Gbps effective, while the RTX 5090 D V2’s memory clock is 1750 MHz with 28 Gbps effective. The MI300X’s release date is 2023-12-05; the RTX 5090 D V2’s release date is 2025-08-14. The RTX 5090 D V2 has a launch MSRP of 2,299 USD. The MI300X has no launch MSRP in the database. The MI300X measures 1017 mm² with 153,000 million transistors; the RTX 5090 D V2 measures 750 mm² with 92,200 million transistors. The RTX 5090 D V2 has physical dimensions of 304 mm length, 137 mm height, and 48 mm width. The MI300X has no recorded dimensions.
Architecture Differences
The MI300X uses the Aqua Vanjaram chip built on CDNA 3.0 architecture, fabricated by TSMC on a 5 nm process. The RTX 5090 D V2 uses the GB202 chip built on Blackwell 2.0 architecture, also fabricated by TSMC on a 5 nm process. Both are 5 nm parts from the same foundry, so the process node is identical. The transistor density differs: the MI300X packs 150.4 million transistors per mm², while the RTX 5090 D V2 packs 122.9 million per mm². This means the MI300X uses its 1017 mm² die more efficiently in terms of transistor count, despite having a larger die overall.
The MI300X belongs to AMD’s Instinct (MIx) generation, with a predecessor of Radeon Instinct. The RTX 5090 D V2 belongs to the GeForce 50-series, with a predecessor of GeForce 40 and a successor of GeForce 60. The MI300X has no listed successor. The architectural lineage splits clearly: CDNA 3.0 is AMD’s compute-optimized design, omitting graphics-specific hardware entirely. The MI300X has no ROPs, no RT cores, no tensor cores, and no display outputs, and it reports N/A for all graphics APIs. This is a pure compute accelerator designed for data center workloads.
Blackwell 2.0 on the RTX 5090 D V2 retains full graphics functionality. It includes 170 RT cores and 680 tensor cores, which are dedicated to ray tracing and AI acceleration respectively. It also supports DirectX 12 Ultimate, OpenGL 4.6, and Vulkan 1.4, making it a complete client GPU. The MI300X has 19,456 shading units, which are more than the RTX 5090 D V2’s 21,760 shading units is incorrect per the data; the RTX 5090 D V2 has 21,760 shading units, which is higher. The MI300X has 1216 TMUs versus 680 for the RTX 5090 D V2, giving the AMD part a higher texture rate despite lower clocks.
The memory architecture is fundamentally different. HBM3 on the MI300X is stacked and placed close to the compute die, enabling an 8192-bit bus. GDDR7 on the RTX 5090 D V2 is discrete memory on a 384-bit bus. The MI300X’s 5.32 TB/s bandwidth is nearly 4 times higher than the RTX 5090 D V2’s 1.34 TB/s. The MI300X also has a 1000 MHz base clock and 2100 MHz boost, while the RTX 5090 D V2 has a 2017 MHz base and 2407 MHz boost. The power delivery differs: the MI300X is an OAM module with no power connectors and a 750 W TDP, while the RTX 5090 D V2 is a dual-slot card with a 1x 16-pin connector and a 575 W TDP.
The Verdict
The data shows two accelerators built for different purposes. The AMD Instinct MI300X is a compute-only device with 192 GB of HBM3, 5.32 TB/s bandwidth, and a 100th percentile OpenCL score of 317,994. It is 7.5% ahead of the NVIDIA L40S and 10.7% ahead of the NVIDIA RTX 6000 Ada Generation, though it trails the H200 NVL by 5% and the B200 by 8%. For workloads that require large memory capacity and high bandwidth, such as large language model inference or scientific simulation, the MI300X is the clear choice from these two parts.
The NVIDIA GeForce RTX 5090 D V2 is a complete graphics card with display outputs, ray tracing, tensor cores, and full API support. Its Steel Nomad score of 16,504 places it at the 59th percentile, and its closest rivals are within 0.9%, indicating that this particular benchmark does not showcase its strengths. Its FP32 throughput of 104.8 TFLOPS exceeds the MI300X’s 81.72 TFLOPS, and its higher clocks (2407 MHz boost versus 2100 MHz) contribute to that lead. For interactive rendering, ray-traced workloads, or any task requiring a display output, the RTX 5090 D V2 is the only viable option.
Neither part is a substitute for the other. The MI300X cannot output video, has no graphics API support, and lacks rasterization hardware. The RTX 5090 D V2 cannot match the MI300X’s memory capacity or bandwidth, and its benchmark percentile is far lower. The database records the MI300X at the 100th percentile and the RTX 5090 D V2 at the 59th, but these are measured on different tests. The correct choice depends entirely on whether the workload is compute-centric or graphics-centric. For data center compute, the MI300X’s specifications and benchmark results dominate. For client-side graphics and general-purpose GPU tasks, the RTX 5090 D V2 is the functional option.