AMD Instinct MI300X vs NVIDIA GeForce RTX 5090 Comparison

AMD
RADEON

AMD Instinct MI300X

CORE STATE Aqua Vanjaram
VRAM 192 GB
CLOCK SPEED 2100 MHz
TDP 750 W
BUS WIDTH 8192 bit
ARCHITECTURE CDNA 3.0
nm
PROCESS 5 nm
LAUNCH DATE 2023
VS
NVIDIA
GEFORCE

GeForce RTX 5090

CORE STATE GB202
VRAM 32 GB
CLOCK SPEED 2407 MHz
TDP 575 W
BUS WIDTH 512 bit
ARCHITECTURE Blackwell 2.0
nm
PROCESS 5 nm
LAUNCH DATE 2025

PERFORMANCE BENCHMARKS

geekbench_opencl
317,994
334,370
3dmark_3dmark_steel_nomad_dx12
N/A
18,355
geekbench_vulkan
N/A
376,728
passmark_directx_10
N/A
226
passmark_directx_11
N/A
341
passmark_directx_12
N/A
185
passmark_directx_9
N/A
395
passmark_g2d
N/A
1,413
passmark_g3d
N/A
39,650
passmark_gpu_compute
N/A
26,756

Analysis: AMD Instinct MI300X vs NVIDIA GeForce RTX 5090

Where Each One Wins

The benchmark data splits these two accelerators along predictable but dramatic lines. The AMD Instinct MI300X is a compute-first accelerator designed for scale-up AI and HPC workloads, and its single recorded benchmark reflects that positioning. The NVIDIA GeForce RTX 5090 is a client graphics card with a full feature set, and its benchmark suite covers both compute and traditional graphics workloads.

In the single directly comparable compute test, Geekbench OpenCL, the RTX 5090 wins outright. The NVIDIA part scores 334370 against the MI300X’s 317994, a delta of 4.9 percent. This is the only head-to-head measurement in the database, so on raw compute throughput as measured by that specific workload, the RTX 5090 holds the advantage.

However, the MI300X’s positioning becomes clear when examining its percentile ranking. The database places the Instinct MI300X in the 100th percentile against all GPUs, meaning it sits at the absolute top of the recorded performance distribution. The RTX 5090, despite its Geekbench OpenCL win, ranks in the 92nd percentile. This discrepancy suggests that the MI300X’s single benchmark score represents a workload class where it dominates far more decisively than the Geekbench result implies, while the RTX 5090’s average benchmark score is dragged down by its many graphics-oriented tests.

The RTX 5090 wins in every category where the database has recorded a test. Beyond Geekbench OpenCL, it posts strong results in Vulkan compute (376728), DirectX 12 (185 in Passmark), DirectX 11 (341), DirectX 10 (226), DirectX 9 (395), and G3D (39650). Its Passmark GPU Compute score is 26756. None of these tests exist for the MI300X, which makes sense given that the Instinct card has no display outputs and no graphics API support. The MI300X is not a graphics card; it is an accelerator with a different job.

The use-case split is therefore stark: if the workload is graphics, rendering, or client-side compute, the RTX 5090 is the only option with recorded data. If the workload is large-scale HPC or AI inference where 192 GB of memory matters, the MI300X is the part the database ranks at the 100th percentile, even if the one shared benchmark does not favor it.

Architecture Differences

The two chips come from different architectural lineages. The MI300X uses AMD’s CDNA 3.0 architecture, specifically the Aqua Vanjaram chip, built on a 5 nm TSMC process. The RTX 5090 uses NVIDIA’s Blackwell 2.0 architecture, implemented as the GB202 chip, also on a 5 nm TSMC process. Both are leading-edge parts, but their design philosophies diverge sharply.

Transistor counts and die sizes tell the first part of the story. The MI300X packs 153,000 million transistors onto a 1017 mm² die, yielding a density of 150.4 million transistors per square millimeter. The RTX 5090 has 92,200 million transistors on a 750 mm² die, for a density of 122.9 million per square millimeter. The MI300X is the larger and denser chip, but its transistor budget is spent differently.

Memory architecture is where the gap becomes enormous. The MI300X carries 192 GB of HBM3 across an 8192-bit bus, delivering 5.32 TB/s of bandwidth. The RTX 5090 has 32 GB of GDDR7 on a 512-bit bus, providing 1.79 TB/s. The MI300X has six times the memory capacity and nearly three times the bandwidth. This is not a minor spec difference; it defines the two products’ roles. The MI300X is built to hold massive models and datasets in memory, while the RTX 5090 is built for latency-sensitive client workloads with far smaller memory footprints.

Compute unit configurations differ as well. The MI300X has 19456 shading units and 1216 texture mapping units, but zero ROPs. Its pixel rate is recorded as 0 MPixel/s because it has no raster output stage. The RTX 5090 has 21760 shading units, 680 TMUs, and 176 ROPs, giving it a pixel rate of 423.6 GPixel/s. The MI300X has more shading units and nearly double the TMUs, but it cannot rasterize at all. The RTX 5090 also includes 170 RT cores and 680 tensor cores, while the MI300X has no recorded RT or tensor core counts in its CDNA 3.0 configuration.

Clock behavior also differs. The MI300X runs at a 1000 MHz base and 2100 MHz boost. The RTX 5090 runs at 2017 MHz base and 2407 MHz boost. The NVIDIA part clocks significantly higher, which contributes to its FP32 advantage despite having only slightly more shading units. The MI300X’s FP32 throughput is 81.72 TFLOPS, while the RTX 5090 reaches 104.8 TFLOPS. Both parts have a 1:1 FP16 to FP32 ratio, meaning their FP16 rates match their FP32 rates.

Power and physical design reflect their intended environments. The MI300X has a 750 W TDP, a suggested PSU of 1150 W, no power connectors of its own (it is an OAM module), and no display outputs. The RTX 5090 has a 575 W TDP, a suggested PSU of 950 W, a single 16-pin connector, and a dual-slot form factor with full display outputs (1x HDMI 2.1b, 3x DisplayPort 2.1b). The MI300X is a data center module; the RTX 5090 is a consumer card that fits in a desktop case at 304 mm length, 137 mm height, and 40 mm width.

Head-to-Head Benchmarks

The database contains exactly one benchmark where both parts have recorded scores: Geekbench OpenCL. The RTX 5090 scores 334370, and the MI300X scores 317994. The delta is 4.9 percent in favor of NVIDIA. This is a meaningful but not overwhelming margin in a compute workload that tends to favor higher clock speeds and general-purpose shader throughput.

Context matters here. The MI300X’s nearest rivals in the database are all NVIDIA data center parts: the B200 at 345482 (8 percent ahead), the H200 NVL at 334891 (5 percent ahead), the L40S at 295763 (7.5 percent behind), and the RTX 6000 Ada Generation at 287237 (10.7 percent behind). The MI300X sits in the middle of that group. Its score is lower than the flagship B200 and H200, but higher than the L40S and RTX 6000 Ada. The RTX 5090’s nearest rivals, by contrast, are older or lower-tier parts: the Tesla P100 PCIe 16 GB (0.3 percent behind), the Tesla P100 PCIe 12 GB (0.6 percent behind), the Radeon RX 6850M XT (1.1 percent behind), and the Radeon Pro Vega 64X (1.4 percent ahead). The RTX 5090’s average benchmark score of 79842 reflects all its tests, not just the OpenCL one.

The RTX 5090’s Vulkan score of 376728 is notably higher than its OpenCL score, suggesting that its compute throughput improves under Vulkan’s API. The MI300X has no Vulkan score recorded, so no comparison is possible. The RTX 5090 also shows strong Passmark G3D performance at 39650, but its DirectX scores are low (185 in DX12, 341 in DX11, 226 in DX10, 395 in DX9), which is typical for a modern card running legacy DX workloads. The G2D score of 1413 and GPU Compute score of 26756 round out its profile.

The single head-to-head result tells us that in a general-purpose compute benchmark, the RTX 5090 edges out the MI300X. But the gap is small, and the MI300X’s 100th percentile ranking across all GPUs suggests that its intended workloads, likely those that exploit its 192 GB memory and 5.32 TB/s bandwidth, are not represented in this one comparison.

Specification Differences

The following fields differ between the two parts:

  • Shading units: MI300X has 19456, RTX 5090 has 21760.
  • Texture mapping units: MI300X has 1216, RTX 5090 has 680.
  • ROPs: MI300X has 0, RTX 5090 has 176.
  • RT cores: MI300X has none recorded, RTX 5090 has 170.
  • Tensor cores: MI300X has none recorded, RTX 5090 has 680.
  • Pixel rate: MI300X is 0 MPixel/s, RTX 5090 is 423.6 GPixel/s.
  • Texture rate: MI300X is 2,553.6 GTexel/s, RTX 5090 is 1,636.8 GTexel/s.
  • FP32 throughput: MI300X is 81.72 TFLOPS, RTX 5090 is 104.8 TFLOPS.
  • Base clock: MI300X is 1000 MHz, RTX 5090 is 2017 MHz.
  • Boost clock: MI300X is 2100 MHz, RTX 5090 is 2407 MHz.
  • Memory clock: MI300X is 1300 MHz (5.2 Gbps effective), RTX 5090 is 1750 MHz (28 Gbps effective).
  • Memory size: MI300X is 192 GB, RTX 5090 is 32 GB.
  • Memory type: MI300X is HBM3, RTX 5090 is GDDR7.
  • Memory bus width: MI300X is 8192 bit, RTX 5090 is 512 bit.
  • Memory bandwidth: MI300X is 5.32 TB/s, RTX 5090 is 1.79 TB/s.
  • Transistors: MI300X is 153,000 million, RTX 5090 is 92,200 million.
  • Die size: MI300X is 1017 mm², RTX 5090 is 750 mm².
  • Transistor density: MI300X is 150.4M / mm², RTX 5090 is 122.9M / mm².
  • TDP: MI300X is 750 W, RTX 5090 is 575 W.
  • Suggested PSU: MI300X is 1150 W, RTX 5090 is 950 W.
  • Slot width: MI300X is OAM Module, RTX 5090 is Dual-slot.
  • Power connectors: MI300X has none, RTX 5090 has 1x 16-pin.
  • Display outputs: MI300X has none, RTX 5090 has 1x HDMI 2.1b, 3x DisplayPort 2.1b.
  • API support: MI300X has N/A for DirectX, OpenGL, and Vulkan; RTX 5090 supports DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4.
  • Chip name: MI300X is Aqua Vanjaram, RTX 5090 is GB202.
  • Architecture: MI300X is CDNA 3.0, RTX 5090 is Blackwell 2.0.
  • Generation: MI300X is Instinct (MIx), RTX 5090 is GeForce 50.
  • Release date: MI300X is 2023-12-05, RTX 5090 is 2025-01-29.
  • Predecessor: MI300X is Radeon Instinct, RTX 5090 is GeForce 40.
  • Successor: MI300X has none recorded, RTX 5090 is GeForce 60.
  • Production status: MI300X has none recorded, RTX 5090 is Active.
  • Dimensions: MI300X has none recorded, RTX 5090 is 304 mm x 137 mm x 40 mm.
  • Launch MSRP: MI300X has none recorded, RTX 5090 is 1,999 USD.

FAQ

Q: Which card is faster in Geekbench OpenCL?

A: The RTX 5090 scores 334370 against the MI300X’s 317994, a 4.9 percent advantage for NVIDIA.

Q: Does the MI300X support graphics output?

A: No. The MI300X has no display outputs, no ROPs, a pixel rate of 0 MPixel/s, and N/A ratings for DirectX, OpenGL, and Vulkan. It is an OAM module, not a client graphics card.

Q: How do the memory capacities compare?

A: The MI300X has 192 GB of HBM3 on an 8192-bit bus with 5.32 TB/s bandwidth. The RTX 5090 has 32 GB of GDDR7 on a 512-bit bus with 1.79 TB/s bandwidth. The MI300X offers 6 times the capacity and roughly 3 times the bandwidth.

Q: What is the percentile ranking for each card?

A: The MI300X is in the 100th percentile against all GPUs, while the RTX 5090 is in the 92nd percentile. The MI300X’s average benchmark score is 317994, and the RTX 5090’s is 79842.

Q: Which card has higher FP32 compute throughput?

A: The RTX 5090 reaches 104.8 TFLOPS, while the MI300X delivers 81.72 TFLOPS. Both have a 1:1 FP16 to FP32 ratio.

Q: What are the power requirements?

A: The MI300X has a 750 W TDP and suggests a 1150 W PSU. The RTX 5090 has a 575 W TDP and suggests a 950 W PSU. The MI300X has no power connectors as an OAM module, while the RTX 5090 uses a single 16-pin connector.

DETAILED SPECIFICATIONS

SPECIFICATION
Instinct MI300X
RTX 5090
Core Specs
Shading Units
19,456
21,760 +11.8%
Shaders
19,456
21,760 +11.8%
TMUs
1,216
680 -44.1%
ROPs
0
176 +∞%
Compute Units
304
—
SM Count
—
170
Clocks
Base Clock
1000 MHz
2017 MHz
Boost Clock
2100 MHz
2407 MHz
Memory Clock
1300 MHz 5.2 Gbps effective
1750 MHz 28 Gbps effective
Memory
Memory Size
192 GB
32 GB
VRAM (MB)
196,608
32,768 -83.3%
Memory Type
HBM3
GDDR7
Memory Bus
8192 bit
512 bit
Bandwidth
5.32 TB/s
1.79 TB/s
Cache
L1 Cache
16 KB (per CU)
128 KB (per SM)
L2 Cache
16 MB
96 MB
L3 Cache
256 MB
—
Performance
Pixel Rate
0 MPixel/s
423.6 GPixel/s
Texture Rate
2,553.6 GTexel/s
1,636.8 GTexel/s
FP32 (TFLOPS)
81.72 TFLOPS
104.8 TFLOPS
FP64 (TFLOPS)
40.86 TFLOPS (1:2)
1.637 TFLOPS (1:64)
FP16 (TFLOPS)
81.72 TFLOPS (1:1)
104.8 TFLOPS (1:1)
AI/RT
RT Cores
—
170
Tensor Cores
—
680
Matrix Cores
1,216
—
Power
TDP
750 W
575 W
TDP (W)
750
575 -23.3%
Suggested PSU
1150 W
950 W
Power Connectors
None
1x 16-pin
Architecture
Architecture
CDNA 3.0
Blackwell 2.0
GPU Name
Aqua Vanjaram
GB202
Generation
Instinct (MIx)
GeForce 50
Process Size
5 nm
5 nm
Transistors
153,000 million
92,200 million
Die Size
1017 mm²
750 mm²
Foundry
TSMC
TSMC
Density
150.4M / mm²
122.9M / mm²
AMD MCM
MCM
2
—
API Support
DirectX
—
12 Ultimate (12_2)
OpenGL
—
4.6
Vulkan
—
1.4
OpenCL
3.0
3.0
CUDA
—
12.0
Shader Model
—
6.9
Physical
Slot Width
OAM Module
Dual-slot
Length
—
304 mm 12 inches
Height
—
137 mm 5.4 inches
Outputs
No outputs
1x HDMI 2.1b3x DisplayPort 2.1b
Bus Interface
PCIe 5.0 x16
PCIe 5.0 x16
Other
Launch Price
—
1,999 USD
Production
—
Active
Predecessor
Radeon Instinct
GeForce 40
Successor
—
GeForce 60
View Instinct MI300X Details View GeForce RTX 5090 Details