AMD Instinct MI350X vs NVIDIA GeForce RTX 5090 SE Comparison
AMD Instinct MI350X
GeForce RTX 5090 SE
Analysis: AMD Instinct MI350X vs NVIDIA GeForce RTX 5090 SE
Head-to-Head Benchmarks
The recorded data shows an unusual situation for these two accelerators: there are no direct head-to-head benchmark results in the database for the AMD Instinct MI350X and the NVIDIA GeForce RTX 5090 SE. Both products have an empty benchmark array, a zero average benchmark score, and a percentile versus all GPUs of 50, which places them at the median of the database distribution. This is not a case of one card beating the other by a measurable margin; rather, the database contains no comparative performance measurements between them. The absence of tested workloads means the raw compute figures from the specification sheets must serve as the primary evidence for differentiation.
The strongest single-number comparison comes from FP32 throughput. The MI350X delivers 72.09 TFLOPS, while the RTX 5090 SE delivers 66.94 TFLOPS. That puts the AMD part ahead by roughly 7.7 percent in raw single-precision floating point. The margin is modest, but it is the clearest quantitative win the database offers for the AMD side. Texture rate shows a wider gap: the MI350X reaches 2,252.8 GTexel/s versus 1,045.9 GTexel/s for the NVIDIA card, a lead that exceeds 2.1 times. The MI350X also has a substantially larger shading unit count at 16,384 versus 14,080, and more texture mapping units at 1,024 versus 440.
The NVIDIA part claims the pixel rate category outright. The RTX 5090 SE achieves 380.3 GPixel/s, while the MI350X is listed at 0 MPixel/s. That zero is not a measurement of poor performance; it reflects the absence of raster output pipelines on the AMD accelerator. The MI350X has zero ROPs, meaning it cannot produce conventional pixel output. The RTX 5090 SE has 160 ROPs and a full graphics pipeline, including 110 ray tracing cores and 440 tensor cores. The MI350X lists no ray tracing cores and no tensor cores in its specification, so those categories go entirely to NVIDIA by default. Clock behavior also favors the NVIDIA card: the RTX 5090 SE boosts to 2,377 MHz versus 2,200 MHz for the MI350X, and its base clock sits at 1,740 MHz versus 1,000 MHz. The AMD part compensates with a much larger die and a higher thermal envelope.
The Verdict
The data indicates these products serve entirely different purposes, and neither can be called a universal winner. The AMD Instinct MI350X is a compute-oriented accelerator with no display outputs, no graphics API support (DirectX, OpenGL, and Vulkan are all listed as N/A), and no ROPs. Its strengths are memory capacity and bandwidth: 288 GB of HBM3e on a 8192-bit bus delivers 8.19 TB/s, which is more than six times the 1.34 TB/s of the RTX 5090 SE. The NVIDIA card holds 24 GB of GDDR7 on a 384-bit bus. Anyone selecting a product for data-center-scale memory-bound workloads would find the MI350X specification far more compelling.
The RTX 5090 SE is the only one of the two that functions as a graphics card. It has display outputs (1x HDMI 2.1b and 3x DisplayPort 2.1b), full DirectX 12 Ultimate, OpenGL 4.6, and Vulkan 1.4 support, plus ray tracing and tensor hardware. Its 500 W TDP and 900 W suggested PSU make it far easier to integrate into a conventional workstation. The MI350X demands a 1000 W TDP and a 1400 W suggested PSU, with an OAM module form factor and no power connectors listed, which implies a server chassis with dedicated power delivery. The launch MSRP of the RTX 5090 SE is 1,499 USD. The MI350X has no launch MSRP in the database. The verdict from the recorded specifications is straightforward: the MI350X is for compute without graphics; the RTX 5090 SE is for graphics with substantial compute.
Architecture Differences
The MI350X uses the CDNA 4.0 architecture on AMD's MI350 256CU chip, built on a 3 nm process at TSMC. The RTX 5090 SE uses the Blackwell 2.0 architecture on NVIDIA's GB202 chip, built on a 5 nm process, also at TSMC. The process node difference is significant: the AMD part uses a denser, more advanced fabrication process. Transistor counts reflect the physical scale of each design. The MI350X contains 185,000 million transistors on a die size of 2,380 mm², which works out to a transistor density of 77.7M / mm². The RTX 5090 SE contains 92,200 million transistors on a 750 mm² die, giving a density of 122.9M / mm². The NVIDIA chip is smaller and denser, while the AMD chip is enormous and slightly less dense per square millimeter.
Memory architecture separates the two fundamentally. The MI350X uses HBM3e memory with a 8192-bit bus width, achieving 8.19 TB/s of bandwidth. The RTX 5090 SE uses GDDR7 with a 384-bit bus, achieving 1.34 TB/s. The AMD memory clock is listed as 2000 MHz with 8 Gbps effective, while the NVIDIA memory clock is 1750 MHz with 28 Gbps effective. The higher effective data rate on NVIDIA's GDDR7 does not compensate for the far wider bus on the AMD side. The MI350X also has no graphics API support, no display outputs, and no ROPs, which confirms its design target as a pure compute accelerator. The RTX 5090 SE has full graphics API support, ray tracing cores, tensor cores, and a conventional dual-slot form factor with a 16-pin power connector.
Specification Differences
The two devices differ in nearly every measurable field. The MI350X has a base clock of 1000 MHz and a boost clock of 2200 MHz; the RTX 5090 SE has a base clock of 1740 MHz and a boost clock of 2377 MHz. Shading units: 16,384 on the MI350X versus 14,080 on the RTX 5090 SE. Texture mapping units: 1,024 versus 440. Raster output pipelines: 0 versus 160. Ray tracing cores: none listed versus 110. Tensor cores: none listed versus 440. Pixel rate: 0 MPixel/s versus 380.3 GPixel/s. Texture rate: 2,252.8 GTexel/s versus 1,045.9 GTexel/s. FP32 and FP16 throughput are identical within each card: 72.09 TFLOPS for the MI350X and 66.94 TFLOPS for the RTX 5090 SE.
Memory size: 288 GB versus 24 GB. Memory type: HBM3e versus GDDR7. Bus width: 8192 bit versus 384 bit. Bandwidth: 8.19 TB/s versus 1.34 TB/s. TDP: 1000 W versus 500 W. Suggested PSU: 1400 W versus 900 W. Slot width: OAM Module versus dual-slot. Power connectors: none versus 1x 16-pin. Display outputs: none versus 1x HDMI 2.1b and 3x DisplayPort 2.1b. Bus interface is identical: PCIe 5.0 x16 for both. Dimensions differ heavily: the MI350X is 102 mm long and 165 mm wide, while the RTX 5090 SE is 267 mm long, 111 mm high, and 40 mm wide. The MI350X has no production status listed; the RTX 5090 SE is marked Active. The MI350X released on 2025-06-11, while the RTX 5090 SE has a release date of 2025-12-31. The MI350X predecessor is Radeon Instinct; the RTX 5090 SE predecessor is GeForce 40. The RTX 5090 SE has a successor listed as GeForce 60; the MI350X has none.
FAQ
Q: Which card has higher FP32 compute performance?
A: The AMD Instinct MI350X delivers 72.09 TFLOPS of FP32, while the NVIDIA GeForce RTX 5090 SE delivers 66.94 TFLOPS. The AMD part has a lead of about 5.15 TFLOPS.
Q: Does the MI350X support graphics rendering?
A: No. The MI350X lists DirectX, OpenGL, and Vulkan as N/A, has no display outputs, and has a pixel rate of 0 MPixel/s with zero ROPs. It is not designed for graphics.
Q: How does memory bandwidth compare between the two?
A: The MI350X provides 8.19 TB/s over an 8192-bit HBM3e interface. The RTX 5090 SE provides 1.34 TB/s over a 384-bit GDDR7 interface. The AMD card has roughly 6.1 times the bandwidth.
Q: What is the TDP difference?
A: The MI350X has a TDP of 1000 W and a suggested PSU of 1400 W. The RTX 5090 SE has a TDP of 500 W and a suggested PSU of 900 W.
Q: Does the RTX 5090 SE have ray tracing cores?
A: Yes, the RTX 5090 SE includes 110 ray tracing cores and 440 tensor cores. The MI350X lists neither.
Q: Which card was released first?
A: The MI350X has a release date of 2025-06-11. The RTX 5090 SE has a release date of 2025-12-31.
Where Each One Wins
The AMD Instinct MI350X wins in memory capacity, memory bandwidth, FP32 throughput, texture rate, shading unit count, and TMU count. Its 288 GB of HBM3e and 8.19 TB/s of bandwidth make it the stronger candidate for large model inference, scientific simulation, and any workload where data movement dominates compute. The 72.09 TFLOPS FP32 figure and the 2,252.8 GTexel/s texture rate indicate a device engineered for sustained mathematical throughput rather than frame output. Its 3 nm process node and 185,000 million transistor count also point to a design that prioritizes scale over efficiency in the conventional graphics sense.
The NVIDIA GeForce RTX 5090 SE wins in clock speed, pixel rate, ray tracing, tensor operations, graphics API support, power efficiency, and physical integration. The boost clock of 2,377 MHz exceeds the MI350X's 2,200 MHz. The 380.3 GPixel/s pixel rate and 160 ROPs confirm it can rasterize images. The 110 ray tracing cores and 440 tensor cores give it dedicated hardware for real-time rendering and AI acceleration within a graphics context. The 500 W TDP and 900 W suggested PSU make it manageable in a standard workstation, and its dual-slot design with a 16-pin connector fits conventional PC power delivery. The display outputs and full DirectX 12 Ultimate, OpenGL 4.6, and Vulkan 1.4 support mean it can drive monitors directly, which the MI350X cannot do at all. The transistor density of 122.9M / mm² on the 750 mm² die also shows a more compact, efficient layout than the MI350X's 77.7M / mm² on 2,380 mm².
The benchmark database does not contain any direct performance scores for either device, so all conclusions rest on the recorded specifications. The split is clear: the MI350X owns raw compute scale, the RTX 5090 SE owns graphics capability and system flexibility.