AMD Instinct MI355X vs NVIDIA GeForce RTX 5090 Comparison
AMD Instinct MI355X
GeForce RTX 5090
PERFORMANCE BENCHMARKS
Analysis: AMD Instinct MI355X vs NVIDIA GeForce RTX 5090
AMD Instinct MI355X and NVIDIA GeForce RTX 5090 represent two distinct extremes in the current GPU landscape. The MI355X is a compute-oriented accelerator with a massive memory pool, while the RTX 5090 is a high-performance graphics card with a full feature set for rendering and general compute. The recorded data shows that the RTX 5090 holds a 92nd percentile ranking among all GPUs, whereas the MI355X sits at the 50th percentile, though this comparison is complicated by the fact that the MI355X has no recorded benchmark scores in the database. The RTX 5090’s average benchmark score of 79,842 places it marginally ahead of the nearest rival, the NVIDIA Tesla P100 PCIe 16 GB, by 0.3 percent.
Where Each One Wins
The RTX 5090 wins decisively in every measurable benchmark category, as it is the only one of the two with recorded performance data. Its 3DMark Steel Nomad DX12 score of 18,355 demonstrates strong graphics rendering capability, while its Geekbench OpenCL score of 334,370 and Vulkan score of 376,728 show substantial compute throughput in cross-platform APIs. PassMark results further confirm its versatility: a G3D score of 39,650, a GPU Compute score of 26,756, and a G2D score of 1,413. These figures position the RTX 5090 as a top-tier performer across synthetic workloads, with the nearest rival in the database, the Tesla P100 PCIe 16 GB, trailing by only 0.3 percent in average score.
The MI355X, by contrast, has no recorded benchmarks, so its wins cannot be quantified from the database. Its specifications, however, indicate a design goal entirely different from the RTX 5090. The MI355X uses 288 GB of HBM3e memory with 8.19 TB/s of bandwidth, which is 8 times the capacity and roughly 4.6 times the bandwidth of the RTX 5090’s 32 GB GDDR7 at 1.79 TB/s. For workloads that exceed the RTX 5090’s memory capacity, such as large language model inference or massive scientific simulations, the MI355X would hold the advantage, but this is not reflected in any recorded benchmark score. The RTX 5090 wins in all direct comparisons available, while the MI355X’s strengths remain theoretical in the database.
Architecture Differences
The two GPUs employ fundamentally different architectures. The MI355X uses AMD’s CDNA 4.0 architecture on a 3 nm process from TSMC, packing 185,000 million transistors into a die size of 2,380 mm², yielding a transistor density of 77.7 million per mm². The RTX 5090 uses NVIDIA’s Blackwell 2.0 architecture on a 5 nm process, also from TSMC, with 92,200 million transistors on a 750 mm² die, resulting in a higher density of 122.9 million per mm². The MI355X’s die is over three times larger, but the RTX 5090’s smaller node geometry allows for a tighter packing of transistors.
The MI355X’s chip is designated MI350 256CU, indicating a compute unit count that aligns with its 16,384 shading units. It has 1,024 texture mapping units but no ROPs, pixel rate, or display outputs, confirming it is not designed for rasterized graphics output. Its memory is HBM3e with an 8,192-bit bus, which is 16 times wider than the RTX 5090’s 512-bit bus. The RTX 5090, built on the GB202 chip, has 21,760 shading units, 680 TMUs, 176 ROPs, 170 RT cores, and 680 tensor cores. It includes full graphics APIs: DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4, while the MI355X lists N/A for all graphics APIs.
Clock speeds differ significantly. The MI355X has a base clock of 1,000 MHz and a boost clock of 2,400 MHz, while the RTX 5090 starts at 2,017 MHz and boosts to 2,407 MHz. The MI355X’s memory runs at 2,000 MHz (8 Gbps effective), whereas the RTX 5090’s memory operates at 1,750 MHz (28 Gbps effective). The MI355X delivers 78.64 TFLOPS for both FP32 and FP16 (1:1), while the RTX 5090 achieves 104.8 TFLOPS for both, giving NVIDIA a 33 percent higher peak floating-point throughput. Texture rate favors the MI355X at 2,457.6 GTexel/s versus 1,636.8 GTexel/s for the RTX 5090, but the RTX 5090 has a pixel rate of 423.6 GPixel/s while the MI355X is at zero.
FAQ
Q: Which GPU has the higher FP32 compute throughput?
A: The RTX 5090 delivers 104.8 TFLOPS for FP32, which is 33 percent higher than the MI355X’s 78.64 TFLOPS.
Q: How do the memory capacities compare?
A: The MI355X has 288 GB of HBM3e memory, while the RTX 5090 has 32 GB of GDDR7. The MI355X’s capacity is 9 times larger.
Q: What is the memory bandwidth difference?
A: The MI355X provides 8.19 TB/s of bandwidth, while the RTX 5090 provides 1.79 TB/s, making the MI355X approximately 4.6 times faster in bandwidth.
Q: Does the MI355X support any display outputs?
A: No, the MI355X lists no display outputs, while the RTX 5090 offers 1x HDMI 2.1b and 3x DisplayPort 2.1b.
Q: Which GPU has a higher transistor density?
A: The RTX 5090 has a transistor density of 122.9 million per mm², compared to the MI355X’s 77.7 million per mm², despite the MI355X having more total transistors.
Q: What is the power consumption difference?
A: The MI355X has a TDP of 1,400 W with a suggested PSU of 1,800 W, whereas the RTX 5090 has a TDP of 575 W with a suggested PSU of 950 W.
Specification Differences
The two cards differ across nearly every core specification. The MI355X has 16,384 shading units, while the RTX 5090 has 21,760, a 33 percent advantage for NVIDIA. Texture mapping units favor the MI355X at 1,024 versus 680, but ROPs exist only on the RTX 5090 with 176 units. The RTX 5090 includes 170 RT cores and 680 tensor cores, both absent from the MI355X. Clock speeds show a higher base for the RTX 5090 at 2,017 MHz versus 1,000 MHz, but the boost clocks are close at 2,407 MHz and 2,400 MHz respectively.
Memory configurations are starkly different. The MI355X uses 288 GB of HBM3e on an 8,192-bit bus, while the RTX 5090 uses 32 GB of GDDR7 on a 512-bit bus. Bandwidth is 8.19 TB/s versus 1.79 TB/s. The MI355X has no ROPs and a pixel rate of zero, while the RTX 5090 achieves 423.6 GPixel/s. Texture rate is higher on the MI355X at 2,457.6 GTexel/s, but the RTX 5090 wins in FP32 and FP16 throughput. The MI355X is an OAM module with no power connectors, while the RTX 5090 is a dual-slot card with a single 16-pin connector. The MI355X measures 102 mm in length and 165 mm in width, while the RTX 5090 is 304 mm long, 137 mm high, and 40 mm wide.
Head-to-Head Benchmarks
There are no direct head-to-head benchmark results recorded between the MI355X and the RTX 5090. The database lists zero wins for each side in this comparison. The RTX 5090’s own benchmark suite, however, provides a reference point for its performance tier. Its 3DMark Steel Nomad DX12 score of 18,355 is a strong result for modern graphics workloads. The Geekbench Vulkan score of 376,728 exceeds the OpenCL score of 334,370 by 12.7 percent, indicating better optimization for Vulkan in the recorded tests. PassMark results show a G3D score of 39,650 and a GPU Compute score of 26,756, with the compute score trailing the graphics score by 32.5 percent.
The RTX 5090’s nearest rival in the database, the NVIDIA Tesla P100 PCIe 16 GB, has an average score of 79,605, which is 0.3 percent lower than the RTX 5090’s 79,842. The Tesla P100 PCIe 12 GB is 0.6 percent behind, while the AMD Radeon RX 6850M XT is 1.1 percent behind. On the other side, the AMD Radeon Pro Vega 64X scores 80,959, which is 1.4 percent higher than the RTX 5090. These deltas place the RTX 5090 in a tight cluster of high-end GPUs, none of which approach the MI355X’s memory capacity. The MI355X’s 78.64 TFLOPS FP32 is lower than the RTX 5090’s 104.8 TFLOPS, so even without direct tests, the compute throughput data favors NVIDIA in raw arithmetic operations. The MI355X’s advantage lies solely in memory capacity and bandwidth, which are not captured in the benchmark suite available for the RTX 5090.