AMD Instinct MI355X vs NVIDIA GeForce RTX 5090 SE Comparison
AMD Instinct MI355X
GeForce RTX 5090 SE
Analysis: AMD Instinct MI355X vs NVIDIA GeForce RTX 5090 SE
Head-to-Head Benchmarks
The recorded data shows both accelerators in a 50th percentile position against all GPUs in the database, with an average benchmark score of zero for each. This means neither part has completed standardized testing runs in the current dataset, so direct performance comparisons must be derived from their architectural specifications and measured peak rates rather than application-level scores. The database contains no head-to-head benchmark entries, no wins for either side, and no rival comparison data, so the analysis below relies entirely on the recorded peak throughput, memory subsystem, and feature-set figures.
The most decisive specification advantage belongs to the AMD Instinct MI355X in memory capacity and bandwidth. It delivers 288 GB of HBM3e across an 8192-bit bus, producing 8.19 TB/s of memory bandwidth. The NVIDIA GeForce RTX 5090 SE offers 24 GB of GDDR7 on a 384-bit bus, reaching 1.34 TB/s. The AMD part holds a 6.11x bandwidth advantage (8.19 divided by 1.34) and a 12x capacity advantage (288 divided by 24). These are not minor margins; they place the MI355X in a different memory class entirely, one suited for workloads where resident datasets exceed the entire frame buffer of the RTX 5090 SE.
In raw compute throughput, the AMD Instinct MI355X also leads. Its FP32 peak is 78.64 TFLOPS, and its FP16 peak is identical at 78.64 TFLOPS (1:1 ratio). The NVIDIA GeForce RTX 5090 SE reaches 66.94 TFLOPS in both FP32 and FP16 (also 1:1). The AMD part is 17.5% ahead in FP32 (78.64 divided by 66.94 equals 1.175) and equally ahead in FP16. Texture rate follows the same pattern: the MI355X outputs 2,457.6 GTexel/s versus 1,045.9 GTexel/s for the RTX 5090 SE, a 2.35x lead for AMD. However, the pixel rate reverses this trend. The RTX 5090 SE records 380.3 GPixel/s, while the MI355X records 0 MPixel/s, which indicates the AMD part has no raster output stage, a fundamental difference in rendering capability.
Clock behavior differs substantially. The NVIDIA part has a much higher base clock at 1740 MHz versus 1000 MHz, and its boost clock of 2377 MHz slightly exceeds the AMD boost of 2400 MHz. The AMD part relies on a lower base frequency but achieves a similar boost ceiling, suggesting a wider architecture with more parallel units operating at moderate clocks. The RTX 5090 SE uses 14080 shading units, 440 texture mapping units, 160 ROPs, 110 RT cores, and 440 tensor cores. The MI355X specifies 16384 shading units, 1024 TMUs, zero ROPs, and no RT or tensor core counts in the record.
Where Each One Wins
The AMD Instinct MI355X wins in every category tied to memory-bound and compute-bound workloads. Its 288 GB HBM3e frame buffer allows it to hold large neural network weights, massive simulation grids, or extensive scientific datasets without spilling to slower storage. The 8.19 TB/s bandwidth means those large datasets can be streamed into the compute units at rates the NVIDIA part cannot match. The FP32 and FP16 throughput of 78.64 TFLOPS gives it an edge in general-purpose matrix math, and the 2,457.6 GTexel/s texture rate supports high-volume texture sampling, though the zero pixel rate confirms it cannot perform traditional rasterization.
The NVIDIA GeForce RTX 5090 SE wins in rendering and graphics-specific tasks. Its 380.3 GPixel/s pixel rate, 160 ROPs, and 110 RT cores provide hardware support for ray tracing and pixel output that the MI355X entirely lacks. The RTX 5090 SE also supports DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4, while the MI355X has no recorded API support at all. The NVIDIA part includes display outputs (1x HDMI 2.1b and 3x DisplayPort 2.1b), making it functional as a graphics card for interactive use, whereas the MI355X has no display outputs and is an OAM module, meaning it cannot drive a monitor.
Power consumption and physical form factor further separate their intended uses. The RTX 5090 SE has a TDP of 500 W, a dual-slot design, a 267 mm length, and a single 16-pin power connector. The MI355X consumes 1400 W, occupies an OAM module slot, measures 102 mm by 165 mm, and has no power connectors of its own, relying on the host system's power delivery. The database suggests a 900 W PSU for the NVIDIA part and an 1800 W PSU for the AMD part, confirming the MI355X is designed for server racks with dedicated power infrastructure, not desktop systems.
Architecture Differences
The AMD Instinct MI355X uses the CDNA 4.0 architecture on a 3 nm TSMC process, packing 185,000 million transistors into a 2380 mm² die. The transistor density is 77.7 million per square millimeter. The NVIDIA GeForce RTX 5090 SE uses the Blackwell 2.0 architecture on a 5 nm TSMC process, with 92,200 million transistors on a 750 mm² die, yielding a higher density of 122.9 million per square millimeter. The AMD chip is 3.17x larger in die area (2380 divided by 750) and has 2.01x more transistors (185,000 divided by 92,200), but the NVIDIA chip achieves better packing density due to its smaller feature size.
The memory architectures are fundamentally different. The MI355X uses HBM3e, a stacked high-bandwidth memory design, with a 8192-bit bus. The RTX 5090 SE uses GDDR7, a discrete memory type, with a 384-bit bus. The bus width difference is 21.33x in favor of AMD (8192 divided by 384), which explains the massive bandwidth gap despite the NVIDIA memory running at a higher effective clock (28 Gbps versus 8 Gbps). The MI355X memory clock is 2000 MHz with 8 Gbps effective, while the RTX 5090 SE memory clock is 1750 MHz with 28 Gbps effective. The GDDR7's faster signaling per pin cannot compensate for the HBM3e's much wider interface.
The compute unit composition differs sharply. The MI355X has 16384 shading units and 1024 TMUs, with zero ROPs. The RTX 5090 SE has 14080 shading units, 440 TMUs, and 160 ROPs. The shading unit count favors AMD by 16.4% (16384 divided by 14080), and the TMU count favors AMD by 2.33x (1024 divided by 440). The NVIDIA part adds dedicated RT cores (110) and tensor cores (440) for ray tracing and AI acceleration, while the AMD part has no recorded equivalents. This indicates the MI355X relies on its general-purpose shader array for all compute, whereas the RTX 5090 SE has specialized hardware for graphics and inference.
The feature sets diverge entirely. The MI355X has no display outputs, no DirectX, OpenGL, or Vulkan support, and no rasterization capability (0 MPixel/s). The RTX 5090 SE has full graphics API support, multiple display outputs, and a pixel rate of 380.3 GPixel/s. The MI355X is a compute-only accelerator, while the RTX 5090 SE is a complete graphics processor. The RTX 5090 SE also has a release date of 2025-12-31 and an active production status, while the MI355X has a release date of 2025-06-11 and no recorded production status. The NVIDIA part has a listed predecessor (GeForce 40) and successor (GeForce 60), while the AMD part lists only a predecessor (Radeon Instinct).
The power delivery differs by design. The MI355X has no power connectors because it is an OAM module, receiving power through the baseboard. The RTX 5090 SE uses a single 16-pin connector. The TDP gap is 900 W (1400 minus 500), and the suggested PSU gap is 900 W (1800 minus 900). The NVIDIA part is a dual-slot card at 267 mm long, 111 mm high, and 40 mm wide. The AMD part is a compact 102 mm by 165 mm module, fitting into a different physical ecosystem.
FAQ
Q: Which accelerator has more memory bandwidth?
A: The AMD Instinct MI355X has 8.19 TB/s from HBM3e on an 8192-bit bus, while the NVIDIA GeForce RTX 5090 SE has 1.34 TB/s from GDDR7 on a 384-bit bus.
Q: Can the AMD Instinct MI355X render graphics?
A: No. It has a pixel rate of 0 MPixel/s, zero ROPs, and no display outputs, making it a compute-only accelerator.
Q: What is the FP32 performance difference?
A: The AMD Instinct MI355X delivers 78.64 TFLOPS FP32, while the NVIDIA GeForce RTX 5090 SE delivers 66.94 TFLOPS FP32, a 17.5% advantage for AMD.
Q: Does the NVIDIA card support ray tracing?
A: Yes. The RTX 5090 SE has 110 RT cores and 440 tensor cores, plus a 380.3 GPixel/s pixel rate, indicating full rasterization and ray tracing capability.
Q: Which part consumes more power?
A: The AMD Instinct MI355X has a 1400 W TDP and suggests an 1800 W PSU. The NVIDIA GeForce RTX 5090 SE has a 500 W TDP and suggests a 900 W PSU.
Q: What is the memory capacity difference?
A: The AMD Instinct MI355X has 288 GB of HBM3e, which is 12 times the 24 GB of GDDR7 on the NVIDIA GeForce RTX 5090 SE.
Specification Differences
| Specification | AMD Instinct MI355X | NVIDIA GeForce RTX 5090 SE |
|---|---|---|
| Architecture | CDNA 4.0 | Blackwell 2.0 |
| Process node | 3 nm | 5 nm |
| Transistors | 185,000 million | 92,200 million |
| Die size | 2380 mm² | 750 mm² |
| Transistor density | 77.7M / mm² | 122.9M / mm² |
| Base clock | 1000 MHz | 1740 MHz |
| Boost clock | 2400 MHz | 2377 MHz |
| Memory clock | 2000 MHz, 8 Gbps effective | 1750 MHz, 28 Gbps effective |
| Memory size | 288 GB | 24 GB |
| Memory type | HBM3e | GDDR7 |
| Memory bus | 8192 bit | 384 bit |
| Memory bandwidth | 8.19 TB/s | 1.34 TB/s |
| Shading units | 16384 | 14080 |
| TMUs | 1024 | 440 |
| ROPs | 0 | 160 |
| RT cores | Not recorded | 110 |
| Tensor cores | Not recorded | 440 |
| Pixel rate | 0 MPixel/s | 380.3 GPixel/s |
| Texture rate | 2,457.6 GTexel/s | 1,045.9 GTexel/s |
| FP32 | 78.64 TFLOPS | 66.94 TFLOPS |
| FP16 | 78.64 TFLOPS (1:1) | 66.94 TFLOPS (1:1) |
| TDP | 1400 W | 500 W |
| Slot width | OAM Module | Dual-slot |
| Power connectors | None | 1x 16-pin |
| Suggested PSU | 1800 W | 900 W |
| Display outputs | No outputs | 1x HDMI 2.1b, 3x DisplayPort 2.1b |
| DirectX | N/A | 12 Ultimate (12_2) |
| OpenGL | N/A | 4.6 |
| Vulkan | N/A | 1.4 |
| Dimensions | 102 mm x 165 mm | 267 mm x 111 mm x 40 mm |
| Release date | 2025-06-11 | 2025-12-31 |
| Production status | Not recorded | Active |
| Predecessor | Radeon Instinct | GeForce 40 |
| Successor | Not recorded | GeForce 60 |
The Verdict
The data indicates a clear division of purpose. The AMD Instinct MI355X is a compute accelerator for large-scale data processing, with 288 GB of HBM3e, 8.19 TB/s bandwidth, and 78.64 TFLOPS FP32/FP16 throughput. Its zero pixel rate, absence of display outputs, and lack of graphics API support confirm it cannot function as a rendering device. The 1400 W TDP and OAM form factor place it in server infrastructure where power delivery and cooling are managed at the rack level.
The NVIDIA GeForce RTX 5090 SE is a graphics card that also performs compute. Its 380.3 GPixel/s pixel rate, 110 RT cores, and 440 tensor cores enable ray-traced rendering, rasterization, and AI-accelerated graphics. The 500 W TDP and dual-slot design allow it to operate in a standard desktop chassis with a 900 W PSU. Its 24 GB GDDR7 memory and 1.34 TB/s bandwidth are sufficient for interactive workloads and moderate-sized datasets but fall short of the MI355X by an order of magnitude.
Users with workloads that require holding very large datasets in fast memory, such as training large neural networks or processing scientific simulations with multi-hundred-gigabyte working sets, should select the AMD Instinct MI355X. Its 12x memory capacity and 6.11x bandwidth advantage over the RTX 5090 SE are decisive for such tasks. Users who need a GPU for real-time rendering, ray tracing, or general desktop graphics should select the NVIDIA GeForce RTX 5090 SE, as it is the only one of the two with any display capability or graphics API support. The recorded data offers no overlap: the MI355X cannot output pixels, and the RTX 5090 SE cannot match the MI355X's memory or raw compute scale. The choice is determined entirely by whether the workload requires rendering or massive memory residency.