AMD Instinct MI355X vs NVIDIA GeForce RTX 5080 SUPER Comparison
AMD Instinct MI355X
GeForce RTX 5080 SUPER
PERFORMANCE BENCHMARKS
Analysis: AMD Instinct MI355X vs NVIDIA GeForce RTX 5080 SUPER
Head-to-Head Benchmarks
The recorded data for these two accelerators tells a remarkably lopsided story, though the headline numbers can be misleading without proper context. The NVIDIA GeForce RTX 5080 SUPER has a single benchmark entry in the database: a 3DMark Steel Nomad DX12 result of 3075 points. The AMD Instinct MI355X has no benchmark scores recorded at all, which places it at the 50th percentile among all GPUs in the database, while the RTX 5080 SUPER sits at the 19th percentile. That percentile gap is substantial, but it reflects the absence of gaming-oriented test data for the MI355X rather than an inherent performance deficit.
Looking at the nearest rivals for the RTX 5080 SUPER, its 3075 score places it 2.8% behind the NVIDIA Quadro P1000 (3163 points) and 3.4% behind the Intel Arc Pro B60 (3182 points). It edges out the NVIDIA GeForce 820A by 3.1% (2983 points) and the NVIDIA GeForce GTX 860M by 3.6% (2967 points). These are narrow margins, clustering the RTX 5080 SUPER in a tight performance band with older professional and mobile parts. The MI355X, lacking any comparable benchmark entries, cannot be placed against these same rivals in the database.
The raw compute specifications tell a different story. The MI355X delivers 78.64 TFLOPS of FP32 throughput, which is 39.7% higher than the RTX 5080 SUPER's 56.28 TFLOPS. Texture fill rates follow a similar pattern: the MI355X reaches 2,457.6 GTexel/s versus 879.3 GTexel/s for the RTX 5080 SUPER, a 179.5% advantage. The pixel rate comparison is unusual, as the MI355X reports 0 MPixel/s due to its lack of traditional ROPs, while the RTX 5080 SUPER manages 293.1 GPixel/s.
Architecture Differences
The architectural split between these two parts is fundamental. The AMD Instinct MI355X uses CDNA 4.0 architecture built on a 3 nm process at TSMC, while the NVIDIA GeForce RTX 5080 SUPER uses Blackwell 2.0 on a 5 nm process, also from TSMC. The MI355X packs 185,000 million transistors onto a 2380 mm² die, achieving a transistor density of 77.7 million per mm². The RTX 5080 SUPER has 45,600 million transistors on a 378 mm² die, with a higher density of 120.6 million per mm². The MI355X is clearly a compute-oriented monolithic design, while the RTX 5080 SUPER is a more conventional graphics processor.
Memory subsystems could not be more different. The MI355X uses 288 GB of HBM3e across an 8192-bit bus, delivering 8.19 TB/s of bandwidth. The RTX 5080 SUPER uses 24 GB of GDDR7 on a 256-bit bus, with 1.02 TB/s of bandwidth. That is a 12x capacity difference and an 8x bandwidth difference in favor of the MI355X. The MI355X also has far more shading units: 16,384 versus 10,752 for the RTX 5080 SUPER. Texture mapping units number 1,024 on the MI355X versus 336 on the RTX 5080 SUPER. The RTX 5080 SUPER has 112 ROPs, while the MI355X reports zero, reflecting its lack of traditional rasterization hardware.
Clock speeds favor the NVIDIA part substantially. The RTX 5080 SUPER has a base clock of 2295 MHz and a boost clock of 2617 MHz, while the MI355X runs at 1000 MHz base and 2400 MHz boost. Memory clocks are nominally similar at 2000 MHz, but effective data rates differ: 8 Gbps for the MI355X versus 32 Gbps for the RTX 5080 SUPER. The RTX 5080 SUPER also includes 84 RT cores and 336 tensor cores, features that the MI355X does not list at all. API support diverges completely: the RTX 5080 SUPER supports DirectX 12 Ultimate, OpenGL 4.6, and Vulkan 1.4, while the MI355X reports N/A for all three.
Where Each One Wins
The NVIDIA GeForce RTX 5080 SUPER wins in any scenario that requires traditional graphics rendering, rasterization, or real-time ray tracing. Its 112 ROPs, 84 RT cores, and full DirectX 12 Ultimate support make it suitable for gaming and interactive visualization workloads. The 293.1 GPixel/s pixel fill rate and 879.3 GTexel/s texture rate confirm this part is built for display output and frame generation. The 24 GB of GDDR7 memory, while modest compared to the MI355X, is paired with a 256-bit bus and 1.02 TB/s bandwidth, sufficient for high-resolution textures and modern game assets. The dual-slot form factor, 1x 16-pin power connector, and display outputs (1x HDMI 2.1b, 3x DisplayPort 2.1b) make it a drop-in component for a desktop workstation.
The AMD Instinct MI355X wins in compute density and memory capacity. The 288 GB HBM3e pool with 8.19 TB/s bandwidth is built for large model inference and training datasets that would never fit in the RTX 5080 SUPER's 24 GB. The 16,384 shading units and 78.64 TFLOPS FP32 throughput give it a raw arithmetic advantage that matters for scientific simulation, AI training, and other non-graphics workloads. The 2,457.6 GTexel/s texture rate, while nominally a graphics metric, also indicates massive parallel throughput for compute shaders and data processing. Its OAM Module form factor and lack of display outputs signal a server or data center orientation, not a desktop one.
The 3DMark Steel Nomad result for the RTX 5080 SUPER, 3075 points, places it in a competitive band with the Quadro P1000 and Arc Pro B60. That benchmark does not exist for the MI355X, so no direct comparison is possible. The MI355X's 50th percentile ranking versus the RTX 5080 SUPER's 19th percentile suggests the database has far more comparable data points for the NVIDIA part, likely because gaming and workstation GPUs are tested more frequently than data center accelerators.
FAQ
Q: Which GPU has more memory bandwidth?
A: The AMD Instinct MI355X has 8.19 TB/s of bandwidth from its HBM3e memory, while the NVIDIA GeForce RTX 5080 SUPER has 1.02 TB/s from GDDR7. That is an 8x difference in favor of the MI355X.
Q: Does the AMD Instinct MI355X support DirectX 12?
A: No. The MI355X reports N/A for DirectX, OpenGL, and Vulkan API support. The RTX 5080 SUPER supports DirectX 12 Ultimate, OpenGL 4.6, and Vulkan 1.4.
Q: What is the FP32 compute performance difference?
A: The MI355X delivers 78.64 TFLOPS of FP32 throughput, while the RTX 5080 SUPER delivers 56.28 TFLOPS. The MI355X is 39.7% higher.
Q: How does the RTX 5080 SUPER compare to its nearest database rivals?
A: Its 3DMark Steel Nomad DX12 score of 3075 is 2.8% behind the NVIDIA Quadro P1000 (3163), 3.4% behind the Intel Arc Pro B60 (3182), 3.1% ahead of the NVIDIA GeForce 820A (2983), and 3.6% ahead of the NVIDIA GeForce GTX 860M (2967).
Q: Can the MI355X be used in a desktop PC?
A: The MI355X is an OAM Module with no display outputs and no power connectors listed. It requires external power delivery through its module interface and is not designed for conventional desktop installation. The RTX 5080 SUPER is a dual-slot card with a 1x 16-pin power connector and standard display outputs.
Q: What are the process node sizes for each chip?
A: The MI355X uses a 3 nm process at TSMC, while the RTX 5080 SUPER uses a 5 nm process at TSMC. Despite the larger node, the RTX 5080 SUPER has a higher transistor density at 120.6 million per mm² versus 77.7 million per mm² for the MI355X.
The Verdict
The data points to two entirely different products that happen to share a PCIe 5.0 x16 interface. The AMD Instinct MI355X is a compute accelerator with 288 GB of HBM3e, 8.19 TB/s of bandwidth, and 78.64 TFLOPS FP32. It has no display outputs, no graphics API support, and no conventional ROPs. The NVIDIA GeForce RTX 5080 SUPER is a graphics card with 24 GB of GDDR7, 293.1 GPixel/s pixel fill, and full DirectX 12 Ultimate support. Its 3DMark Steel Nomad score of 3075 positions it in a cluster with professional and mobile GPUs from previous generations.
A builder or researcher needing massive memory capacity for large datasets should choose the MI355X. Its 288 GB frame buffer is 12x larger than the RTX 5080 SUPER's 24 GB, and its bandwidth advantage is 8x. A user needing a functioning graphics output for gaming, CAD, or real-time rendering should choose the RTX 5080 SUPER. It has the ROPs, RT cores, and API compatibility that the MI355X lacks entirely.
The transistor counts reflect different design philosophies: the MI355X uses 185,000 million transistors on a 2380 mm² die for maximum compute throughput, while the RTX 5080 SUPER uses 45,600 million on 378 mm² for balanced graphics and compute. The RTX 5080 SUPER achieves higher clock speeds (2617 MHz boost versus 2400 MHz) and higher transistor density, but the MI355X compensates with 4x the transistor budget and 6.3x the die area.
The TDP figures underscore the use case split: 1400 W for the MI355X versus 415 W for the RTX 5080 SUPER. The MI355X requires a suggested 1800 W PSU, while the RTX 5080 SUPER lists no suggested PSU. These are not competing products in the traditional sense. They serve different markets, and the database records confirm that the RTX 5080 SUPER has gaming and workstation benchmark data while the MI355X has none.
For compute workloads that fit within 24 GB of memory, the RTX 5080 SUPER offers a 56.28 TFLOPS FP32 baseline with the flexibility of a standard PCIe card. For workloads that need hundreds of gigabytes of high-bandwidth memory, the MI355X is the only option between these two. The RTX 5080 SUPER records a 50th percentile position among all GPUs, while the MI355X sits at the 19th percentile, but that ranking reflects benchmark availability, not absolute capability.
Specification Differences
| Field | AMD Instinct MI355X | NVIDIA GeForce RTX 5080 SUPER |
|-------|---------------------|-------------------------------|
| Architecture | CDNA 4.0 | Blackwell 2.0 |
| Process Node | 3 nm | 5 nm |
| Transistors | 185,000 million | 45,600 million |
| Die Size | 2380 mm² | 378 mm² |
| Transistor Density | 77.7M / mm² | 120.6M / mm² |
| Base Clock | 1000 MHz | 2295 MHz |
| Boost Clock | 2400 MHz | 2617 MHz |
| Memory Clock | 2000 MHz 8 Gbps effective | 2000 MHz 32 Gbps effective |
| Memory Size | 288 GB | 24 GB |
| Memory Type | HBM3e | GDDR7 |
| Memory Bus Width | 8192 bit | 256 bit |
| Memory Bandwidth | 8.19 TB/s | 1.02 TB/s |
| Shading Units | 16384 | 10752 |
| TMUs | 1024 | 336 |
| ROPs | 0 | 112 |
| RT Cores | Not listed | 84 |
| Tensor Cores | Not listed | 336 |
| Pixel Rate | 0 MPixel/s | 293.1 GPixel/s |
| Texture Rate | 2,457.6 GTexel/s | 879.3 GTexel/s |
| FP32 | 78.64 TFLOPS | 56.28 TFLOPS |
| FP16 | 78.64 TFLOPS (1:1) | 56.28 TFLOPS (1:1) |
| TDP | 1400 W | 415 W |
| Slot Width | OAM Module | Dual-slot |
| Power Connectors | None | 1x 16-pin |
| Suggested PSU | 1800 W | Not listed |
| Display Outputs | No outputs | 1x HDMI 2.1b 3x DisplayPort 2.1b |
| DirectX | N/A | 12 Ultimate (12_2) |
| OpenGL | N/A | 4.6 |
| Vulkan | N/A | 1.4 |
| Length | 102 mm 4 inches | 304 mm 12 inches |
| Height | Not listed | 137 mm 5.4 inches |
| Width | 165 mm 6.5 inches | 40 mm 1.6 inches |
| Production Status | Not listed | Active |
| Release Date | 2025-06-11 | 2025-12-31 |
| Launch MSRP | Not listed | 999 USD |