AMD Instinct MI355X vs NVIDIA GeForce RTX 5090 D V2 Comparison

AMD
RADEON

AMD Instinct MI355X

CORE STATE MI350 256CU
VRAM 288 GB
CLOCK SPEED 2400 MHz
TDP 1400 W
BUS WIDTH 8192 bit
ARCHITECTURE CDNA 4.0
nm
PROCESS 3 nm
LAUNCH DATE 2025
VS
NVIDIA
GEFORCE

GeForce RTX 5090 D V2

CORE STATE GB202
VRAM 24 GB
CLOCK SPEED 2407 MHz
TDP 575 W
BUS WIDTH 384 bit
ARCHITECTURE Blackwell 2.0
nm
PROCESS 5 nm
LAUNCH DATE 2025

PERFORMANCE BENCHMARKS

3dmark_3dmark_steel_nomad_dx12
N/A
16,504

Analysis: AMD Instinct MI355X vs NVIDIA GeForce RTX 5090 D V2

AMD Instinct MI355X and NVIDIA GeForce RTX 5090 D V2 represent two diametrically opposed philosophies in GPU design, one aimed at massive-scale compute workloads and the other at high-performance client graphics. The recorded data shows a stark contrast in physical scale, memory capacity, and architectural intent, with the MI355X built as an OAM module for server racks and the RTX 5090 D V2 as a dual-slot consumer card. The MI355X carries a 1400 W TDP and requires an 1800 W suggested PSU, while the RTX 5090 D V2 draws 575 W and asks for a 950 W PSU. These numbers alone frame the discussion: one device is a compute accelerator, the other is a graphics card, and their benchmark footprints reflect that divide.

Where Each One Wins

The AMD Instinct MI355X claims superiority in raw memory capacity and memory bandwidth, fields that dominate large-model inference and data-parallel workloads. Its 288 GB of HBM3e memory at 8.19 TB/s dwarfs the RTX 5090 D V2's 24 GB of GDDR7 at 1.34 TB/s, producing a 12x capacity advantage and a 6.1x bandwidth advantage. This positions the MI355X for problems that cannot fit in local memory on the NVIDIA card, such as very large neural network parameter sets or simulation datasets that need constant high-speed access.

The NVIDIA GeForce RTX 5090 D V2 wins in conventional graphics and client-side compute metrics. Its 104.8 TFLOPS FP32 throughput surpasses the MI355X's 78.64 TFLOPS by 33%. The RTX 5090 D V2 also delivers a real pixel rate of 423.6 GPixel/s, while the MI355X registers 0 MPixel/s, meaning the AMD part has no display outputs and cannot rasterize frames in the traditional sense. The presence of 170 RT cores and 680 tensor cores on the NVIDIA card versus the absence of those counts in the MI355X record indicates that graphics acceleration, ray tracing, and client AI features are the NVIDIA card's domain.

The benchmark data recorded for the RTX 5090 D V2 shows a single entry: 3DMark Steel Nomad DX12 score of 16504. This places it at the 59th percentile among all GPUs in the database, with its nearest rival being the NVIDIA T400 at 16508, a 0% delta. The AMD MI355X has no benchmark entries and sits at the 50th percentile with a score of zero, so its competitive position cannot be measured through the same direct comparison. The database indicates that the MI355X wins where memory scale matters, and the RTX 5090 D V2 wins where graphics rendering and single-device compute throughput matter.

Architecture Differences

The MI355X uses the MI350 256CU chip built on CDNA 4.0 architecture, fabricated on a 3 nm process at TSMC. This is a compute-optimized design with 16,384 shading units, 1,024 texture mapping units, and zero ROPs, reflecting that it does not output pixels. Its transistor count reaches 185,000 million on a die size of 2380 mm², yielding a density of 77.7 million transistors per square millimeter. The huge die and 3 nm node point to a design that prioritizes memory controllers and compute arrays over fixed-function graphics hardware.

The RTX 5090 D V2 uses the GB202 chip on Blackwell 2.0 architecture, also from TSMC but on a 5 nm process. It contains 92,200 million transistors on a 750 mm² die, giving a higher density of 122.9 million transistors per square millimeter. The architecture includes 21,760 shading units, 680 TMUs, 176 ROPs, 170 RT cores, and 680 tensor cores, a full suite of graphics and compute features. The MI355X lists no RT cores or tensor cores in its record, while the NVIDIA card supports DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4, APIs the AMD part entirely lacks.

Clock behavior also differs sharply. The MI355X runs a 1000 MHz base and 2400 MHz boost, while the RTX 5090 D V2 runs a 2017 MHz base and 2407 MHz boost. The NVIDIA card also reaches a higher FP16 rate at 104.8 TFLOPS (1:1) versus the MI355X's 78.64 TFLOPS (1:1), despite the AMD part having a wider 8192-bit memory bus. The 8192-bit bus on the MI355X supports its HBM3e memory, while the 384-bit bus on the RTX 5090 D V2 feeds GDDR7, a narrower but faster-per-pin configuration.

Head-to-Head Benchmarks

Direct head-to-head benchmark results between these two devices are absent from the database, which is expected given the MI355X has no recorded benchmark scores. The only comparable data point comes from the RTX 5090 D V2's 3DMark Steel Nomad DX12 result of 16504, which exceeds the MI355X's recorded average score of zero by definition, but that comparison is meaningless in practical terms because the AMD part is not designed for that workload.

The nearest rival data for the RTX 5090 D V2 provides useful context. The NVIDIA T400 scores 16508, a 0% delta from the RTX 5090 D V2, meaning the two effectively tie in this specific test. The AMD Radeon PRO W7500 scores 16415, a 0.5% delta, while the NVIDIA RTX PRO 6000 Blackwell scores 16408, a 0.6% delta. The AMD Radeon RX 5700 XT scores 16361, a 0.9% delta. These margins are extremely tight, all within 1% of each other, suggesting that in the Steel Nomad DX12 test, the RTX 5090 D V2 performs statistically at parity with a wide range of other GPUs, including older and workstation-oriented parts.

For the MI355X, the absence of any benchmark entries means its nearest rivals list is empty, and its percentile rank of 50 is a placeholder rather than a measured result. The data indicates that the MI355X's strengths lie outside the 3DMark-style rendering tests, in domains like memory-bound compute where the 8.19 TB/s bandwidth and 288 GB capacity would dominate. The largest memory advantage, 12x capacity and 6.1x bandwidth, is the only quantitative metric where the MI355X clearly leads, and that lead does not translate to any recorded graphics benchmark.

FAQ

Q: Which GPU has more memory bandwidth?

A: The AMD Instinct MI355X has 8.19 TB/s of bandwidth from 288 GB of HBM3e memory, while the NVIDIA GeForce RTX 5090 D V2 has 1.34 TB/s from 24 GB of GDDR7, making the MI355X 6.1x faster in bandwidth.

Q: Which card can output video to a display?

A: The NVIDIA GeForce RTX 5090 D V2 has 1x HDMI 2.1b and 3x DisplayPort 2.1b outputs, while the AMD Instinct MI355X lists no display outputs and a pixel rate of 0 MPixel/s, so it cannot drive a monitor.

Q: What is the FP32 compute difference?

A: The NVIDIA GeForce RTX 5090 D V2 delivers 104.8 TFLOPS FP32, which is 33% higher than the AMD Instinct MI355X's 78.64 TFLOPS FP32.

Q: How does the RTX 5090 D V2 compare to its nearest rivals in the recorded benchmark?

A: Its 3DMark Steel Nomad DX12 score of 16504 ties the NVIDIA T400 at 16508 (0% delta), edges the AMD Radeon PRO W7500 at 16415 (0.5% delta), and leads the NVIDIA RTX PRO 6000 Blackwell at 16408 (0.6% delta) and AMD Radeon RX 5700 XT at 16361 (0.9% delta).

Q: Which GPU has a higher boost clock?

A: The NVIDIA GeForce RTX 5090 D V2 boosts to 2407 MHz, slightly higher than the AMD Instinct MI355X's 2400 MHz boost, though the MI355X has a much lower base clock of 1000 MHz versus 2017 MHz.

Q: What is the transistor density comparison?

A: The NVIDIA GeForce RTX 5090 D V2 packs 122.9 million transistors per square millimeter on a 750 mm² die, while the AMD Instinct MI355X has 77.7 million per square millimeter on a 2380 mm² die, giving the NVIDIA part higher density despite fewer total transistors.

Specification Differences

| Specification | AMD Instinct MI355X | NVIDIA GeForce RTX 5090 D V2 |

|----------------|---------------------|------------------------------|

| Chip | MI350 256CU | GB202 |

| Architecture | CDNA 4.0 | Blackwell 2.0 |

| Process Node | 3 nm (TSMC) | 5 nm (TSMC) |

| Transistors | 185,000 million | 92,200 million |

| Die Size | 2380 mm² | 750 mm² |

| Transistor Density | 77.7M / mm² | 122.9M / mm² |

| Base Clock | 1000 MHz | 2017 MHz |

| Boost Clock | 2400 MHz | 2407 MHz |

| Memory Size | 288 GB HBM3e | 24 GB GDDR7 |

| Memory Bus Width | 8192 bit | 384 bit |

| Memory Bandwidth | 8.19 TB/s | 1.34 TB/s |

| Shading Units | 16384 | 21760 |

| TMUs | 1024 | 680 |

| ROPs | 0 | 176 |

| RT Cores | None listed | 170 |

| Tensor Cores | None listed | 680 |

| Pixel Rate | 0 MPixel/s | 423.6 GPixel/s |

| Texture Rate | 2,457.6 GTexel/s | 1,636.8 GTexel/s |

| FP32 | 78.64 TFLOPS | 104.8 TFLOPS |

| FP16 | 78.64 TFLOPS (1:1) | 104.8 TFLOPS (1:1) |

| TDP | 1400 W | 575 W |

| Slot Width | OAM Module | Dual-slot |

| Power Connectors | None | 1x 16-pin |

| Suggested PSU | 1800 W | 950 W |

| Display Outputs | No outputs | 1x HDMI 2.1b, 3x DisplayPort 2.1b |

| DirectX | N/A | 12 Ultimate (12_2) |

| OpenGL | N/A | 4.6 |

| Vulkan | N/A | 1.4 |

| Dimensions | 102 mm x 165 mm | 304 mm x 137 mm x 48 mm |

| Release Date | 2025-06-11 | 2025-08-14 |

| Production Status | Not listed | Active |

The Verdict

The data separates these two GPUs by function, not by merit. The AMD Instinct MI355X is a server accelerator with 288 GB of HBM3e, 8.19 TB/s bandwidth, and a 1400 W TDP in an OAM module format, suited for workloads that require holding enormous datasets in local memory. Its 78.64 TFLOPS FP32 is lower than the NVIDIA card, but its memory subsystem provides the only clear quantitative advantage in the entire record, and that advantage matters for inference and simulation tasks that the RTX 5090 D V2 cannot accommodate due to its 24 GB limit.

The NVIDIA GeForce RTX 5090 D V2 is a client graphics card with 104.8 TFLOPS FP32, 423.6 GPixel/s rendering, 170 RT cores, and full display output support. Its 575 W TDP and dual-slot design fit into a desktop chassis, and its 3DMark Steel Nomad DX12 score of 16504, ranking at the 59th percentile, ties with the NVIDIA T400 and edges out several workstation and older gaming GPUs by less than 1%. The RTX 5090 D V2 also supports DirectX 12 Ultimate, OpenGL 4.6, and Vulkan 1.4, while the MI355X supports none of those APIs.

The MI355X has no recorded benchmarks, so its performance in rendering workloads cannot be verified, and its zero ROP count confirms it is not built for graphics output. The RTX 5090 D V2 has no memory capacity remotely close to the MI355X, making it unsuitable for the largest memory-bound compute tasks. The choice reduces to workload: the MI355X for memory-scale compute in a server rack, the RTX 5090 D V2 for graphics, ray tracing, and client-side compute where its higher FP32 and FP16 throughput and established API support give it a measurable edge. The RTX 5090 D V2 has a launch MSRP of 2,299 USD, while the MI355X has no listed launch price.

DETAILED SPECIFICATIONS

SPECIFICATION
Instinct MI355X
RTX 5090 D V2
Core Specs
Shading Units
16,384
21,760 +32.8%
Shaders
16,384
21,760 +32.8%
TMUs
1,024
680 -33.6%
ROPs
0
176 +∞%
Compute Units
256
SM Count
170
Clocks
Base Clock
1000 MHz
2017 MHz
Boost Clock
2400 MHz
2407 MHz
Memory Clock
2000 MHz 8 Gbps effective
1750 MHz 28 Gbps effective
Memory
Memory Size
288 GB
24 GB
VRAM (MB)
294,912
24,576 -91.7%
Memory Type
HBM3e
GDDR7
Memory Bus
8192 bit
384 bit
Bandwidth
8.19 TB/s
1.34 TB/s
Cache
L1 Cache
32 KB (per CU)
128 KB (per SM)
L2 Cache
32 MB
96 MB
L3 Cache
256 MB
Performance
Pixel Rate
0 MPixel/s
423.6 GPixel/s
Texture Rate
2,457.6 GTexel/s
1,636.8 GTexel/s
FP32 (TFLOPS)
78.64 TFLOPS
104.8 TFLOPS
FP64 (TFLOPS)
39.32 TFLOPS (1:2)
1.637 TFLOPS (1:64)
FP16 (TFLOPS)
78.64 TFLOPS (1:1)
104.8 TFLOPS (1:1)
AI/RT
RT Cores
170
Tensor Cores
680
Matrix Cores
1,024
Power
TDP
1400 W
575 W
TDP (W)
1,400
575 -58.9%
Suggested PSU
1800 W
950 W
Power Connectors
None
1x 16-pin
Architecture
Architecture
CDNA 4.0
Blackwell 2.0
GPU Name
MI350 256CU
GB202
Generation
Instinct (MIx)
GeForce 50
Process Size
3 nm
5 nm
Transistors
185,000 million
92,200 million
Die Size
2380 mm²
750 mm²
Foundry
TSMC
TSMC
Density
77.7M / mm²
122.9M / mm²
API Support
DirectX
12 Ultimate (12_2)
OpenGL
4.6
Vulkan
1.4
OpenCL
3.0
3.0
CUDA
12.0
Shader Model
6.9
Physical
Slot Width
OAM Module
Dual-slot
Length
102 mm 4 inches
304 mm 12 inches
Height
137 mm 5.4 inches
Outputs
No outputs
1x HDMI 2.1b3x DisplayPort 2.1b
Bus Interface
PCIe 5.0 x16
PCIe 5.0 x16
Other
Launch Price
2,299 USD
Production
Active
Predecessor
Radeon Instinct
GeForce 40
Successor
GeForce 60
View Instinct MI355X Details View GeForce RTX 5090 D V2 Details