AMD Instinct MI350P vs NVIDIA GeForce RTX 5090 SE Comparison

AMD
RADEON

AMD Instinct MI350P

CORE STATE MI350 128CU
VRAM 144 GB
CLOCK SPEED 2200 MHz
TDP 600 W
BUS WIDTH 8192 bit
ARCHITECTURE CDNA 4.0
nm
PROCESS 3 nm
LAUNCH DATE 2026
VS
NVIDIA
GEFORCE

GeForce RTX 5090 SE

CORE STATE GB202
VRAM 24 GB
CLOCK SPEED 2377 MHz
TDP 500 W
BUS WIDTH 384 bit
ARCHITECTURE Blackwell 2.0
nm
PROCESS 5 nm
LAUNCH DATE 2026

Analysis: AMD Instinct MI350P vs NVIDIA GeForce RTX 5090 SE

Where Each One Wins

The recorded data draws a sharp line between these two accelerators. The AMD Instinct MI350P is built exclusively for compute workloads with no display outputs, no graphics API support, and a zero pixel rate. Its 144 GB of HBM3e memory and 8.19 TB/s bandwidth position it for large-scale data processing, model training, and memory-bound scientific tasks. The NVIDIA GeForce RTX 5090 SE, by contrast, is a fully featured graphics card with DirectX 12 Ultimate, OpenGL 4.6, Vulkan 1.4, and three DisplayPort 2.1b outputs plus one HDMI 2.1b. It delivers 380.3 GPixel/s pixel throughput and 1,045.9 GTexel/s texture rate, which makes it the only one of the two suitable for rendering, rasterization, or any visual output.

The win split in the database is clear: the MI350P wins on memory capacity, memory bandwidth, and raw texture rate, while the RTX 5090 SE wins on shading units, clock speeds, pixel rate, FP32 throughput, and API compatibility. The MI350P has 512 TMUs versus 440 on the RTX 5090 SE, giving it a 1,126.4 GTexel/s texture rate that edges out the NVIDIA card's 1,045.9 GTexel/s. However, the RTX 5090 SE counters with 14,080 shading units against 8,192, and a boost clock of 2377 MHz versus 2200 MHz, producing 66.94 TFLOPS FP32 against 36.04 TFLOPS on the AMD part. For any workload that relies on FP32 compute, the NVIDIA card holds a decisive edge. For workloads that need massive on-card memory or maximal texture throughput, the AMD card is the clear selection.

FAQ

Q: Which card has more memory bandwidth?

A: The AMD Instinct MI350P has 8.19 TB/s of bandwidth from its 8192-bit HBM3e interface, while the NVIDIA GeForce RTX 5090 SE has 1.34 TB/s from a 384-bit GDDR7 bus. The AMD card offers more than six times the bandwidth.

Q: Does the AMD Instinct MI350P support any display outputs?

A: No. The MI350P lists "No outputs" for display connections and has N/A for DirectX, OpenGL, and Vulkan support. The RTX 5090 SE has 1x HDMI 2.1b and 3x DisplayPort 2.1b outputs.

Q: What is the FP32 compute difference?

A: The NVIDIA GeForce RTX 5090 SE delivers 66.94 TFLOPS FP32, which is 85.7% higher than the AMD Instinct MI350P's 36.04 TFLOPS. Both cards run FP16 at the same rate as FP32 (1:1 ratio).

Q: Which card has more memory capacity?

A: The AMD Instinct MI350P has 144 GB of HBM3e, six times the 24 GB of GDDR7 on the NVIDIA GeForce RTX 5090 SE.

Q: What are the power requirements?

A: The AMD Instinct MI350P has a 600 W TDP and suggests a 1000 W power supply, while the NVIDIA GeForce RTX 5090 SE has a 500 W TDP and suggests a 900 W power supply. Both use a single 16-pin power connector.

Q: Which card has a higher transistor count?

A: The NVIDIA GeForce RTX 5090 SE has 92,200 million transistors on a 750 mm² die, while the AMD Instinct MI350P has 73,000 million transistors on a 1190 mm² die. The NVIDIA chip achieves a higher transistor density of 122.9M per mm² versus 61.3M per mm².

Head-to-Head Benchmarks

The head-to-head benchmark array in the database is empty, so the comparison relies on the recorded specification data and derived performance metrics. The largest single win for the NVIDIA GeForce RTX 5090 SE is in FP32 compute: 66.94 TFLOPS versus 36.04 TFLOPS, a margin of 30.90 TFLOPS or roughly 85.7% higher. This stems from its 14,080 shading units, 5,888 more than the MI350P's 8,192, combined with a boost clock of 2377 MHz versus 2200 MHz. The pixel rate difference is absolute: the RTX 5090 SE produces 380.3 GPixel/s while the MI350P produces 0 MPixel/s, because the AMD card has zero ROPs and no display pipeline.

The AMD Instinct MI350P counters with memory capacity and bandwidth. Its 144 GB of HBM3e is 120 GB more than the RTX 5090 SE's 24 GB of GDDR7. The bandwidth gap is even more pronounced: 8.19 TB/s versus 1.34 TB/s, a difference of 6.85 TB/s. In texture throughput, the MI350P's 512 TMUs at a 2200 MHz boost produce 1,126.4 GTexel/s, which is 80.5 GTexel/s ahead of the RTX 5090 SE's 1,045.9 GTexel/s from 440 TMUs. The AMD card also has a wider memory bus at 8192 bits versus 384 bits, which directly enables its bandwidth advantage.

Clock speed comparison shows the NVIDIA card operating at a higher frequency: 1740 MHz base and 2377 MHz boost versus 1000 MHz base and 2200 MHz boost for the AMD card. The memory clock also favors NVIDIA in effective rate: 28 Gbps effective versus 8 Gbps effective, though the AMD HBM3e memory compensates with its vastly wider bus. Die size and transistor density tell another part of the story: the GB202 chip on the RTX 5090 SE packs 92,200 million transistors into 750 mm², while the MI350P's 128CU chip spreads 73,000 million transistors across 1190 mm².

Specification Differences

The two cards share identical physical dimensions: 267 mm length, 111 mm height, 40 mm width, both dual-slot with a single 16-pin power connector and PCIe 5.0 x16 interface. The differences begin with memory: the MI350P uses 144 GB HBM3e on an 8192-bit bus, while the RTX 5090 SE uses 24 GB GDDR7 on a 384-bit bus. The power envelope differs by 100 W, with the AMD card at 600 W TDP and a 1000 W suggested PSU, versus 500 W TDP and 900 W suggested PSU for NVIDIA.

Shading units diverge sharply: 8,192 on the MI350P versus 14,080 on the RTX 5090 SE. Texture mapping units are 512 versus 440, but ROPs are 0 on the AMD card versus 160 on the NVIDIA card. The RTX 5090 SE includes 110 ray tracing cores and 440 tensor cores, while the MI350P lists neither. Clock speeds favor NVIDIA: 1740 MHz base and 2377 MHz boost against 1000 MHz base and 2200 MHz boost. The AMD card has no display outputs and no graphics API support, while the RTX 5090 SE supports DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4.

Architecture Differences

The AMD Instinct MI350P uses the CDNA 4.0 architecture built on TSMC's 3 nm process, with the MI350 128CU chip. The NVIDIA GeForce RTX 5090 SE uses the Blackwell 2.0 architecture on TSMC's 5 nm process, with the GB202 chip. The process node difference gives AMD a density advantage in terms of manufacturing, but the actual transistor density recorded is much higher on the NVIDIA chip: 122.9M transistors per mm² versus 61.3M per mm². This is because the MI350P die is 1190 mm², the largest in this comparison, while the GB202 die is 750 mm².

Memory architecture diverges completely: HBM3e on a 8192-bit bus for AMD versus GDDR7 on a 384-bit bus for NVIDIA. The AMD card provides no ray tracing cores, no tensor cores, and no graphics pipeline, reflecting its compute-only design. The NVIDIA card includes 110 RT cores and 440 tensor cores, alongside a full ROP setup of 160 units. The MI350P belongs to the Instinct (MIx) generation with a Radeon Instinct predecessor, while the RTX 5090 SE belongs to the GeForce 50-series with a GeForce 40 predecessor and a GeForce 60 successor. Release timing also differs: the MI350P has a recorded release date of 2026-05-06, while the RTX 5090 SE has a release date of 2025-12-31.

The Verdict

The data directs each card to a distinct audience. The AMD Instinct MI350P serves workloads that demand enormous memory capacity and bandwidth: 144 GB of HBM3e at 8.19 TB/s is the defining feature. Its 512 TMUs and 1,126.4 GTexel/s texture rate also make it strong for texture-heavy compute. However, it has no display outputs, no graphics API support, zero pixel rate, and only 36.04 TFLOPS FP32. It is not a graphics card in any functional sense.

The NVIDIA GeForce RTX 5090 SE is the general-purpose option. It delivers 66.94 TFLOPS FP32, 380.3 GPixel/s, full DirectX 12 Ultimate support, and four display outputs. Its 24 GB GDDR7 memory and 1.34 TB/s bandwidth are modest relative to the AMD card, but its 14,080 shading units and 440 tensor cores provide broad compute capability. The RTX 5090 SE also runs at higher clocks (2377 MHz boost versus 2200 MHz) and consumes 100 W less power.

For memory-bound training or inference workloads where data residency on the card matters more than raw FP32 throughput, the MI350P's 120 GB extra memory and 6.85 TB/s extra bandwidth are decisive. For rendering, graphics, or any workload requiring FP32 compute, display output, or ray tracing, the RTX 5090 SE is the only functional choice. The database records no head-to-head benchmark wins for either card, so the verdict rests on the specification split: the MI350P wins on memory and texture throughput, the RTX 5090 SE wins on compute, graphics, and power efficiency. The RTX 5090 SE has a launch MSRP of 1,499 USD.

DETAILED SPECIFICATIONS

SPECIFICATION
Instinct MI350P
RTX 5090 SE
Core Specs
Shading Units
8,192
14,080 +71.9%
Shaders
8,192
14,080 +71.9%
TMUs
512
440 -14.1%
ROPs
0
160 +∞%
Compute Units
128
—
SM Count
—
110
Clocks
Base Clock
1000 MHz
1740 MHz
Boost Clock
2200 MHz
2377 MHz
Memory Clock
2000 MHz 8 Gbps effective
1750 MHz 28 Gbps effective
Memory
Memory Size
144 GB
24 GB
VRAM (MB)
147,456
24,576 -83.3%
Memory Type
HBM3e
GDDR7
Memory Bus
8192 bit
384 bit
Bandwidth
8.19 TB/s
1.34 TB/s
Cache
L1 Cache
16 KB (per CU)
128 KB (per SM)
L2 Cache
16 MB
96 MB
L3 Cache
128 MB
—
Performance
Pixel Rate
0 MPixel/s
380.3 GPixel/s
Texture Rate
1,126.4 GTexel/s
1,045.9 GTexel/s
FP32 (TFLOPS)
36.04 TFLOPS
66.94 TFLOPS
FP64 (TFLOPS)
18.02 TFLOPS (1:2)
1,045.9 GFLOPS (1:64)
FP16 (TFLOPS)
36.04 TFLOPS (1:1)
66.94 TFLOPS (1:1)
AI/RT
RT Cores
—
110
Tensor Cores
—
440
Matrix Cores
512
—
Power
TDP
600 W
500 W
TDP (W)
600
500 -16.7%
Suggested PSU
1000 W
900 W
Power Connectors
1x 16-pin
1x 16-pin
Architecture
Architecture
CDNA 4.0
Blackwell 2.0
GPU Name
MI350 128CU
GB202
Generation
Instinct (MIx)
GeForce 50
Process Size
3 nm
5 nm
Transistors
73,000 million
92,200 million
Die Size
1190 mm²
750 mm²
Foundry
TSMC
TSMC
Density
61.3M / mm²
122.9M / mm²
AMD MCM
MCM
2
—
API Support
DirectX
—
12 Ultimate (12_2)
OpenGL
—
4.6
Vulkan
—
1.4
OpenCL
3.0
3.0
CUDA
—
12.0
Shader Model
—
6.9
Physical
Slot Width
Dual-slot
Dual-slot
Length
267 mm 10.5 inches
267 mm 10.5 inches
Height
111 mm 4.4 inches
111 mm 4.4 inches
Outputs
No outputs
1x HDMI 2.1b3x DisplayPort 2.1b
Bus Interface
PCIe 5.0 x16
PCIe 5.0 x16
Other
Launch Price
—
1,499 USD
Production
—
Active
Predecessor
Radeon Instinct
GeForce 40
Successor
—
GeForce 60
View Instinct MI350P Details View GeForce RTX 5090 SE Details