AMD Instinct MI350P vs NVIDIA GeForce RTX 4080 SUPER Comparison

AMD
RADEON

AMD Instinct MI350P

CORE STATE MI350 128CU
VRAM 144 GB
CLOCK SPEED 2200 MHz
TDP 600 W
BUS WIDTH 8192 bit
ARCHITECTURE CDNA 4.0
nm
PROCESS 3 nm
LAUNCH DATE 2026
VS
NVIDIA
GEFORCE

GeForce RTX 4080 SUPER

CORE STATE AD103
VRAM 16 GB
CLOCK SPEED 2550 MHz
TDP 320 W
BUS WIDTH 256 bit
ARCHITECTURE Ada Lovelace
nm
PROCESS 5 nm
LAUNCH DATE 2024

PERFORMANCE BENCHMARKS

3dmark_3dmark_steel_nomad_dx12
N/A
6,600
geekbench_opencl
N/A
219,065
geekbench_vulkan
N/A
260,075
passmark_directx_10
N/A
193
passmark_directx_11
N/A
301
passmark_directx_12
N/A
134
passmark_directx_9
N/A
381
passmark_g2d
N/A
1,270
passmark_g3d
N/A
34,245
passmark_gpu_compute
N/A
19,822

Analysis: AMD Instinct MI350P vs NVIDIA GeForce RTX 4080 SUPER

Head-to-Head Benchmarks

The recorded database contains benchmark scores for the NVIDIA GeForce RTX 4080 SUPER, while the AMD Instinct MI350P has no benchmark entries in the database. This makes a direct score-by-score comparison impossible. The RTX 4080 SUPER’s average benchmark score is 54,209, placing it in the 86th percentile of all GPUs tracked. Its nearest rivals in the database are the NVIDIA GeForce RTX 4080 with an average score of 54,247 (0.1% higher), the AMD Radeon Pro W5700X at 54,828 (1.1% higher), the AMD Radeon RX 6750 GRE 12 GB at 55,698 (2.7% higher), and the AMD Radeon 8060S at 55,757 (2.8% higher). The RTX 4080 SUPER trails these four cards by small margins, indicating its performance sits just below that cluster.

Looking at individual tests for the RTX 4080 SUPER, the 3DMark Steel Nomad DX12 result is 6,600. In Geekbench OpenCL, it scores 219,065, and in Geekbench Vulkan, 260,075. Passmark results show DirectX 10 at 193, DirectX 11 at 301, DirectX 12 at 134, DirectX 9 at 381, G2D at 1,270, G3D at 34,245, and GPU Compute at 19,822. The G3D score of 34,245 is the highest among the Passmark sub-tests, reflecting strong rasterization performance. The GPU Compute score of 19,822 is notably lower than G3D, suggesting compute workloads do not scale as well as graphics workloads on this card.

The AMD Instinct MI350P, by contrast, carries no benchmark scores in the database. Its average benchmark score is recorded as 0, and its percentile versus all GPUs is 50. This means the database has no measured evidence of its real-world performance. The head-to-head benchmark section is therefore empty, with zero wins recorded for either product. Any comparison of raw performance must rely on architectural specifications rather than measured results.

Architecture Differences

The two products use fundamentally different designs. The AMD Instinct MI350P is built on CDNA 4.0 architecture, manufactured on a 3 nm process at TSMC, with 73,000 million transistors on a die size of 1,190 mm². The transistor density is 61.3M per mm². Its chip is labeled MI350 128CU, indicating 128 compute units. The NVIDIA GeForce RTX 4080 SUPER uses Ada Lovelace architecture on a 5 nm process, also at TSMC, with 45,900 million transistors on a 379 mm² die. Its transistor density is 121.1M per mm², which is roughly double that of the MI350P despite the older process node, because the smaller die concentrates fewer transistors more densely.

The MI350P has 8,192 shading units, 512 texture mapping units, and 0 ROPs. Its pixel rate is recorded as 0 MPixel/s, and its texture rate is 1,126.4 GTexel/s. The RTX 4080 SUPER has 10,240 shading units, 320 TMUs, and 112 ROPs. Its pixel rate is 285.6 GPixel/s, and its texture rate is 816.0 GTexel/s. The MI350P has more TMUs and a higher texture rate, but no ROPs, meaning it cannot output pixels to a display. The RTX 4080 SUPER has dedicated ROPs and a functional pixel pipeline.

Memory configurations differ sharply. The MI350P uses 144 GB of HBM3e on an 8192-bit bus, delivering 8.19 TB/s of bandwidth. Memory clock is 2000 MHz, with 8 Gbps effective. The RTX 4080 SUPER has 16 GB of GDDR6X on a 256-bit bus, with bandwidth of 736.3 GB/s. Memory clock is 1438 MHz, with 23 Gbps effective. The MI350P’s bandwidth is over 11 times higher, but the RTX 4080 SUPER’s GDDR6X runs at a higher effective speed per pin.

Compute capabilities: the MI350P delivers 36.04 TFLOPS for both FP32 and FP16 (1:1 ratio). The RTX 4080 SUPER delivers 52.22 TFLOPS for FP32 and 52.22 TFLOPS for FP16 (1:1). The RTX 4080 SUPER has a clear lead in raw floating-point throughput. The MI350P does not list dedicated ray tracing cores or tensor cores, while the RTX 4080 SUPER has 80 RT cores and 320 tensor cores.

Power and physical specs: the MI350P has a TDP of 600 W, uses a single 16-pin power connector, and requires a suggested PSU of 1000 W. It is dual-slot, measuring 267 mm long, 111 mm high, and 40 mm wide. The RTX 4080 SUPER has a TDP of 320 W, also uses a single 16-pin connector, and requires a suggested PSU of 700 W. It is triple-slot, measuring 310 mm long, 140 mm high, and 61 mm wide. The MI350P is shorter and thinner but draws nearly double the power.

Bus interfaces: the MI350P uses PCIe 5.0 x16, while the RTX 4080 SUPER uses PCIe 4.0 x16. The MI350P has no display outputs, while the RTX 4080 SUPER has 1x HDMI 2.1 and 3x DisplayPort 1.4a. API support: the MI350P lists N/A for DirectX, OpenGL, and Vulkan. The RTX 4080 SUPER supports DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4.

Release timing: the MI350P is dated for May 6, 2026, while the RTX 4080 SUPER launched January 30, 2024. The RTX 4080 SUPER is marked end-of-life, with predecessor GeForce 30 and successor GeForce 50. The MI350P has predecessor Radeon Instinct and no successor listed.

Where Each One Wins

Based on the available data, the RTX 4080 SUPER wins in every measurable benchmark category because it is the only product with recorded scores. Its FP32 throughput of 52.22 TFLOPS exceeds the MI350P’s 36.04 TFLOPS by 45%. Its pixel rate of 285.6 GPixel/s is functional, whereas the MI350P’s is 0, meaning the RTX 4080 SUPER can render and output graphics, while the MI350P cannot. The RTX 4080 SUPER’s API support includes DirectX 12 Ultimate, OpenGL 4.6, and Vulkan 1.4, making it suitable for gaming and general graphics workloads. The MI350P has no API support listed, confirming its role is not consumer graphics.

The MI350P wins in memory capacity and bandwidth. Its 144 GB of HBM3e dwarfs the RTX 4080 SUPER’s 16 GB. Its bandwidth of 8.19 TB/s is 11.1 times higher than the RTX 4080 SUPER’s 736.3 GB/s. For workloads that require holding large datasets in on-card memory, such as large model inference or scientific simulation, the MI350P’s capacity is decisive. Its texture rate of 1,126.4 GTexel/s is 38% higher than the RTX 4080 SUPER’s 816.0 GTexel/s, indicating faster texture fetch operations. The MI350P also uses PCIe 5.0 x16, doubling the bus bandwidth versus the RTX 4080 SUPER’s PCIe 4.0 x16, which helps with data transfer to and from the host system.

The RTX 4080 SUPER has more shading units (10,240 vs 8,192) and higher FP32 throughput, so for general compute that is not memory-bound, it holds an advantage. The MI350P’s lack of ROPs and display outputs means it cannot drive monitors, so any use case requiring visual output falls to the RTX 4080 SUPER.

FAQ

Q: Which product has higher FP32 performance?

A: The NVIDIA GeForce RTX 4080 SUPER delivers 52.22 TFLOPS FP32, compared to the AMD Instinct MI350P’s 36.04 TFLOPS. The RTX 4080 SUPER is 45% faster in raw single-precision compute.

Q: What is the memory capacity difference?

A: The AMD Instinct MI350P has 144 GB of HBM3e, while the NVIDIA GeForce RTX 4080 SUPER has 16 GB of GDDR6X. The MI350P offers 9 times more memory capacity.

Q: Can the AMD Instinct MI350P output to a display?

A: No. The MI350P has no display outputs and 0 ROPs, with a pixel rate of 0 MPixel/s. Its API support for DirectX, OpenGL, and Vulkan is listed as N/A. The RTX 4080 SUPER has 1x HDMI 2.1 and 3x DisplayPort 1.4a outputs.

Q: What are the power requirements?

A: The MI350P has a TDP of 600 W and a suggested PSU of 1000 W. The RTX 4080 SUPER has a TDP of 320 W and a suggested PSU of 700 W. Both use a single 16-pin power connector.

Q: Which architecture is newer?

A: The AMD Instinct MI350P uses CDNA 4.0 on a 3 nm process, while the NVIDIA GeForce RTX 4080 SUPER uses Ada Lovelace on a 5 nm process. The MI350P’s release date is May 6, 2026, versus January 30, 2024 for the RTX 4080 SUPER.

Q: Does the MI350P have ray tracing or tensor cores?

A: The database lists no RT cores or tensor cores for the MI350P. The RTX 4080 SUPER has 80 RT cores and 320 tensor cores.

The Verdict

The data shows two products aimed at entirely different workloads. The RTX 4080 SUPER is a consumer graphics card with end-of-life status, a launch MSRP of 999 USD, and a full suite of benchmark scores. Its average benchmark score of 54,209 places it in the 86th percentile, and its nearest rivals are all within 2.8% of its performance, indicating a tightly grouped competitive field. It offers display outputs, full API support, and 52.22 TFLOPS FP32, making it suitable for gaming, content creation, and general GPU compute.

The MI350P is a datacenter compute accelerator with no benchmarks recorded. Its 144 GB HBM3e and 8.19 TB/s bandwidth are unmatched by the RTX 4080 SUPER, but its 36.04 TFLOPS FP32 is lower. It has no ROPs, no display outputs, no API support, and no ray tracing or tensor cores. Its performance percentile is 50, which is the median, but that is based on zero measured scores. The MI350P’s advantage lies in memory capacity and bandwidth, not in raw compute throughput.

For a user needing a graphics card that outputs video, runs DirectX 12 Ultimate, and has proven benchmark performance, the RTX 4080 SUPER is the only choice from the recorded data. For a user needing to process very large datasets entirely on-card, the MI350P’s 144 GB memory and 8.19 TB/s bandwidth are compelling, but its lack of measured performance means its compute capability is unverified. The RTX 4080 SUPER’s 86th percentile standing is based on real scores; the MI350P’s 50th percentile is a placeholder.

Specification Differences

| Field | AMD Instinct MI350P | NVIDIA GeForce RTX 4080 SUPER |

|---|---|---|

| Architecture | CDNA 4.0 | Ada Lovelace |

| Process Node | 3 nm | 5 nm |

| Transistors | 73,000 million | 45,900 million |

| Die Size | 1190 mm² | 379 mm² |

| Transistor Density | 61.3M / mm² | 121.1M / mm² |

| Base Clock | 1000 MHz | 2295 MHz |

| Boost Clock | 2200 MHz | 2550 MHz |

| Memory Clock | 2000 MHz 8 Gbps effective | 1438 MHz 23 Gbps effective |

| Memory Size | 144 GB | 16 GB |

| Memory Type | HBM3e | GDDR6X |

| Memory Bus Width | 8192 bit | 256 bit |

| Memory Bandwidth | 8.19 TB/s | 736.3 GB/s |

| Shading Units | 8192 | 10240 |

| TMUs | 512 | 320 |

| ROPs | 0 | 112 |

| RT Cores | N/A | 80 |

| Tensor Cores | N/A | 320 |

| Pixel Rate | 0 MPixel/s | 285.6 GPixel/s |

| Texture Rate | 1,126.4 GTexel/s | 816.0 GTexel/s |

| FP32 | 36.04 TFLOPS | 52.22 TFLOPS |

| FP16 | 36.04 TFLOPS (1:1) | 52.22 TFLOPS (1:1) |

| TDP | 600 W | 320 W |

| Slot Width | Dual-slot | Triple-slot |

| Power Connectors | 1x 16-pin | 1x 16-pin |

| Suggested PSU | 1000 W | 700 W |

| Bus Interface | PCIe 5.0 x16 | PCIe 4.0 x16 |

| Display Outputs | No outputs | 1x HDMI 2.1, 3x DisplayPort 1.4a |

| DirectX | N/A | 12 Ultimate (12_2) |

| OpenGL | N/A | 4.6 |

| Vulkan | N/A | 1.4 |

| Length | 267 mm | 310 mm |

| Height | 111 mm | 140 mm |

| Width | 40 mm | 61 mm |

| Release Date | 2026-05-06 | 2024-01-30 |

| Production Status | Not listed | End-of-life |

| Predecessor | Radeon Instinct | GeForce 30 |

| Successor | Not listed | GeForce 50 |

| Launch MSRP | Not listed | 999 USD |

DETAILED SPECIFICATIONS

SPECIFICATION
Instinct MI350P
RTX 4080 SUPER
Core Specs
Shading Units
8,192
10,240 +25.0%
Shaders
8,192
10,240 +25.0%
TMUs
512
320 -37.5%
ROPs
0
112 +∞%
Compute Units
128
—
SM Count
—
80
Clocks
Base Clock
1000 MHz
2295 MHz
Boost Clock
2200 MHz
2550 MHz
Memory Clock
2000 MHz 8 Gbps effective
1438 MHz 23 Gbps effective
Memory
Memory Size
144 GB
16 GB
VRAM (MB)
147,456
16,384 -88.9%
Memory Type
HBM3e
GDDR6X
Memory Bus
8192 bit
256 bit
Bandwidth
8.19 TB/s
736.3 GB/s
Cache
L1 Cache
16 KB (per CU)
128 KB (per SM)
L2 Cache
16 MB
64 MB
L3 Cache
128 MB
—
Performance
Pixel Rate
0 MPixel/s
285.6 GPixel/s
Texture Rate
1,126.4 GTexel/s
816.0 GTexel/s
FP32 (TFLOPS)
36.04 TFLOPS
52.22 TFLOPS
FP64 (TFLOPS)
18.02 TFLOPS (1:2)
816.0 GFLOPS (1:64)
FP16 (TFLOPS)
36.04 TFLOPS (1:1)
52.22 TFLOPS (1:1)
AI/RT
RT Cores
—
80
Tensor Cores
—
320
Matrix Cores
512
—
Power
TDP
600 W
320 W
TDP (W)
600
320 -46.7%
Suggested PSU
1000 W
700 W
Power Connectors
1x 16-pin
1x 16-pin
Architecture
Architecture
CDNA 4.0
Ada Lovelace
GPU Name
MI350 128CU
AD103
Generation
Instinct (MIx)
GeForce 40
Process Size
3 nm
5 nm
Transistors
73,000 million
45,900 million
Die Size
1190 mm²
379 mm²
Foundry
TSMC
TSMC
Density
61.3M / mm²
121.1M / mm²
AMD MCM
MCM
2
—
API Support
DirectX
—
12 Ultimate (12_2)
OpenGL
—
4.6
Vulkan
—
1.4
OpenCL
3.0
3.0
CUDA
—
8.9
Shader Model
—
6.9
Physical
Slot Width
Dual-slot
Triple-slot
Length
267 mm 10.5 inches
310 mm 12.2 inches
Height
111 mm 4.4 inches
140 mm 5.5 inches
Outputs
No outputs
1x HDMI 2.13x DisplayPort 1.4a
Bus Interface
PCIe 5.0 x16
PCIe 4.0 x16
Other
Launch Price
—
999 USD
Production
—
End-of-life
Predecessor
Radeon Instinct
GeForce 30
Successor
—
GeForce 50
View Instinct MI350P Details View GeForce RTX 4080 SUPER Details