AMD Instinct MI355X vs NVIDIA GeForce RTX 4080 SUPER Comparison

AMD
RADEON

AMD Instinct MI355X

CORE STATE MI350 256CU
VRAM 288 GB
CLOCK SPEED 2400 MHz
TDP 1400 W
BUS WIDTH 8192 bit
ARCHITECTURE CDNA 4.0
nm
PROCESS 3 nm
LAUNCH DATE 2025
VS
NVIDIA
GEFORCE

GeForce RTX 4080 SUPER

CORE STATE AD103
VRAM 16 GB
CLOCK SPEED 2550 MHz
TDP 320 W
BUS WIDTH 256 bit
ARCHITECTURE Ada Lovelace
nm
PROCESS 5 nm
LAUNCH DATE 2024

PERFORMANCE BENCHMARKS

3dmark_3dmark_steel_nomad_dx12
N/A
6,600
geekbench_opencl
N/A
219,065
geekbench_vulkan
N/A
260,075
passmark_directx_10
N/A
193
passmark_directx_11
N/A
301
passmark_directx_12
N/A
134
passmark_directx_9
N/A
381
passmark_g2d
N/A
1,270
passmark_g3d
N/A
34,245
passmark_gpu_compute
N/A
19,822

Analysis: AMD Instinct MI355X vs NVIDIA GeForce RTX 4080 SUPER

The Verdict

The AMD Instinct MI355X and NVIDIA GeForce RTX 4080 SUPER are fundamentally different products with different intended workloads. The recorded data places the RTX 4080 SUPER at the 86th percentile among all GPUs, with an average benchmark score of 54,209. The MI355X has no registered benchmark scores, sits at the 50th percentile, and its average benchmark score is recorded as zero. This makes direct numerical comparison impossible, but the specification data clarifies the separation.

The RTX 4080 SUPER is the only one of the two with measured performance data. Its nearest rivals in the database are the NVIDIA GeForce RTX 4080 at 54,247 (0.1% slower), the AMD Radeon Pro W5700X at 54,828 (1.1% ahead), the AMD Radeon RX 6750 GRE 12 GB at 55,698 (2.7% ahead), and the AMD Radeon 8060S at 55,757 (2.8% ahead). This places the RTX 4080 SUPER in a tightly packed performance cluster where the largest gap between it and the nearest rival is only 2.8%.

The MI355X is a different class of device entirely. It is an OAM module with no display outputs, a 1400 W TDP, and 288 GB of HBM3e memory. The RTX 4080 SUPER is a triple-slot consumer graphics card with 16 GB of GDDR6X memory, a 320 W TDP, and display outputs. The data indicates the MI355X targets acceleration workloads where memory capacity and bandwidth dominate, while the RTX 4080 SUPER targets rendering and general compute with proven software support. The MI355X has no launch MSRP recorded, while the RTX 4080 SUPER launched at 999 USD.

FAQ

Q: Which GPU has a higher FP32 compute throughput?

A: The AMD Instinct MI355X delivers 78.64 TFLOPS of FP32 compute, which is 50.6% higher than the NVIDIA GeForce RTX 4080 SUPER at 52.22 TFLOPS. Both cards also deliver identical FP16 throughput at a 1:1 ratio with their FP32 numbers.

Q: How do the memory subsystems compare?

A: The MI355X uses 288 GB of HBM3e memory on an 8192-bit bus with 8.19 TB/s of bandwidth. The RTX 4080 SUPER uses 16 GB of GDDR6X memory on a 256-bit bus with 736.3 GB/s of bandwidth. The MI355X has 18 times the memory capacity and over 11 times the bandwidth.

Q: Does the MI355X support DirectX or Vulkan?

A: No. The database records all graphics APIs for the MI355X as N/A. The RTX 4080 SUPER supports DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4. The MI355X also has no display outputs.

Q: What is the power consumption difference?

A: The MI355X has a TDP of 1400 W with a suggested power supply of 1800 W. The RTX 4080 SUPER has a TDP of 320 W with a suggested power supply of 700 W. The MI355X uses no power connectors because it is an OAM module, while the RTX 4080 SUPER uses a single 16-pin connector.

Q: Which card has a higher transistor density?

A: The RTX 4080 SUPER has a transistor density of 121.1M per mm² across 45,900 million transistors on a 379 mm² die. The MI355X has 77.7M per mm² across 185,000 million transistors on a 2380 mm² die. The RTX 4080 SUPER packs transistors more densely despite being built on a 5 nm process versus the MI355X's 3 nm process.

Q: What does the average benchmark score say about each card?

A: The RTX 4080 SUPER has an average benchmark score of 54,209 across ten recorded tests, placing it in the 86th percentile. The MI355X has an average benchmark score of zero with no recorded tests, placing it in the 50th percentile. The database contains no performance measurements for the MI355X.

Specification Differences

The two GPUs differ in nearly every measured specification. The MI355X uses the MI350 256CU chip with CDNA 4.0 architecture, built on a 3 nm process at TSMC with 185,000 million transistors on a 2380 mm² die. The RTX 4080 SUPER uses the AD103 chip with Ada Lovelace architecture, built on a 5 nm process at TSMC with 45,900 million transistors on a 379 mm² die. Transistor density favors the RTX 4080 SUPER at 121.1M per mm² versus 77.7M per mm² for the MI355X.

Clock speeds differ substantially. The MI355X runs at a 1000 MHz base clock and 2400 MHz boost clock, with memory clocked at 2000 MHz (8 Gbps effective). The RTX 4080 SUPER runs at 2295 MHz base and 2550 MHz boost, with memory at 1438 MHz (23 Gbps effective).

The compute configuration diverges sharply. The MI355X has 16,384 shading units and 1,024 texture mapping units, with zero ROPs and a pixel rate of 0 MPixel/s. The RTX 4080 SUPER has 10,240 shading units, 320 TMUs, 112 ROPs, 80 RT cores, and 320 tensor cores, with a pixel rate of 285.6 GPixel/s and a texture rate of 816.0 GTexel/s. The MI355X's texture rate of 2,457.6 GTexel/s is roughly three times higher.

Memory capacity, type, bus width, and bandwidth all differ. The MI355X has 288 GB of HBM3e on an 8192-bit bus with 8.19 TB/s bandwidth. The RTX 4080 SUPER has 16 GB of GDDR6X on a 256-bit bus with 736.3 GB/s bandwidth.

Power and physical specifications are in different categories. The MI355X has a 1400 W TDP, no power connectors, and is an OAM module measuring 102 mm by 165 mm. The RTX 4080 SUPER has a 320 W TDP, one 16-pin connector, and is a triple-slot card measuring 310 mm by 140 mm by 61 mm. The MI355X uses PCIe 5.0 x16, while the RTX 4080 SUPER uses PCIe 4.0 x16. The MI355X has no display outputs; the RTX 4080 SUPER has one HDMI 2.1 and three DisplayPort 1.4a outputs.

The release dates are separated by roughly 17 months. The RTX 4080 SUPER was released on 2024-01-30, is marked end-of-life, and has the GeForce 50 series as its successor. The MI355X was released on 2025-06-11 with no production status or successor recorded.

Head-to-Head Benchmarks

The database contains no head-to-head benchmark results between the MI355X and the RTX 4080 SUPER. The MI355X has zero recorded benchmark entries, while the RTX 4080 SUPER has ten. The wins counters show zero for both cards in direct competition.

The RTX 4080 SUPER's recorded benchmarks provide the only available performance data. In 3DMark Steel Nomad DX12, it scores 6,600. Geekbench OpenCL returns 219,065, and Geekbench Vulkan returns 260,075. Passmark results include 34,245 in G3D, 19,822 in GPU compute, 1,270 in G2D, 381 in DirectX 9, 301 in DirectX 11, 193 in DirectX 10, and 134 in DirectX 12.

The MI355X's compute specifications suggest capabilities that the RTX 4080 SUPER cannot match, but no benchmark measurements exist to confirm this. The FP32 throughput gap is 26.42 TFLOPS in favor of the MI355X, and the memory bandwidth gap is 7.45 TB/s in favor of the MI355X. Texture rate favors the MI355X by 1,641.6 GTexel/s. The RTX 4080 SUPER counters with a pixel rate of 285.6 GPixel/s versus zero for the MI355X, and it has 112 ROPs where the MI355X has none.

The RTX 4080 SUPER's nearest rival data shows it performs within a narrow band of comparable products. It is 0.1% behind the RTX 4080, 1.1% behind the Radeon Pro W5700X, 2.7% behind the Radeon RX 6750 GRE 12 GB, and 2.8% behind the Radeon 8060S. The spread of 2.8 percentage points across four rivals indicates the RTX 4080 SUPER sits at the lower end of a tight performance cluster in the database's average scores.

Architecture Differences

The MI355X uses CDNA 4.0 architecture, AMD's compute-focused design lineage. It has no RT cores and no tensor cores listed, no ROPs, and no graphics API support. Its design targets data-center acceleration with 16,384 shading units and a 3 nm process. The 2380 mm² die with 185,000 million transistors makes it one of the largest GPUs in the database by die area.

The RTX 4080 SUPER uses Ada Lovelace architecture, NVIDIA's graphics and ray tracing design. It includes 80 RT cores and 320 tensor cores specifically for ray tracing and AI acceleration. It supports DirectX 12 Ultimate, OpenGL 4.6, and Vulkan 1.4. The 5 nm process with 45,900 million transistors on a 379 mm² die gives it a higher transistor density than the MI355X despite the older process node.

The memory architectures reflect different design goals. The MI355X's HBM3e stack on an 8192-bit bus delivers 8.19 TB/s, suitable for large model weights and high-bandwidth data streaming. The RTX 4080 SUPER's GDDR6X on a 256-bit bus delivers 736.3 GB/s, sufficient for graphics rendering but far below the MI355X's capacity. The MI355X has no pixel pipeline, confirming it is not designed for rasterization. The RTX 4080 SUPER's 112 ROPs and 285.6 GPixel/s pixel rate confirm its graphics focus.

The power architecture differs as well. The MI355X uses no power connectors because it is an OAM module, a board form factor designed for server integration. The RTX 4080 SUPER uses a single 16-pin connector and a suggested 700 W power supply. The MI355X's suggested 1800 W power supply indicates a system-level power requirement far beyond any single consumer GPU.

Where Each One Wins

The MI355X wins in every metric related to raw compute capacity and memory. FP32 throughput of 78.64 TFLOPS exceeds the RTX 4080 SUPER's 52.22 TFLOPS by 50.6%. Memory bandwidth of 8.19 TB/s is more than 11 times the RTX 4080 SUPER's 736.3 GB/s. Memory capacity of 288 GB is 18 times the 16 GB available on the RTX 4080 SUPER. Texture rate of 2,457.6 GTexel/s is roughly three times the RTX 4080 SUPER's 816.0 GTexel/s. The MI355X also uses a newer 3 nm process and a PCIe 5.0 interface versus PCIe 4.0 on the RTX 4080 SUPER.

The RTX 4080 SUPER wins in every metric related to graphics output and measured performance. It has 112 ROPs and a 285.6 GPixel/s pixel rate, while the MI355X has zero ROPs and zero pixel rate. It supports DirectX 12 Ultimate, OpenGL 4.6, and Vulkan 1.4, while the MI355X lists N/A for all three. It has display outputs, while the MI355X has none. It has a higher base clock (2295 MHz versus 1000 MHz), a higher boost clock (2550 MHz versus 2400 MHz), and a higher transistor density (121.1M per mm² versus 77.7M per mm²). It carries a 320 W TDP versus 1400 W, and its suggested power supply is 700 W versus 1800 W.

The RTX 4080 SUPER is the only one with recorded benchmark results. Its average score of 54,209 places it in the 86th percentile, while the MI355X sits at the 50th percentile with no scores. The RTX 4080 SUPER also has a release date 17 months earlier and a successor already recorded in the GeForce 50 series.

The data supports a clear use-case split. The MI355X is built for memory-bound and compute-heavy acceleration workloads, with its 288 GB HBM3e pool and 8.19 TB/s bandwidth. The RTX 4080 SUPER is built for graphics rendering, ray tracing, and general compute, with its ROPs, RT cores, tensor cores, API support, and measured benchmark scores. The MI355X has no graphics capabilities, no API support, and no measured performance. The RTX 4080 SUPER has no memory capacity or bandwidth approaching the MI355X's level. Each card wins in the domain its specifications target.

DETAILED SPECIFICATIONS

SPECIFICATION
Instinct MI355X
RTX 4080 SUPER
Core Specs
Shading Units
16,384
10,240 -37.5%
Shaders
16,384
10,240 -37.5%
TMUs
1,024
320 -68.8%
ROPs
0
112 +∞%
Compute Units
256
SM Count
80
Clocks
Base Clock
1000 MHz
2295 MHz
Boost Clock
2400 MHz
2550 MHz
Memory Clock
2000 MHz 8 Gbps effective
1438 MHz 23 Gbps effective
Memory
Memory Size
288 GB
16 GB
VRAM (MB)
294,912
16,384 -94.4%
Memory Type
HBM3e
GDDR6X
Memory Bus
8192 bit
256 bit
Bandwidth
8.19 TB/s
736.3 GB/s
Cache
L1 Cache
32 KB (per CU)
128 KB (per SM)
L2 Cache
32 MB
64 MB
L3 Cache
256 MB
Performance
Pixel Rate
0 MPixel/s
285.6 GPixel/s
Texture Rate
2,457.6 GTexel/s
816.0 GTexel/s
FP32 (TFLOPS)
78.64 TFLOPS
52.22 TFLOPS
FP64 (TFLOPS)
39.32 TFLOPS (1:2)
816.0 GFLOPS (1:64)
FP16 (TFLOPS)
78.64 TFLOPS (1:1)
52.22 TFLOPS (1:1)
AI/RT
RT Cores
80
Tensor Cores
320
Matrix Cores
1,024
Power
TDP
1400 W
320 W
TDP (W)
1,400
320 -77.1%
Suggested PSU
1800 W
700 W
Power Connectors
None
1x 16-pin
Architecture
Architecture
CDNA 4.0
Ada Lovelace
GPU Name
MI350 256CU
AD103
Generation
Instinct (MIx)
GeForce 40
Process Size
3 nm
5 nm
Transistors
185,000 million
45,900 million
Die Size
2380 mm²
379 mm²
Foundry
TSMC
TSMC
Density
77.7M / mm²
121.1M / mm²
API Support
DirectX
12 Ultimate (12_2)
OpenGL
4.6
Vulkan
1.4
OpenCL
3.0
3.0
CUDA
8.9
Shader Model
6.9
Physical
Slot Width
OAM Module
Triple-slot
Length
102 mm 4 inches
310 mm 12.2 inches
Height
140 mm 5.5 inches
Outputs
No outputs
1x HDMI 2.13x DisplayPort 1.4a
Bus Interface
PCIe 5.0 x16
PCIe 4.0 x16
Other
Launch Price
999 USD
Production
End-of-life
Predecessor
Radeon Instinct
GeForce 30
Successor
GeForce 50
View Instinct MI355X Details View GeForce RTX 4080 SUPER Details