AMD Radeon Instinct MI60 vs NVIDIA GeForce RTX 4090 Comparison

AMD
RADEON

AMD Radeon Instinct MI60

CORE STATE Vega 20
VRAM 32 GB
CLOCK SPEED 1800 MHz
TDP 300 W
BUS WIDTH 4096 bit
ARCHITECTURE GCN 5.1
nm
PROCESS 7 nm
LAUNCH DATE 2018
VS
NVIDIA
GEFORCE

GeForce RTX 4090

CORE STATE AD102
VRAM 24 GB
CLOCK SPEED 2520 MHz
TDP 450 W
BUS WIDTH 384 bit
ARCHITECTURE Ada Lovelace
nm
PROCESS 5 nm
LAUNCH DATE 2022

PERFORMANCE BENCHMARKS

geekbench_opencl
92,488
255,416
geekbench_vulkan
92,444
271,631
3dmark_3dmark_steel_nomad_dx12
N/A
9,223
passmark_directx_10
N/A
224
passmark_directx_11
N/A
326
passmark_directx_12
N/A
150
passmark_directx_9
N/A
397
passmark_g2d
N/A
1,299
passmark_g3d
N/A
38,194
passmark_gpu_compute
N/A
26,613

Analysis: AMD Radeon Instinct MI60 vs NVIDIA GeForce RTX 4090

The AMD Radeon Instinct MI60 and the NVIDIA GeForce RTX 4090 represent two very different points in the GPU landscape, separated by nearly four years of architecture evolution and aimed at distinct workloads. The database records show the MI60 as a 2018-era data center compute card built on GCN 5.1, while the RTX 4090 is a 2022 consumer flagship built on Ada Lovelace. Their benchmark scores, however, tell a clear story of generational performance shifts, with the RTX 4090 dominating the recorded head-to-head results. This analysis examines where each card wins, how their architectures diverge, and what the recorded data implies for different use cases.

Where Each One Wins

Based on the recorded benchmark data, the NVIDIA GeForce RTX 4090 is the clear winner in every head-to-head comparison available in the database. The MI60 wins zero benchmark comparisons, while the RTX 4090 wins two. This is not a close contest in raw compute throughput. The RTX 4090 delivers a Geekbench OpenCL score of 255,416 compared to the MI60’s 92,488, and a Geekbench Vulkan score of 271,631 versus 92,444. The deltas are massive: the RTX 4090 leads by 63.8% in OpenCL and 66% in Vulkan. These are not marginal improvements; they represent a fundamental shift in compute capability.

However, the MI60 is not without its own positioning. Its average benchmark score of 92,466 places it in the 93rd percentile of all GPUs in the database, which is higher than the RTX 4090’s 88th percentile despite the latter’s far higher raw scores. This apparent contradiction stems from the benchmark pool: the MI60’s nearest rivals include the NVIDIA RTX A4500 (average score 91,671, only 0.9% lower) and the AMD Radeon Pro VII (97,131, which beats the MI60 by 4.8%). The RTX 4090’s nearest rivals, by contrast, include the Intel Arc Pro A60 (60,326, essentially tied at 0% delta) and the AMD Radeon Pro W6600M (61,896, which beats it by 2.5%). This suggests the RTX 4090’s average is dragged down by a much broader benchmark suite, while the MI60’s limited test set (only two Geekbench entries) may not capture its full data center capabilities.

For use-case splits, the data indicates the RTX 4090 is the choice for any workload measured by Geekbench OpenCL or Vulkan, which typically include general-purpose compute, rendering, and modern API-based graphics tasks. The MI60, with its 32 GB of HBM2 memory and 1.02 TB/s bandwidth, is built for large-memory data center tasks, but the recorded benchmarks do not show it winning any compute tests. The RTX 4090’s 24 GB of GDDR6X with 1.01 TB/s bandwidth is nearly equivalent in bandwidth, while offering far higher shader throughput.

Architecture Differences

The architectural chasm between these two GPUs is stark. The MI60 uses the Vega 20 chip on AMD’s GCN 5.1 architecture, fabricated on TSMC’s 7 nm process. The RTX 4090 uses the AD102 chip on NVIDIA’s Ada Lovelace architecture, fabricated on TSMC’s 5 nm process. This process shrink alone allows the RTX 4090 to pack 76,300 million transistors into a 609 mm² die, versus the MI60’s 13,230 million transistors on a 331 mm² die. Transistor density tells the story: the RTX 4090 achieves 125.3 million transistors per square millimeter, while the MI60 sits at 40.0 million per square millimeter. That is a threefold density advantage, enabling the RTX 4090 to house 16,384 shading units versus the MI60’s 4,096, and 512 texture mapping units versus 256.

The memory subsystems differ fundamentally. The MI60 uses HBM2 with a 4096-bit bus, yielding 1.02 TB/s bandwidth across 32 GB. The RTX 4090 uses GDDR6X with a 384-bit bus, yielding 1.01 TB/s bandwidth across 24 GB. The bandwidth is nearly identical, but the bus widths and memory types reflect different design philosophies: HBM2 for capacity and density in compute, GDDR6X for speed and cost in consumer graphics.

Feature sets diverge completely. The RTX 4090 includes 128 ray tracing cores and 512 tensor cores, neither of which exist on the MI60. The MI60’s FP16 throughput is 29.49 TFLOPS at a 2:1 ratio relative to FP32, while the RTX 4090 achieves 82.58 TFLOPS FP16 at a 1:1 ratio. This means the RTX 4090 does not double FP16 rate; it runs FP16 at the same rate as FP32, which is a hallmark of modern architectures that treat FP16 as a first-class compute type. The MI60’s FP32 is 14.75 TFLOPS, while the RTX 4090 reaches 82.58 TFLOPS, a 5.6x difference.

Power and physical design also differ. The MI60 is rated at 300 W TDP, uses a dual-slot cooler, and requires a 1x 6-pin plus 1x 8-pin power connector with a suggested 700 W PSU. The RTX 4090 is rated at 450 W TDP, uses a triple-slot cooler, and requires a single 16-pin connector with a suggested 850 W PSU. The RTX 4090 is longer at 304 mm versus 267 mm, and taller at 137 mm versus 111 mm. The MI60 has one mini-DisplayPort 1.4a output, while the RTX 4090 has one HDMI 2.1 and three DisplayPort 1.4a outputs.

Head-to-Head Benchmarks

The only two head-to-head benchmarks recorded are Geekbench OpenCL and Geekbench Vulkan, and the RTX 4090 wins both decisively. In Geekbench OpenCL, the RTX 4090 scores 255,416 against the MI60’s 92,488, a delta of 63.8%. This test measures raw compute performance across a variety of workloads, including integer and floating-point operations, memory access patterns, and parallel execution. The RTX 4090’s 16,384 shading units and 82.58 TFLOPS FP32 clearly overwhelm the MI60’s 4,096 shading units and 14.75 TFLOPS.

In Geekbench Vulkan, the gap narrows slightly but remains enormous: the RTX 4090 scores 271,631 versus the MI60’s 92,444, a delta of 66%. Vulkan is a low-level graphics and compute API, and the RTX 4090’s support for Vulkan 1.4 versus the MI60’s Vulkan 1.3 may contribute to its advantage, though the raw compute difference dominates. The RTX 4090 also has ray tracing and tensor cores, which can accelerate certain Vulkan workloads, but the baseline compute disparity is sufficient to explain the result.

The MI60’s nearest rival comparisons provide context. Its average score of 92,466 is only 0.9% above the NVIDIA RTX A4500 and 1.5% above the RTX A4500 Mobile, but it trails the AMD Radeon Pro VII by 4.8% and the AMD Radeon RX 7900M by 5.2%. This places the MI60 in a mid-pack position among professional and workstation GPUs, not at the top. The RTX 4090’s average score of 60,347 is effectively tied with the Intel Arc Pro A60 (0% delta) and 0.3% above the AMD Radeon Pro Vega 48, but it trails the AMD Radeon Pro W6600M by 2.5% and leads the AMD Radeon PRO V710 by 2.9%. This suggests the RTX 4090’s average is pulled down by a wider test suite that includes DirectX and Passmark tests where it may not excel relative to professional cards.

FAQ

Q: Which GPU has a higher average benchmark score?

A: The AMD Radeon Instinct MI60 has an average benchmark score of 92,466, while the NVIDIA GeForce RTX 4090 has an average of 60,347. However, this average is based on different test sets; the MI60 has only two Geekbench entries, while the RTX 4090 has ten entries including Passmark tests.

Q: What is the memory bandwidth difference between the two?

A: The MI60 provides 1.02 TB/s bandwidth from 32 GB of HBM2 on a 4096-bit bus. The RTX 4090 provides 1.01 TB/s from 24 GB of GDDR6X on a 384-bit bus. The bandwidth is nearly identical, differing by only 0.01 TB/s.

Q: Does the RTX 4090 support ray tracing?

A: Yes, the RTX 4090 includes 128 ray tracing cores. The MI60 has no ray tracing cores listed in the database.

Q: Which GPU has more shading units?

A: The RTX 4090 has 16,384 shading units, exactly four times the MI60’s 4,096 shading units.

Q: What are the release dates for these GPUs?

A: The AMD Radeon Instinct MI60 was released on 2018-11-17, and the NVIDIA GeForce RTX 4090 was released on 2022-09-19.

Q: How do their FP32 compute rates compare?

A: The RTX 4090 achieves 82.58 TFLOPS FP32, while the MI60 achieves 14.75 TFLOPS FP32, a difference of 5.6x in favor of the RTX 4090.

Specification Differences

The following fields differ between the two GPUs, based on the recorded data:

| Specification | AMD Radeon Instinct MI60 | NVIDIA GeForce RTX 4090 |

|----------------|--------------------------|--------------------------|

| Architecture | GCN 5.1 | Ada Lovelace |

| Process Node | 7 nm | 5 nm |

| Transistors | 13,230 million | 76,300 million |

| Die Size | 331 mm² | 609 mm² |

| Transistor Density | 40.0M / mm² | 125.3M / mm² |

| Base Clock | 1200 MHz | 2235 MHz |

| Boost Clock | 1800 MHz | 2520 MHz |

| Memory Clock | 1000 MHz (2 Gbps effective) | 1313 MHz (21 Gbps effective) |

| Memory Size | 32 GB | 24 GB |

| Memory Type | HBM2 | GDDR6X |

| Memory Bus Width | 4096 bit | 384 bit |

| Shading Units | 4096 | 16384 |

| TMUs | 256 | 512 |

| ROPs | 64 | 176 |

| RT Cores | None | 128 |

| Tensor Cores | None | 512 |

| Pixel Rate | 115.2 GPixel/s | 443.5 GPixel/s |

| Texture Rate | 460.8 GTexel/s | 1,290.2 GTexel/s |

| FP32 | 14.75 TFLOPS | 82.58 TFLOPS |

| FP16 | 29.49 TFLOPS (2:1) | 82.58 TFLOPS (1:1) |

| TDP | 300 W | 450 W |

| Slot Width | Dual-slot | Triple-slot |

| Power Connectors | 1x 6-pin + 1x 8-pin | 1x 16-pin |

| Suggested PSU | 700 W | 850 W |

| Display Outputs | 1x mini-DisplayPort 1.4a | 1x HDMI 2.1, 3x DisplayPort 1.4a |

| DirectX Version | 12 (12_1) | 12 Ultimate (12_2) |

| Vulkan Version | 1.3 | 1.4 |

| Dimensions | 267 mm x 111 mm | 304 mm x 137 mm x 61 mm |

| Release Date | 2018-11-17 | 2022-09-19 |

| Predecessor | FirePro Data Center | GeForce 30 |

| Successor | None | GeForce 50 |

| Launch MSRP | None listed | 1,599 USD |

The Verdict

The recorded data points to the NVIDIA GeForce RTX 4090 as the superior performer in every benchmark where both are tested. Its Geekbench OpenCL score is 63.8% higher, and its Geekbench Vulkan score is 66% higher, reflecting a 5.6x advantage in FP32 compute, a 4x advantage in shading units, and a 3.9x advantage in pixel rate. For any workload that relies on raw compute, ray tracing, or tensor operations, the RTX 4090 is the clear choice.

The AMD Radeon Instinct MI60, however, retains relevance for specific data center scenarios. Its 32 GB of HBM2 memory on a 4096-bit bus provides capacity that the RTX 4090 cannot match, even though bandwidth is nearly identical. The MI60 also has a lower TDP of 300 W versus 450 W, which may be attractive for dense server deployments where power density is a constraint. Its 93rd percentile ranking among all GPUs indicates it remains a competent compute card, even if its nearest rivals include the AMD Radeon Pro VII and RX 7900M that outperform it by around 5%.

For buyers, the decision hinges on workload. If the task involves modern graphics APIs, ray tracing, or tensor-heavy AI inference, the RTX 4090’s 128 RT cores, 512 tensor cores, and 82.58 TFLOPS FP16 make it unmatched in this comparison. If the task requires large memory buffers exceeding 24 GB, or if the deployment environment cannot accommodate a 450 W triple-slot card, the MI60’s 32 GB capacity and 300 W dual-slot design are the only advantages the data supports. The RTX 4090 also carries a launch MSRP of 1,599 USD, while the MI60 has no recorded launch MSRP, but both cards are end-of-life, so availability will determine practical choices. The data overwhelmingly favors the RTX 4090 for performance, with the MI60’s memory capacity as its sole differentiating strength.

DETAILED SPECIFICATIONS

SPECIFICATION
Instinct MI60
RTX 4090
Core Specs
Shading Units
4,096
16,384 +300.0%
Shaders
4,096
16,384 +300.0%
TMUs
256
512 +100.0%
ROPs
64
176 +175.0%
Compute Units
64
SM Count
128
Clocks
Base Clock
1200 MHz
2235 MHz
Boost Clock
1800 MHz
2520 MHz
Memory Clock
1000 MHz 2 Gbps effective
1313 MHz 21 Gbps effective
Memory
Memory Size
32 GB
24 GB
VRAM (MB)
32,768
24,576 -25.0%
Memory Type
HBM2
GDDR6X
Memory Bus
4096 bit
384 bit
Bandwidth
1.02 TB/s
1.01 TB/s
Cache
L1 Cache
16 KB (per CU)
128 KB (per SM)
L2 Cache
4 MB
72 MB
Performance
Pixel Rate
115.2 GPixel/s
443.5 GPixel/s
Texture Rate
460.8 GTexel/s
1,290.2 GTexel/s
FP32 (TFLOPS)
14.75 TFLOPS
82.58 TFLOPS
FP64 (TFLOPS)
7.373 TFLOPS (1:2)
1,290.2 GFLOPS (1:64)
FP16 (TFLOPS)
29.49 TFLOPS (2:1)
82.58 TFLOPS (1:1)
AI/RT
RT Cores
128
Tensor Cores
512
Power
TDP
300 W
450 W
TDP (W)
300
450 +50.0%
Suggested PSU
700 W
850 W
Power Connectors
1x 6-pin + 1x 8-pin
1x 16-pin
Architecture
Architecture
GCN 5.1
Ada Lovelace
GPU Name
Vega 20
AD102
Generation
Radeon Instinct (MIx)
GeForce 40
Process Size
7 nm
5 nm
Transistors
13,230 million
76,300 million
Die Size
331 mm²
609 mm²
Foundry
TSMC
TSMC
Density
40.0M / mm²
125.3M / mm²
API Support
DirectX
12 (12_1)
12 Ultimate (12_2)
OpenGL
4.6
4.6
Vulkan
1.3
1.4
OpenCL
2.1
3.0
CUDA
8.9
Shader Model
6.7
6.8
Physical
Slot Width
Dual-slot
Triple-slot
Length
267 mm 10.5 inches
304 mm 12 inches
Height
111 mm 4.4 inches
137 mm 5.4 inches
Outputs
1x mini-DisplayPort 1.4a
1x HDMI 2.13x DisplayPort 1.4a
Bus Interface
PCIe 4.0 x16
PCIe 4.0 x16
Other
Launch Price
1,599 USD
Production
End-of-life
End-of-life
Predecessor
FirePro Data Center
GeForce 30
Successor
GeForce 50
View Radeon Instinct MI60 Details View GeForce RTX 4090 Details