AMD Radeon Instinct MI60 vs NVIDIA GeForce RTX 4090 D Comparison

AMD
RADEON

AMD Radeon Instinct MI60

CORE STATE Vega 20
VRAM 32 GB
CLOCK SPEED 1800 MHz
TDP 300 W
BUS WIDTH 4096 bit
ARCHITECTURE GCN 5.1
nm
PROCESS 7 nm
LAUNCH DATE 2018
VS
NVIDIA
GEFORCE

GeForce RTX 4090 D

CORE STATE AD102
VRAM 24 GB
CLOCK SPEED 2520 MHz
TDP 425 W
BUS WIDTH 384 bit
ARCHITECTURE Ada Lovelace
nm
PROCESS 5 nm
LAUNCH DATE 2023

PERFORMANCE BENCHMARKS

geekbench_opencl
92,488
278,621
geekbench_vulkan
92,444
246,941
3dmark_3dmark_steel_nomad_dx12
N/A
8,587

Analysis: AMD Radeon Instinct MI60 vs NVIDIA GeForce RTX 4090 D

Head-to-Head Benchmarks

The recorded data shows a dominant performance gap between the NVIDIA GeForce RTX 4090 D and the AMD Radeon Instinct MI60 across the two shared benchmark tests. In Geekbench OpenCL, the RTX 4090 D scores 278,621 against 92,488 for the MI60, a delta of 201.3%. That is the single largest margin in the entire comparison. The NVIDIA card more than triples the AMD card's compute output in this test, which reflects raw parallel processing capability under the OpenCL API.

The Vulkan result narrows slightly but remains lopsided. The RTX 4090 D posts 246,941, while the MI60 manages 92,444, putting the NVIDIA card ahead by 167.1%. The gap shrinks by roughly 34 percentage points compared to OpenCL, which suggests the MI60's GCN architecture handles Vulkan's command and scheduling model somewhat more efficiently relative to its own OpenCL showing. Still, it is not enough to change the outcome. The NVIDIA card wins both head-to-head tests, giving it a 2-0 record in the direct comparison.

Looking at the broader database context, the RTX 4090 D sits at the 98th percentile among all GPUs, with an average benchmark score of 178,050. Its nearest rivals in the database are all NVIDIA data center and workstation parts: the RTX PRO 5000 Blackwell trails by 2.2%, the A100 SXM4 80 GB trails by 3.1%, the RTX 5000 Ada Generation trails by 3.6%, and the A100 SXM4 40 GB trails by 4.9%. The RTX 4090 D is therefore not just ahead of the MI60, it is within a small margin of some of the most capable accelerators NVIDIA has produced. The MI60, by contrast, sits at the 93rd percentile with an average score of 92,466. Its nearest rivals include the RTX A4500, which it leads by 0.9%, and the RTX A4500 Mobile, which it leads by 1.5%. It trails the AMD Radeon Pro VII by 4.8% and the AMD Radeon RX 7900M by 5.2%. So while the MI60 is competitive within its own performance tier, that tier is roughly half the performance level of the RTX 4090 D.

The average benchmark score difference is stark: 178,050 versus 92,466. That is a 92.5% gap in the database's aggregate metric, which combines multiple tests into a single figure. The head-to-head deltas are even larger than the average score delta, meaning the two tests shared by both cards are among the more favorable comparisons for the NVIDIA part. The data does not show any test where the MI60 closes the gap to within double digits.

Architecture Differences

The two cards come from fundamentally different design eras and philosophies. The NVIDIA GeForce RTX 4090 D uses the AD102 chip built on Ada Lovelace architecture, fabricated on a 5 nm process at TSMC. It packs 76,300 million transistors into a 609 mm² die, yielding a transistor density of 125.3 million per square millimeter. The AMD Radeon Instinct MI60 uses the Vega 20 chip on GCN 5.1 architecture, also built at TSMC but on a 7 nm process. It contains 13,230 million transistors across a 331 mm² die, for a density of 40.0 million per square millimeter. The NVIDIA chip has 5.8 times the transistor count and nearly double the die area, while the process node advantage gives it more than three times the transistor density. These are not incremental differences; they represent a generational leap in manufacturing and design.

The compute resources diverge sharply. The RTX 4090 D has 14,592 shading units, 456 texture mapping units, and 176 render output units. It also includes 114 ray tracing cores and 456 tensor cores, features entirely absent from the MI60, which has no ray tracing cores and no tensor cores. The MI60 fields 4,096 shading units, 256 TMUs, and 64 ROPs. That means the NVIDIA card has roughly 3.6 times the shading units, 1.8 times the TMUs, and 2.75 times the ROPs of the AMD card. The FP32 throughput tells the story clearly: 73.54 TFLOPS for the RTX 4090 D versus 14.75 TFLOPS for the MI60. The NVIDIA card is almost exactly five times faster in single-precision floating point. In FP16, the RTX 4090 D again delivers 73.54 TFLOPS with a 1:1 ratio, while the MI60 reaches 29.49 TFLOPS via a 2:1 ratio. The NVIDIA card still holds a 2.5-to-1 advantage in half precision, but the MI60's 2:1 throughput doubling means the gap narrows considerably in FP16 workloads compared to FP32.

Memory configurations reflect different priorities. The RTX 4090 D comes with 24 GB of GDDR6X on a 384-bit bus, delivering 1.01 TB/s of bandwidth. The MI60 offers 32 GB of HBM2 on a 4096-bit bus, delivering 1.02 TB/s. The bandwidth figures are nearly identical, differing by only 0.01 TB/s, but the memory technologies and bus widths could not be more different. The MI60's HBM2 stack uses a massive 4096-bit interface to achieve its bandwidth, while the RTX 4090 D uses a relatively narrow 384-bit bus with high-speed GDDR6X. The MI60 has more capacity by 8 GB, which matters for very large datasets that must fit entirely in VRAM. However, the RTX 4090 D's memory runs at 1313 MHz with 21 Gbps effective speed, while the MI60's memory runs at 1000 MHz with 2 Gbps effective. The NVIDIA part relies on much faster signaling per pin, the AMD part on raw bus width.

Clock speeds also differ substantially. The RTX 4090 D has a base clock of 2280 MHz and a boost clock of 2520 MHz. The MI60 has a base clock of 1200 MHz and a boost clock of 1800 MHz. The NVIDIA card's boost clock is 40% higher than the MI60's, and its base clock is 90% higher. Pixel and texture rates follow: the RTX 4090 D delivers 443.5 GPixel/s and 1,149.1 GTexel/s, versus 115.2 GPixel/s and 460.8 GTexel/s for the MI60. Power draw reflects the performance gap: the RTX 4090 D has a 425 W TDP and requires a 1x 16-pin connector with an 800 W suggested PSU. The MI60 draws 300 W, uses a 1x 6-pin plus 1x 8-pin configuration, and suggests a 700 W PSU. The physical cards also differ in size: the RTX 4090 D is a triple-slot card measuring 304 mm long, 137 mm tall, and 61 mm wide, while the MI60 is a dual-slot card at 267 mm long and 111 mm tall, with no width listed in the database.

The RTX 4090 D supports DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4. The MI60 supports DirectX 12 (12_1), OpenGL 4.6, and Vulkan 1.3. The NVIDIA card has a more modern DirectX feature level and a newer Vulkan revision. Display outputs also differ: the RTX 4090 D has 1x HDMI 2.1 and 3x DisplayPort 1.4a, while the MI60 has only 1x mini-DisplayPort 1.4a. The MI60 is clearly a compute-oriented accelerator with minimal display functionality, while the RTX 4090 D is a full-featured graphics card. Both are end-of-life products, with the RTX 4090 D released on 2023-12-27 and the MI60 released on 2018-11-17, a five-year gap that explains much of the architectural disparity.

The Verdict

The data supports a straightforward conclusion: the NVIDIA GeForce RTX 4090 D is the superior performer by every measured metric in the shared benchmark suite. It wins both head-to-head tests with margins of 201.3% and 167.1%. Its average benchmark score of 178,050 is 92.5% higher than the MI60's 92,466. The RTX 4090 D also holds a 98th percentile ranking among all GPUs, compared to the MI60's 93rd percentile. For any workload that depends on raw compute throughput, FP32 or FP16 performance, or modern API support, the RTX 4090 D is the clear choice.

The MI60's advantages are narrower but real. It offers 32 GB of memory versus 24 GB, which is 33% more capacity. Its bandwidth of 1.02 TB/s is essentially identical to the RTX 4090 D's 1.01 TB/s, despite the much older architecture. It also draws 125 W less power, at 300 W versus 425 W, and fits in a dual-slot form factor versus the RTX 4090 D's triple-slot design. For deployments where memory capacity is the binding constraint, power is limited, or physical space is tight, the MI60 has a legitimate case. But the benchmark data shows that in raw performance, the MI60 is not in the same class.

The RTX 4090 D's nearest rivals in the database are all NVIDIA parts: the RTX PRO 5000 Blackwell, the A100 SXM4 80 GB, the RTX 5000 Ada Generation, and the A100 SXM4 40 GB. The MI60's nearest rivals are the RTX A4500, the RTX A4500 Mobile, the Radeon Pro VII, and the Radeon RX 7900M. That placement alone tells the story: the RTX 4090 D competes with flagship data center accelerators, while the MI60 competes with upper-midrange workstation cards. The delta between the two cards is not a close call.

FAQ

Q: Which card has more memory?

A: The AMD Radeon Instinct MI60 has 32 GB of HBM2, while the NVIDIA GeForce RTX 4090 D has 24 GB of GDDR6X. The MI60 offers 33% more capacity.

Q: How do their memory bandwidths compare?

A: They are nearly identical. The RTX 4090 D delivers 1.01 TB/s over a 384-bit bus, and the MI60 delivers 1.02 TB/s over a 4096-bit bus.

Q: Does the MI60 support ray tracing or tensor operations?

A: No. The MI60 has no ray tracing cores and no tensor cores. The RTX 4090 D includes 114 ray tracing cores and 456 tensor cores.

Q: What is the largest benchmark margin between the two cards?

A: The largest margin is in Geekbench OpenCL, where the RTX 4090 D scores 278,621 versus 92,488 for the MI60, a 201.3% difference.

Q: How do their power requirements differ?

A: The RTX 4090 D has a 425 W TDP and requires a 1x 16-pin connector with an 800 W suggested PSU. The MI60 has a 300 W TDP, uses 1x 6-pin plus 1x 8-pin connectors, and suggests a 700 W PSU.

Q: Which card has better API support?

A: The RTX 4090 D supports DirectX 12 Ultimate (12_2) and Vulkan 1.4, while the MI60 supports DirectX 12 (12_1) and Vulkan 1.3.

Where Each One Wins

The NVIDIA GeForce RTX 4090 D wins in every compute benchmark recorded in the database. Its 73.54 TFLOPS FP32 throughput dwarfs the MI60's 14.75 TFLOPS, making it the clear pick for single-precision compute workloads such as simulation, rendering, and general-purpose GPU compute. Its FP16 performance of 73.54 TFLOPS also exceeds the MI60's 29.49 TFLOPS, though the margin narrows to 2.5-to-1 in half-precision tasks. The 456 tensor cores give it a major advantage in AI inference and training workloads that can use Tensor Core acceleration, while the 114 ray tracing cores enable hardware-accelerated ray tracing, a feature the MI60 cannot offer at all. The RTX 4090 D also has a newer Vulkan revision (1.4 versus 1.3) and a higher DirectX feature level (12_2 versus 12_1), which matters for compatibility with modern graphics applications. Its triple-slot cooler and 425 W TDP are the cost of that performance, but the benchmark results justify the power envelope.

The AMD Radeon Instinct MI60 wins in a few specific scenarios that are not captured by the head-to-head benchmark scores. Its 32 GB memory capacity exceeds the RTX 4090 D's 24 GB, which is an advantage for workloads that need to hold larger models, datasets, or frame buffers entirely in VRAM without spilling to system memory. Its 4096-bit HBM2 bus achieves 1.02 TB/s bandwidth, essentially matching the RTX 4090 D's 1.01 TB/s despite the much older architecture. The MI60's 300 W TDP is 125 W lower, and its dual-slot design takes up less physical space than the RTX 4090 D's triple-slot footprint. For dense server installations where power density and slot count are constrained, the MI60 is the more practical choice. Its 7 nm process node, while older than the RTX 4090 D's 5 nm node, still delivers respectable efficiency for its era, and its GCN 5.1 architecture has mature software support in compute-oriented ecosystems. The MI60 also has no display outputs beyond a single mini-DisplayPort 1.4a, which reflects its accelerator-first positioning, but that is not a drawback in headless compute deployments.

For users choosing between these two cards, the decision hinges on workload priority. If raw compute speed, modern API support, ray tracing, and tensor acceleration matter most, the RTX 4090 D is the only rational choice. If memory capacity, power efficiency, and physical size are the primary constraints, the MI60 remains viable. The benchmark data does not show a single test where the MI60 outperforms the RTX 4090 D, so the AMD card's case rests entirely on its 8 GB memory advantage and lower power draw.

DETAILED SPECIFICATIONS

SPECIFICATION
Instinct MI60
RTX 4090 D
Core Specs
Shading Units
4,096
14,592 +256.3%
Shaders
4,096
14,592 +256.3%
TMUs
256
456 +78.1%
ROPs
64
176 +175.0%
Compute Units
64
—
SM Count
—
114
Clocks
Base Clock
1200 MHz
2280 MHz
Boost Clock
1800 MHz
2520 MHz
Memory Clock
1000 MHz 2 Gbps effective
1313 MHz 21 Gbps effective
Memory
Memory Size
32 GB
24 GB
VRAM (MB)
32,768
24,576 -25.0%
Memory Type
HBM2
GDDR6X
Memory Bus
4096 bit
384 bit
Bandwidth
1.02 TB/s
1.01 TB/s
Cache
L1 Cache
16 KB (per CU)
128 KB (per SM)
L2 Cache
4 MB
72 MB
Performance
Pixel Rate
115.2 GPixel/s
443.5 GPixel/s
Texture Rate
460.8 GTexel/s
1,149.1 GTexel/s
FP32 (TFLOPS)
14.75 TFLOPS
73.54 TFLOPS
FP64 (TFLOPS)
7.373 TFLOPS (1:2)
1,149.1 GFLOPS (1:64)
FP16 (TFLOPS)
29.49 TFLOPS (2:1)
73.54 TFLOPS (1:1)
AI/RT
RT Cores
—
114
Tensor Cores
—
456
Power
TDP
300 W
425 W
TDP (W)
300
425 +41.7%
Suggested PSU
700 W
800 W
Power Connectors
1x 6-pin + 1x 8-pin
1x 16-pin
Architecture
Architecture
GCN 5.1
Ada Lovelace
GPU Name
Vega 20
AD102
Generation
Radeon Instinct (MIx)
GeForce 40
Process Size
7 nm
5 nm
Transistors
13,230 million
76,300 million
Die Size
331 mm²
609 mm²
Foundry
TSMC
TSMC
Density
40.0M / mm²
125.3M / mm²
API Support
DirectX
12 (12_1)
12 Ultimate (12_2)
OpenGL
4.6
4.6
Vulkan
1.3
1.4
OpenCL
2.1
3.0
CUDA
—
8.9
Shader Model
6.7
6.8
Physical
Slot Width
Dual-slot
Triple-slot
Length
267 mm 10.5 inches
304 mm 12 inches
Height
111 mm 4.4 inches
137 mm 5.4 inches
Outputs
1x mini-DisplayPort 1.4a
1x HDMI 2.13x DisplayPort 1.4a
Bus Interface
PCIe 4.0 x16
PCIe 4.0 x16
Other
Launch Price
—
1,599 USD
Production
End-of-life
End-of-life
Predecessor
FirePro Data Center
GeForce 30
Successor
—
GeForce 50
View Radeon Instinct MI60 Details View GeForce RTX 4090 D Details