AMD Radeon Instinct MI60 vs NVIDIA CMP 40HX Comparison

AMD
RADEON

AMD Radeon Instinct MI60

CORE STATE Vega 20
VRAM 32 GB
CLOCK SPEED 1800 MHz
TDP 300 W
BUS WIDTH 4096 bit
ARCHITECTURE GCN 5.1
nm
PROCESS 7 nm
LAUNCH DATE 2018
VS
NVIDIA
GEFORCE

CMP 40HX

CORE STATE TU106
VRAM 8 GB
CLOCK SPEED 1650 MHz
TDP 185 W
BUS WIDTH 256 bit
ARCHITECTURE Turing
nm
PROCESS 12 nm
LAUNCH DATE 2021

PERFORMANCE BENCHMARKS

geekbench_opencl
92,488
93,395
geekbench_vulkan
92,444
77,879

Analysis: AMD Radeon Instinct MI60 vs NVIDIA CMP 40HX

The AMD Radeon Instinct MI60 and NVIDIA CMP 40HX are both end-of-life accelerators, yet they occupy opposite ends of the hardware spectrum. The MI60 is a 7nm data-center compute card with 32 GB of HBM2, while the CMP 40HX is a 12nm Turing-based mining card with 8 GB of GDDR6. Benchmark data shows a razor-thin split: the NVIDIA card wins OpenCL by 1%, while the AMD card dominates Vulkan by 18.7%. Both sit at the 93rd percentile among all GPUs, but their average scores diverge significantly—92,466 for AMD versus 85,637 for NVIDIA—driven by the CMP 40HX’s weak Vulkan showing. The following analysis breaks down where each card excels, their architectural divergences, and what the numbers actually mean for specific workloads.

FAQ

Q: Which card has the higher average benchmark score?

A: The AMD Radeon Instinct MI60 averages 92,466 across its two benchmark tests, which is 6,829 points higher than the NVIDIA CMP 40HX’s average of 85,637. This gap exists despite the CMP 40HX winning one of the two individual tests.

Q: How do the two cards compare in OpenCL performance?

A: In Geekbench OpenCL, the NVIDIA CMP 40HX scores 93,395 versus 92,488 for the MI60, a 1% advantage for NVIDIA. This is the only test the CMP 40HX wins, and the margin is within typical run-to-run variance.

Q: What about Vulkan performance?

A: The AMD MI60 scores 92,444 in Geekbench Vulkan, while the CMP 40HX manages only 77,879. That translates to an 18.7% lead for AMD, a massive gap that defines the overall performance split.

Q: What are the nearest rivals for each card?

A: The MI60’s closest competitor is the NVIDIA RTX A4500, which trails by 0.9%, while the AMD Radeon Pro VII leads the MI60 by 4.8%. For the CMP 40HX, the AMD Radeon PRO W7600 leads by 1.7%, and the AMD Radeon PRO W6600 trails by 4.4%.

Q: Which card has more memory and bandwidth?

A: The MI60 features 32 GB of HBM2 on a 4096-bit bus, delivering 1.02 TB/s of bandwidth. The CMP 40HX has 8 GB of GDDR6 on a 256-bit bus, providing 448.0 GB/s—less than half the bandwidth and a quarter of the capacity.

Q: What is the launch MSRP of the NVIDIA CMP 40HX?

A: The CMP 40HX had a launch MSRP of 699 USD. The MI60 has no listed launch MSRP in the data.

Where Each One Wins

The benchmark results paint a clear split: the NVIDIA CMP 40HX wins in OpenCL workloads, while the AMD Radeon Instinct MI60 dominates in Vulkan. For OpenCL, the CMP 40HX’s 93,395 score edges out the MI60’s 92,488 by 1%. This is a narrow margin, but it positions the Turing card as marginally better for OpenCL compute tasks. The CMP 40HX also holds a significant advantage in raw memory clock speed—1750 MHz versus 1000 MHz for the MI60—and its 1650 MHz boost clock is 150 MHz lower than the MI60’s 1800 MHz, yet the OpenCL result favors NVIDIA.

The Vulkan test is where the MI60 asserts its dominance. Scoring 92,444, it beats the CMP 40HX’s 77,879 by 18.7%. This is not a close contest; it is a categorical victory for AMD’s architecture. The MI60’s FP32 throughput of 14.75 TFLOPS nearly doubles the CMP 40HX’s 7.603 TFLOPS, and its 460.8 GTexel/s texture rate far outpaces the NVIDIA card’s 237.6 GTexel/s. For any application leveraging Vulkan’s explicit control over GPU resources, the MI60 is the clear choice. The CMP 40HX’s Vulkan score is also its weakest result, 15,516 points below its OpenCL score, indicating a fundamental inefficiency in that API.

Architecture Differences

The two cards are built on different process nodes and microarchitectures. The AMD Radeon Instinct MI60 uses the Vega 20 chip on a 7nm TSMC process with GCN 5.1 architecture. It packs 13,230 million transistors into a 331 mm² die, yielding a transistor density of 40.0M per mm². The NVIDIA CMP 40HX uses the TU106 chip on a 12nm TSMC process with Turing architecture. It contains 10,800 million transistors on a larger 445 mm² die, resulting in a lower density of 24.3M per mm². The MI60’s smaller, denser die reflects its newer process node.

The MI60 has 4096 shading units, 256 TMUs, and 64 ROPs. The CMP 40HX has 2304 shading units, 144 TMUs, and 64 ROPs. This means the AMD card has 78% more shading units and 78% more TMUs, though both have identical ROP counts. The NVIDIA card does include 36 RT cores and 288 tensor cores, features absent from the MI60. These are Turing-specific additions for ray tracing and AI acceleration. However, the MI60’s raw compute throughput is far higher: 14.75 TFLOPS FP32 versus 7.603 TFLOPS for the CMP 40HX.

Memory architectures diverge completely. The MI60 uses 32 GB of HBM2 across a 4096-bit bus, achieving 1.02 TB/s bandwidth. The CMP 40HX uses 8 GB of GDDR6 on a 256-bit bus, with 448.0 GB/s bandwidth. The MI60’s memory bandwidth is 2.3 times higher, and its capacity is four times larger. The CMP 40HX compensates with a faster effective memory speed of 14 Gbps versus 2 Gbps for the MI60, but the narrow bus limits overall throughput.

Specification Differences

The most striking difference is memory: the MI60 offers 32 GB HBM2 with 1.02 TB/s bandwidth, while the CMP 40HX offers 8 GB GDDR6 with 448.0 GB/s. The bus width is 4096-bit for AMD versus 256-bit for NVIDIA. Shading units differ at 4096 versus 2304, and TMUs at 256 versus 144. ROPs are equal at 64. The MI60 has higher clocks: 1200 MHz base and 1800 MHz boost versus 1470 MHz base and 1650 MHz boost for the CMP 40HX. However, the NVIDIA card’s memory runs at 1750 MHz (14 Gbps effective) versus 1000 MHz (2 Gbps effective) for the AMD card.

Power requirements differ significantly. The MI60 has a 300 W TDP and requires a 700 W PSU with 1x 6-pin and 1x 8-pin connectors. The CMP 40HX has a 185 W TDP, a 450 W PSU recommendation, and a single 8-pin connector. The MI60 is longer at 267 mm versus 229 mm for the CMP 40HX, and it has one mini-DisplayPort 1.4a output, while the NVIDIA card has no display outputs at all—a deliberate design for mining. The bus interface also differs: PCIe 4.0 x16 for the MI60 versus PCIe 1.0 x4 for the CMP 40HX, a severe bottleneck for the NVIDIA card. API support favors NVIDIA with DirectX 12 Ultimate (12_2) versus DirectX 12 (12_1) for AMD, and Vulkan 1.4 versus 1.3.

Head-to-Head Benchmarks

The head-to-head results show one win apiece, but the magnitude differs. In Geekbench OpenCL, the NVIDIA CMP 40HX scores 93,395 against the MI60’s 92,488, a 1% delta in NVIDIA’s favor. This is the closest result between the two cards. The CMP 40HX’s Turing architecture, with its 288 tensor cores, likely contributes to this OpenCL advantage. However, the MI60 is not far behind, and its 14.75 TFLOPS FP32 throughput suggests it should outperform in compute-heavy OpenCL tasks. The 1% gap is within noise for most real-world applications.

The Vulkan test is decisive. The AMD MI60 scores 92,444, while the CMP 40HX scores 77,879. The 18.7% delta is the largest margin between the two cards in either test. This result aligns with the MI60’s architectural strengths: higher shading unit count, greater texture rate (460.8 GTexel/s versus 237.6 GTexel/s), and significantly higher memory bandwidth. The CMP 40HX’s PCIe 1.0 x4 interface may also hamper Vulkan performance, as the API’s lower-level access to system memory can expose bus limitations.

Looking at the average scores, the MI60’s 92,466 average is 6,829 points higher than the CMP 40HX’s 85,637. This is entirely due to the Vulkan result; without it, the OpenCL scores are nearly identical. The MI60’s consistency across APIs—92,488 OpenCL and 92,444 Vulkan—shows balanced performance. The CMP 40HX’s scores are inconsistent, with Vulkan trailing OpenCL by 15,516 points. This inconsistency is a critical differentiator for buyers.

For nearest rivals, the MI60’s 0.9% lead over the RTX A4500 in average score places it in competitive territory among professional cards. The CMP 40HX’s 1.7% deficit to the Radeon PRO W7600 and its 4.4% lead over the Radeon PRO W6600 show it sits mid-pack. Despite both cards sharing the 93rd percentile, the MI60’s average score is 7.9% higher than the CMP 40HX’s, a meaningful gap in overall performance. The data suggests the MI60 is the more versatile accelerator, while the CMP 40HX is optimized for a narrower set of tasks.

DETAILED SPECIFICATIONS

SPECIFICATION
Instinct MI60
CMP 40HX
Core Specs
Shading Units
4,096
2,304 -43.8%
Shaders
4,096
2,304 -43.8%
TMUs
256
144 -43.8%
ROPs
64
64 0.0%
Compute Units
64
—
SM Count
—
36
Clocks
Base Clock
1200 MHz
1470 MHz
Boost Clock
1800 MHz
1650 MHz
Memory Clock
1000 MHz 2 Gbps effective
1750 MHz 14 Gbps effective
Memory
Memory Size
32 GB
8 GB
VRAM (MB)
32,768
8,192 -75.0%
Memory Type
HBM2
GDDR6
Memory Bus
4096 bit
256 bit
Bandwidth
1.02 TB/s
448.0 GB/s
Cache
L1 Cache
16 KB (per CU)
64 KB (per SM)
L2 Cache
4 MB
4 MB
Performance
Pixel Rate
115.2 GPixel/s
105.6 GPixel/s
Texture Rate
460.8 GTexel/s
237.6 GTexel/s
FP32 (TFLOPS)
14.75 TFLOPS
7.603 TFLOPS
FP64 (TFLOPS)
7.373 TFLOPS (1:2)
237.6 GFLOPS (1:32)
FP16 (TFLOPS)
29.49 TFLOPS (2:1)
15.21 TFLOPS (2:1)
AI/RT
RT Cores
—
36
Tensor Cores
—
288
Power
TDP
300 W
185 W
TDP (W)
300
185 -38.3%
Suggested PSU
700 W
450 W
Power Connectors
1x 6-pin + 1x 8-pin
1x 8-pin
Architecture
Architecture
GCN 5.1
Turing
GPU Name
Vega 20
TU106
Generation
Radeon Instinct (MIx)
Mining GPUs
Process Size
7 nm
12 nm
Transistors
13,230 million
10,800 million
Die Size
331 mm²
445 mm²
Foundry
TSMC
TSMC
Density
40.0M / mm²
24.3M / mm²
API Support
DirectX
12 (12_1)
12 Ultimate (12_2)
OpenGL
4.6
4.6
Vulkan
1.3
1.4
OpenCL
2.1
3.0
CUDA
—
7.5
Shader Model
6.7
6.8
Physical
Slot Width
Dual-slot
Dual-slot
Length
267 mm 10.5 inches
229 mm 9 inches
Height
111 mm 4.4 inches
111 mm 4.4 inches
Outputs
1x mini-DisplayPort 1.4a
No outputs
Bus Interface
PCIe 4.0 x16
PCIe 1.0 x4
Other
Launch Price
—
699 USD
Production
End-of-life
End-of-life
Predecessor
FirePro Data Center
—
View Radeon Instinct MI60 Details View CMP 40HX Details