AMD Radeon Instinct MI60 vs NVIDIA CMP 40HX Comparison
AMD Radeon Instinct MI60
CMP 40HX
PERFORMANCE BENCHMARKS
Analysis: AMD Radeon Instinct MI60 vs NVIDIA CMP 40HX
The AMD Radeon Instinct MI60 and NVIDIA CMP 40HX are both end-of-life accelerators, yet they occupy opposite ends of the hardware spectrum. The MI60 is a 7nm data-center compute card with 32 GB of HBM2, while the CMP 40HX is a 12nm Turing-based mining card with 8 GB of GDDR6. Benchmark data shows a razor-thin split: the NVIDIA card wins OpenCL by 1%, while the AMD card dominates Vulkan by 18.7%. Both sit at the 93rd percentile among all GPUs, but their average scores diverge significantly—92,466 for AMD versus 85,637 for NVIDIA—driven by the CMP 40HX’s weak Vulkan showing. The following analysis breaks down where each card excels, their architectural divergences, and what the numbers actually mean for specific workloads.
FAQ
Q: Which card has the higher average benchmark score?
A: The AMD Radeon Instinct MI60 averages 92,466 across its two benchmark tests, which is 6,829 points higher than the NVIDIA CMP 40HX’s average of 85,637. This gap exists despite the CMP 40HX winning one of the two individual tests.
Q: How do the two cards compare in OpenCL performance?
A: In Geekbench OpenCL, the NVIDIA CMP 40HX scores 93,395 versus 92,488 for the MI60, a 1% advantage for NVIDIA. This is the only test the CMP 40HX wins, and the margin is within typical run-to-run variance.
Q: What about Vulkan performance?
A: The AMD MI60 scores 92,444 in Geekbench Vulkan, while the CMP 40HX manages only 77,879. That translates to an 18.7% lead for AMD, a massive gap that defines the overall performance split.
Q: What are the nearest rivals for each card?
A: The MI60’s closest competitor is the NVIDIA RTX A4500, which trails by 0.9%, while the AMD Radeon Pro VII leads the MI60 by 4.8%. For the CMP 40HX, the AMD Radeon PRO W7600 leads by 1.7%, and the AMD Radeon PRO W6600 trails by 4.4%.
Q: Which card has more memory and bandwidth?
A: The MI60 features 32 GB of HBM2 on a 4096-bit bus, delivering 1.02 TB/s of bandwidth. The CMP 40HX has 8 GB of GDDR6 on a 256-bit bus, providing 448.0 GB/s—less than half the bandwidth and a quarter of the capacity.
Q: What is the launch MSRP of the NVIDIA CMP 40HX?
A: The CMP 40HX had a launch MSRP of 699 USD. The MI60 has no listed launch MSRP in the data.
Where Each One Wins
The benchmark results paint a clear split: the NVIDIA CMP 40HX wins in OpenCL workloads, while the AMD Radeon Instinct MI60 dominates in Vulkan. For OpenCL, the CMP 40HX’s 93,395 score edges out the MI60’s 92,488 by 1%. This is a narrow margin, but it positions the Turing card as marginally better for OpenCL compute tasks. The CMP 40HX also holds a significant advantage in raw memory clock speed—1750 MHz versus 1000 MHz for the MI60—and its 1650 MHz boost clock is 150 MHz lower than the MI60’s 1800 MHz, yet the OpenCL result favors NVIDIA.
The Vulkan test is where the MI60 asserts its dominance. Scoring 92,444, it beats the CMP 40HX’s 77,879 by 18.7%. This is not a close contest; it is a categorical victory for AMD’s architecture. The MI60’s FP32 throughput of 14.75 TFLOPS nearly doubles the CMP 40HX’s 7.603 TFLOPS, and its 460.8 GTexel/s texture rate far outpaces the NVIDIA card’s 237.6 GTexel/s. For any application leveraging Vulkan’s explicit control over GPU resources, the MI60 is the clear choice. The CMP 40HX’s Vulkan score is also its weakest result, 15,516 points below its OpenCL score, indicating a fundamental inefficiency in that API.
Architecture Differences
The two cards are built on different process nodes and microarchitectures. The AMD Radeon Instinct MI60 uses the Vega 20 chip on a 7nm TSMC process with GCN 5.1 architecture. It packs 13,230 million transistors into a 331 mm² die, yielding a transistor density of 40.0M per mm². The NVIDIA CMP 40HX uses the TU106 chip on a 12nm TSMC process with Turing architecture. It contains 10,800 million transistors on a larger 445 mm² die, resulting in a lower density of 24.3M per mm². The MI60’s smaller, denser die reflects its newer process node.
The MI60 has 4096 shading units, 256 TMUs, and 64 ROPs. The CMP 40HX has 2304 shading units, 144 TMUs, and 64 ROPs. This means the AMD card has 78% more shading units and 78% more TMUs, though both have identical ROP counts. The NVIDIA card does include 36 RT cores and 288 tensor cores, features absent from the MI60. These are Turing-specific additions for ray tracing and AI acceleration. However, the MI60’s raw compute throughput is far higher: 14.75 TFLOPS FP32 versus 7.603 TFLOPS for the CMP 40HX.
Memory architectures diverge completely. The MI60 uses 32 GB of HBM2 across a 4096-bit bus, achieving 1.02 TB/s bandwidth. The CMP 40HX uses 8 GB of GDDR6 on a 256-bit bus, with 448.0 GB/s bandwidth. The MI60’s memory bandwidth is 2.3 times higher, and its capacity is four times larger. The CMP 40HX compensates with a faster effective memory speed of 14 Gbps versus 2 Gbps for the MI60, but the narrow bus limits overall throughput.
Specification Differences
The most striking difference is memory: the MI60 offers 32 GB HBM2 with 1.02 TB/s bandwidth, while the CMP 40HX offers 8 GB GDDR6 with 448.0 GB/s. The bus width is 4096-bit for AMD versus 256-bit for NVIDIA. Shading units differ at 4096 versus 2304, and TMUs at 256 versus 144. ROPs are equal at 64. The MI60 has higher clocks: 1200 MHz base and 1800 MHz boost versus 1470 MHz base and 1650 MHz boost for the CMP 40HX. However, the NVIDIA card’s memory runs at 1750 MHz (14 Gbps effective) versus 1000 MHz (2 Gbps effective) for the AMD card.
Power requirements differ significantly. The MI60 has a 300 W TDP and requires a 700 W PSU with 1x 6-pin and 1x 8-pin connectors. The CMP 40HX has a 185 W TDP, a 450 W PSU recommendation, and a single 8-pin connector. The MI60 is longer at 267 mm versus 229 mm for the CMP 40HX, and it has one mini-DisplayPort 1.4a output, while the NVIDIA card has no display outputs at all—a deliberate design for mining. The bus interface also differs: PCIe 4.0 x16 for the MI60 versus PCIe 1.0 x4 for the CMP 40HX, a severe bottleneck for the NVIDIA card. API support favors NVIDIA with DirectX 12 Ultimate (12_2) versus DirectX 12 (12_1) for AMD, and Vulkan 1.4 versus 1.3.
Head-to-Head Benchmarks
The head-to-head results show one win apiece, but the magnitude differs. In Geekbench OpenCL, the NVIDIA CMP 40HX scores 93,395 against the MI60’s 92,488, a 1% delta in NVIDIA’s favor. This is the closest result between the two cards. The CMP 40HX’s Turing architecture, with its 288 tensor cores, likely contributes to this OpenCL advantage. However, the MI60 is not far behind, and its 14.75 TFLOPS FP32 throughput suggests it should outperform in compute-heavy OpenCL tasks. The 1% gap is within noise for most real-world applications.
The Vulkan test is decisive. The AMD MI60 scores 92,444, while the CMP 40HX scores 77,879. The 18.7% delta is the largest margin between the two cards in either test. This result aligns with the MI60’s architectural strengths: higher shading unit count, greater texture rate (460.8 GTexel/s versus 237.6 GTexel/s), and significantly higher memory bandwidth. The CMP 40HX’s PCIe 1.0 x4 interface may also hamper Vulkan performance, as the API’s lower-level access to system memory can expose bus limitations.
Looking at the average scores, the MI60’s 92,466 average is 6,829 points higher than the CMP 40HX’s 85,637. This is entirely due to the Vulkan result; without it, the OpenCL scores are nearly identical. The MI60’s consistency across APIs—92,488 OpenCL and 92,444 Vulkan—shows balanced performance. The CMP 40HX’s scores are inconsistent, with Vulkan trailing OpenCL by 15,516 points. This inconsistency is a critical differentiator for buyers.
For nearest rivals, the MI60’s 0.9% lead over the RTX A4500 in average score places it in competitive territory among professional cards. The CMP 40HX’s 1.7% deficit to the Radeon PRO W7600 and its 4.4% lead over the Radeon PRO W6600 show it sits mid-pack. Despite both cards sharing the 93rd percentile, the MI60’s average score is 7.9% higher than the CMP 40HX’s, a meaningful gap in overall performance. The data suggests the MI60 is the more versatile accelerator, while the CMP 40HX is optimized for a narrower set of tasks.