AMD Radeon Pro Vega 64 vs NVIDIA CMP 40HX Comparison

AMD
RADEON

AMD Radeon Pro Vega 64

CORE STATE Vega 10
VRAM 16 GB
CLOCK SPEED 1350 MHz
TDP 250 W
BUS WIDTH 2048 bit
ARCHITECTURE GCN 5.0
nm
PROCESS 14 nm
LAUNCH DATE 2017
VS
NVIDIA
GEFORCE

CMP 40HX

CORE STATE TU106
VRAM 8 GB
CLOCK SPEED 1650 MHz
TDP 185 W
BUS WIDTH 256 bit
ARCHITECTURE Turing
nm
PROCESS 12 nm
LAUNCH DATE 2021

PERFORMANCE BENCHMARKS

geekbench_metal
71,868
N/A
geekbench_opencl
71,094
93,395
geekbench_vulkan
74,174
77,879

Analysis: AMD Radeon Pro Vega 64 vs NVIDIA CMP 40HX

NVIDIA CMP 40HX and AMD Radeon Pro Vega 64 occupy very different positions in the GPU landscape, despite both being end-of-life products. The CMP 40HX is a Turing-based mining card with no display outputs, while the Radeon Pro Vega 64 is a GCN 5.0 workstation part designed for Mac integration. Benchmark data shows the NVIDIA part leading in both shared tests, but the AMD card counters with double the memory and a wider compute pipeline. Below is a breakdown of what the numbers actually mean for each card.

FAQ

Q: Which card wins in Geekbench OpenCL performance?

A: The NVIDIA CMP 40HX scores 93,395 versus the AMD Radeon Pro Vega 64's 71,094, giving NVIDIA a 31.4% lead in this test. This is the largest performance gap between the two cards in any benchmark.

Q: How do the cards compare in Vulkan compute?

A: The NVIDIA CMP 40HX again takes the win, scoring 77,879 against the Radeon Pro Vega 64's 74,174. The margin is much narrower here, with NVIDIA ahead by just 5%.

Q: What is the average benchmark score for each card?

A: The NVIDIA CMP 40HX averages 85,637 across its benchmark results, while the AMD Radeon Pro Vega 64 averages 72,379. This places the NVIDIA card in the 93rd percentile of all GPUs, versus the 91st percentile for the AMD card.

Q: Which card has more memory and what type is it?

A: The AMD Radeon Pro Vega 64 has 16 GB of HBM2 memory on a 2048-bit bus, while the NVIDIA CMP 40HX has 8 GB of GDDR6 on a 256-bit bus. Despite the smaller capacity, the NVIDIA card achieves higher memory bandwidth at 448.0 GB/s versus 402.4 GB/s for AMD.

Q: Are there any benchmarks where the AMD card wins?

A: The head-to-head data shows zero wins for the AMD Radeon Pro Vega 64. NVIDIA wins both shared tests (OpenCL and Vulkan), although the AMD card does have a Geekbench Metal score of 71,868, which is not directly comparable since the NVIDIA card lacks a Metal result.

Q: How does the Radeon Pro Vega 64 compare to its closest rivals?

A: The AMD card sits within 2.2% of the AMD Radeon RX 6600 LE (which scores 70,829) and is 0.4% behind the NVIDIA TITAN X Pascal (72,098). It is 1.4% behind the AMD Radeon Vega Frontier Edition (73,370), showing it is tightly clustered with those parts.

The Verdict

The data clearly favors the NVIDIA CMP 40HX for raw compute throughput. It leads by 31.4% in OpenCL and by 5% in Vulkan, and its average benchmark score of 85,637 is 18.3% higher than the AMD card's 72,379. For anyone prioritizing compute performance in OpenCL-heavy workloads, the CMP 40HX is the stronger choice based on these results.

However, the AMD Radeon Pro Vega 64 has its own advantages that the benchmark scores do not capture. It offers 16 GB of HBM2 memory, which is double the CMP 40HX's 8 GB, and its 2048-bit bus width is eight times larger than NVIDIA's 256-bit interface. This makes the AMD card more suitable for memory-capacity-sensitive tasks, even if its bandwidth is slightly lower (402.4 GB/s versus 448.0 GB/s).

The Radeon Pro Vega 64 is also the only one of the two with display outputs (classified as "Portable Device Dependent") and is designed as an integrated graphics processor (IGP) for Mac systems. The CMP 40HX has no display outputs at all, making it useless for any visual output. For users who need a functional workstation GPU with display capability, the AMD card is the only option.

In terms of positioning, the CMP 40HX sits at the 93rd percentile of all GPUs, whereas the Radeon Pro Vega 64 sits at the 91st. The NVIDIA card's nearest rivals include the AMD Radeon PRO W7600 (87,108, 1.7% behind) and the NVIDIA Quadro GP100 (87,445, 2.1% behind), while the AMD card competes with the TITAN X Pascal and Vega Frontier Edition. Neither card is a clear winner across all criteria; the choice depends on whether compute speed or memory capacity and display support matter more.

Head-to-Head Benchmarks

The most decisive result comes from Geekbench OpenCL, where the NVIDIA CMP 40HX posts 93,395 against the AMD Radeon Pro Vega 64's 71,094. This 31.4% delta is substantial and indicates a major advantage in OpenCL compute workloads. The CMP 40HX's score also exceeds its own average benchmark score of 85,637, suggesting OpenCL is a strong suite for this card.

In Geekbench Vulkan, the gap narrows considerably. The NVIDIA card scores 77,879 versus 74,174 for AMD, a 5% difference. Both cards perform above their respective averages in this test — the CMP 40HX's Vulkan score is 9.1% below its OpenCL score, while the Radeon Pro Vega 64's Vulkan result is 4.3% above its OpenCL score, indicating the AMD card is relatively stronger in Vulkan than in OpenCL.

The AMD card also has a Geekbench Metal score of 71,868, which is its second-best result after Vulkan. However, because the NVIDIA card has no Metal benchmark listed, no direct comparison can be made. Across the two shared tests, NVIDIA wins both, but the magnitude of the OpenCL win (31.4%) dwarfs the Vulkan margin (5%). The data suggests that if a workload is Vulkan-based, the two cards perform similarly, but OpenCL workloads will heavily favor the CMP 40HX.

Looking at the broader context, the CMP 40HX's average score of 85,637 places it 18.3% above the Radeon Pro Vega 64's 72,379 average. This aligns with the percentile rankings: 93rd versus 91st. The NVIDIA card's nearest rival, the AMD Radeon PRO W7600, is only 1.7% behind at 87,108, while the AMD card's closest competitor, the TITAN X Pascal, is 0.4% ahead at 72,098. This shows both cards are competitive within their respective performance tiers.

Specification Differences

The most obvious difference is memory: the AMD Radeon Pro Vega 64 has 16 GB of HBM2 on a 2048-bit bus, while the NVIDIA CMP 40HX has 8 GB of GDDR6 on a 256-bit bus. Despite the smaller capacity, NVIDIA's memory runs at 1750 MHz (14 Gbps effective), achieving 448.0 GB/s bandwidth, whereas AMD's HBM2 runs at 786 MHz (1572 Mbps effective), yielding 402.4 GB/s. So the CMP 40HX has higher bandwidth but half the capacity.

Clock speeds also differ. The NVIDIA card has a base clock of 1470 MHz and a boost clock of 1650 MHz, while the AMD card runs at 1250 MHz base and 1350 MHz boost. The CMP 40HX is therefore clocked 220 MHz higher at base and 300 MHz higher at boost. Power consumption reflects this: the CMP 40HX draws 185 W with a 450 W suggested PSU, while the Radeon Pro Vega 64 draws 250 W and has no suggested PSU listed.

Form factor and connectivity diverge sharply. The NVIDIA card is a dual-slot, 229 mm long (9 inches) card requiring a 1x 8-pin power connector, with a PCIe 1.0 x4 bus interface. The AMD card is an IGP (integrated graphics processor) with no power connectors and a PCIe 3.0 x16 interface. The NVIDIA card has no display outputs; the AMD card's outputs are "Portable Device Dependent." The AMD card has no listed dimensions, while the NVIDIA card measures 111 mm (4.4 inches) in height and 35 mm (1.4 inches) in width.

The AMD card has significantly more shading units (4096 versus 2304) and texture mapping units (256 versus 144), but both have 64 ROPs. This gives the AMD card a higher texture rate of 345.6 GTexel/s versus 237.6 GTexel/s for NVIDIA, although NVIDIA's pixel rate is higher at 105.6 GPixel/s versus 86.40 GPixel/s.

Architecture Differences

The NVIDIA CMP 40HX is built on the TU106 chip using Turing architecture on a 12 nm process from TSMC, with 10,800 million transistors on a 445 mm² die. The AMD Radeon Pro Vega 64 uses the Vega 10 chip with GCN 5.0 architecture on a 14 nm process from GlobalFoundries, packing 12,500 million transistors on a 495 mm² die. AMD's chip is larger and has more transistors, but the transistor density is similar: 25.3M per mm² for AMD versus 24.3M per mm² for NVIDIA.

Compute capabilities differ significantly. The AMD card has 4096 shading units and delivers 11.06 TFLOPS FP32 and 22.12 TFLOPS FP16 (2:1 ratio). The NVIDIA card has 2304 shading units and delivers 7.603 TFLOPS FP32 and 15.21 TFLOPS FP16 (2:1 ratio). Despite having fewer shaders, the NVIDIA card's higher clocks partially compensate, but AMD still leads in raw FP32 throughput by 45.5%.

Feature sets are where the architectures diverge most. The NVIDIA CMP 40HX includes 36 RT cores and 288 tensor cores, which are absent from the AMD card entirely. This means the NVIDIA card supports hardware-accelerated ray tracing and tensor operations, though these features are irrelevant for a mining card with no display outputs. In terms of API support, the NVIDIA card supports DirectX 12 Ultimate (12_2) and Vulkan 1.4, while the AMD card supports DirectX 12 (12_1) and Vulkan 1.3. Both support OpenGL 4.6.

The AMD card's GCN 5.0 architecture is older, dating from its 2017 release, while the NVIDIA card is from a 2021 release in the Mining GPUs generation. This generational gap explains the RT and tensor core presence on NVIDIA and the newer DirectX 12 Ultimate support. However, the AMD card's larger memory pool and wider bus (2048-bit versus 256-bit) are architectural choices tailored for high-bandwidth applications like professional visualization in Mac environments, whereas the CMP 40HX is stripped of display functionality for mining use. The NVIDIA card's PCIe 1.0 x4 interface is notably older than the AMD card's PCIe 3.0 x16, likely reflecting the mining card's focus on compute rather than data transfer.

DETAILED SPECIFICATIONS

SPECIFICATION
Pro Vega 64
CMP 40HX
Core Specs
Shading Units
4,096
2,304 -43.8%
Shaders
4,096
2,304 -43.8%
TMUs
256
144 -43.8%
ROPs
64
64 0.0%
Compute Units
64
SM Count
36
Clocks
Base Clock
1250 MHz
1470 MHz
Boost Clock
1350 MHz
1650 MHz
Memory Clock
786 MHz 1572 Mbps effective
1750 MHz 14 Gbps effective
Memory
Memory Size
16 GB
8 GB
VRAM (MB)
16,384
8,192 -50.0%
Memory Type
HBM2
GDDR6
Memory Bus
2048 bit
256 bit
Bandwidth
402.4 GB/s
448.0 GB/s
Cache
L1 Cache
16 KB (per CU)
64 KB (per SM)
L2 Cache
4 MB
4 MB
Performance
Pixel Rate
86.40 GPixel/s
105.6 GPixel/s
Texture Rate
345.6 GTexel/s
237.6 GTexel/s
FP32 (TFLOPS)
11.06 TFLOPS
7.603 TFLOPS
FP64 (TFLOPS)
691.2 GFLOPS (1:16)
237.6 GFLOPS (1:32)
FP16 (TFLOPS)
22.12 TFLOPS (2:1)
15.21 TFLOPS (2:1)
AI/RT
RT Cores
36
Tensor Cores
288
Power
TDP
250 W
185 W
TDP (W)
250
185 -26.0%
Suggested PSU
450 W
Power Connectors
None
1x 8-pin
Architecture
Architecture
GCN 5.0
Turing
GPU Name
Vega 10
TU106
Generation
Radeon Pro Mac (Vega Series)
Mining GPUs
Process Size
14 nm
12 nm
Transistors
12,500 million
10,800 million
Die Size
495 mm²
445 mm²
Foundry
GlobalFoundries
TSMC
Density
25.3M / mm²
24.3M / mm²
API Support
DirectX
12 (12_1)
12 Ultimate (12_2)
OpenGL
4.6
4.6
Vulkan
1.3
1.4
OpenCL
2.1
3.0
CUDA
7.5
Shader Model
6.7
6.8
Physical
Slot Width
IGP
Dual-slot
Length
229 mm 9 inches
Height
111 mm 4.4 inches
Outputs
Portable Device Dependent
No outputs
Bus Interface
PCIe 3.0 x16
PCIe 1.0 x4
Other
Launch Price
699 USD
Production
End-of-life
End-of-life
View Radeon Pro Vega 64 Details View CMP 40HX Details