AMD Radeon Pro Vega 56 vs NVIDIA CMP 40HX Comparison

AMD
RADEON

AMD Radeon Pro Vega 56

CORE STATE Vega 10
VRAM 8 GB
CLOCK SPEED 1250 MHz
TDP 210 W
BUS WIDTH 2048 bit
ARCHITECTURE GCN 5.0
nm
PROCESS 14 nm
LAUNCH DATE 2017
VS
NVIDIA
GEFORCE

CMP 40HX

CORE STATE TU106
VRAM 8 GB
CLOCK SPEED 1650 MHz
TDP 185 W
BUS WIDTH 256 bit
ARCHITECTURE Turing
nm
PROCESS 12 nm
LAUNCH DATE 2021

PERFORMANCE BENCHMARKS

geekbench_metal
63,145
N/A
geekbench_opencl
61,930
93,395
geekbench_vulkan
66,004
77,879

Analysis: AMD Radeon Pro Vega 56 vs NVIDIA CMP 40HX

NVIDIA CMP 40HX and AMD Radeon Pro Vega 56 are two very different products that happen to sit close in the current database rankings. The CMP 40HX, a Turing-based mining card from 2021, holds an average benchmark score of 85,637, placing it in the 93rd percentile of all GPUs. The Radeon Pro Vega 56, a GCN 5.0 workstation part from 2017, averages 63,693 and sits in the 89th percentile. The gap in average score is roughly 34%, a substantial margin. However, the nature of each card, one with no display outputs and another designed for Mac Pro integration, means the comparison is as much about capability and context as raw performance.

Head-to-Head Benchmarks

The database records two shared benchmark tests for these GPUs, and the NVIDIA CMP 40HX wins both. In Geekbench OpenCL, the CMP 40HX scores 93,395 against the Radeon Pro Vega 56's 61,930. That is a delta of 50.8%, meaning the NVIDIA card delivers over half again as much compute throughput in this workload. This is a decisive victory, and the margin is far larger than the average score gap would suggest. In Geekbench Vulkan, the CMP 40HX again takes the win, scoring 77,879 versus 66,004, a delta of 18%. While still a clear win, this is a much tighter contest than the OpenCL result, indicating the AMD card is comparatively stronger under Vulkan than under OpenCL.

The overall head-to-head record is 2 wins for NVIDIA and 0 for AMD. The NVIDIA card's average score of 85,637 is also well ahead of the AMD card's 63,693. But the database's nearest rival data adds nuance. The CMP 40HX's closest competitor is the AMD Radeon PRO W7600, with an average score of 87,108, putting the NVIDIA card just 1.7% behind. The NVIDIA Quadro GP100 is similar, averaging 87,445, with the CMP 40HX trailing by 2.1%. Meanwhile, the CMP 40HX holds a 4.4% lead over the AMD Radeon PRO W6600 and a 5.8% lead over the AMD Radeon Pro Vega 64X. So, while it loses to the top of its rival group, it beats the rest by a meaningful margin.

The Radeon Pro Vega 56's rival group is clustered much tighter. Its nearest rival, the AMD Radeon RX 7600M, averages 63,775, a deficit of just 0.1%. The AMD Radeon RX 9060 XT LP and NVIDIA CMP 30HX are both essentially tied with it, at 63,830 and 63,842 respectively, with deltas of -0.2%. Even the AMD Radeon Pro WX 9100, the top of its rival group at 64,212, is only 0.8% ahead. This shows the Radeon Pro Vega 56 is in a highly competitive performance band, where any of these cards could trade places depending on the specific workload. The CMP 40HX, by contrast, has a more spread-out rival field, with its nearest competitors either slightly ahead or several points behind.

The recorded data also includes a Geekbench Metal score for the Radeon Pro Vega 56 of 63,145, a test the NVIDIA card cannot run due to its lack of display outputs and different driver ecosystem. This is worth noting because it shows the AMD card has a dedicated, high-scoring path in Apple's compute framework. The CMP 40HX has no corresponding Metal result in the database.

Where Each One Wins

The CMP 40HX wins in raw compute density per watt in the shared tests. Its OpenCL score of 93,395 at a TDP of 185 W gives it a far better performance-per-watt figure than the Radeon Pro Vega 56's 61,930 at 210 W. The NVIDIA card also wins in Vulkan, scoring 77,879 versus 66,004, so for any cross-platform compute or rendering workload using these APIs, the choice is clear. The CMP 40HX also has a higher pixel rate, 105.6 GPixel/s versus 80.00 GPixel/s, which indicates faster fill-rate-bound operations. Its memory bandwidth of 448.0 GB/s is also higher than the AMD card's 402.4 GB/s, giving it an edge in bandwidth-sensitive tasks.

The Radeon Pro Vega 56 wins in areas that the CMP 40HX cannot compete in at all. It has display outputs, specifically 1x HDMI 2.0b and 3x DisplayPort 1.4a, meaning it can drive monitors. The CMP 40HX has no outputs, making it useless for any interactive or display-oriented task. The AMD card also has a higher texture rate, 280.0 GTexel/s versus 237.6 GTexel/s, so in pure texture-fetch-bound workloads it is faster. Its FP32 throughput is higher as well, 8.960 TFLOPS versus 7.603 TFLOPS, and its FP16 rate of 17.92 TFLOPS beats the NVIDIA card's 15.21 TFLOPS. These raw shader throughput numbers suggest the AMD card is not simply a lesser product; it is a different kind of product, optimized for certain compute patterns.

The Radeon Pro Vega 56 also has a higher transistor count, 12,500 million, and a larger die, 495 mm², compared to the CMP 40HX's 10,800 million transistors on 445 mm². Its memory bus is vastly wider at 2048 bit versus 256 bit, though the NVIDIA card's GDDR6 at 14 Gbps effective compensates with higher overall bandwidth. In power terms, the AMD card draws 210 W compared to 185 W for the NVIDIA card, so the CMP 40HX is more efficient in the shared benchmarks. There is no suggested PSU for the AMD card, while the NVIDIA card lists 450 W. The NVIDIA card uses a single 8-pin power connector and fits in a dual-slot layout; the AMD card is listed as an IGP, meaning it is integrated into a system board or chassis, with no power connectors of its own.

The Verdict

If the workload is compute-only, using OpenCL or Vulkan, the data points squarely to the NVIDIA CMP 40HX. It is 50.8% ahead in OpenCL and 18% ahead in Vulkan, with a higher average score, higher pixel rate, and higher memory bandwidth. It also does so at a lower TDP, 185 W versus 210 W, and its launch MSRP was 699 USD. For a mining or compute farm where display output is irrelevant, this card is the stronger choice on every measured metric.

If the workload requires a display, or if the system depends on Apple's Metal API, the Radeon Pro Vega 56 is the only option between the two. It has HDMI and DisplayPort outputs, a Metal score of 63,145, and a higher texture rate and FP32 throughput. It also uses HBM2 memory with a 2048-bit bus, which may be preferable for certain memory-latency-sensitive tasks, even though its total bandwidth is lower. The AMD card is also an IGP, so it can be integrated into a Mac Pro environment where the CMP 40HX cannot exist at all.

For a builder assembling a dedicated compute node with no monitors attached, the CMP 40HX is the clear winner. For anyone needing a functioning graphics card with outputs, or working within a Mac ecosystem, the Radeon Pro Vega 56 is the only viable pick. The 34% average score gap is real, but it is irrelevant if the card cannot do the job required.

FAQ

Q: Which card has a higher average benchmark score?

A: The NVIDIA CMP 40HX, with an average score of 85,637, compared to the AMD Radeon Pro Vega 56's 63,693.

Q: What is the largest performance gap between the two cards?

A: In Geekbench OpenCL, the CMP 40HX leads by 50.8%, scoring 93,395 versus 61,930.

Q: Can the NVIDIA CMP 40HX output video to a monitor?

A: No, it has no display outputs. The AMD Radeon Pro Vega 56 has 1x HDMI 2.0b and 3x DisplayPort 1.4a.

Q: Does the Radeon Pro Vega 56 support any API that the CMP 40HX lacks?

A: The database records a Geekbench Metal score of 63,145 for the AMD card, while no Metal result is listed for the NVIDIA card.

Q: Which card is more efficient in terms of TDP?

A: The NVIDIA CMP 40HX has a TDP of 185 W, while the AMD Radeon Pro Vega 56 has a TDP of 210 W.

Q: How does the CMP 40HX compare to its nearest rival, the AMD Radeon PRO W7600?

A: The CMP 40HX trails the PRO W7600 by 1.7%, with average scores of 85,637 and 87,108 respectively.

Architecture Differences

The NVIDIA CMP 40HX is built on the Turing architecture using the TU106 chip, manufactured on a 12 nm process at TSMC. The AMD Radeon Pro Vega 56 uses the GCN 5.0 architecture with the Vega 10 chip, built on a 14 nm process at GlobalFoundries. The NVIDIA card has 10,800 million transistors on a 445 mm² die, giving a transistor density of 24.3M per mm². The AMD card has 12,500 million transistors on a 495 mm² die, for a density of 25.3M per mm². The AMD die is larger and denser, but it is also on an older process node.

In terms of compute resources, the CMP 40HX has 2,304 shading units, 144 texture mapping units, and 64 ROPs. It also includes 36 ray tracing cores and 288 tensor cores, which are Turing-specific features. The Radeon Pro Vega 56 has 3,584 shading units, 224 TMUs, and 64 ROPs, but it has no ray tracing cores and no tensor cores. The AMD card has more raw shader hardware, which explains its higher FP32 and FP16 throughput, but it lacks the specialized acceleration blocks found in the NVIDIA chip.

Memory architecture differs fundamentally. The CMP 40HX uses 8 GB of GDDR6 on a 256-bit bus, with a memory clock of 1750 MHz and 14 Gbps effective, yielding 448.0 GB/s of bandwidth. The Radeon Pro Vega 56 uses 8 GB of HBM2 on a 2048-bit bus, with a memory clock of 786 MHz and 1572 Mbps effective, yielding 402.4 GB/s. The NVIDIA card has higher bandwidth but a much narrower bus; the AMD card has a massive bus width but lower effective speed and lower total bandwidth.

The CMP 40HX supports DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4. The Radeon Pro Vega 56 supports DirectX 12 (12_1), OpenGL 4.6, and Vulkan 1.3. The NVIDIA card has a higher API ceiling, particularly in DirectX feature level. The AMD card is limited to the older DirectX 12_1 feature set. Both are end-of-life products, but the NVIDIA card was released in February 2021, while the AMD card came out in August 2017.

Specification Differences

The two cards differ on nearly every specification field. The NVIDIA CMP 40HX has a base clock of 1470 MHz and a boost clock of 1650 MHz. The AMD Radeon Pro Vega 56 has a base clock of 1138 MHz and a boost clock of 1250 MHz. The NVIDIA card runs at significantly higher clocks.

Process nodes differ: 12 nm for NVIDIA versus 14 nm for AMD. Transistor counts are 10,800 million versus 12,500 million, and die sizes are 445 mm² versus 495 mm². The NVIDIA card has a higher pixel rate, 105.6 GPixel/s versus 80.00 GPixel/s, but the AMD card has a higher texture rate, 280.0 GTexel/s versus 237.6 GTexel/s. FP32 is 7.603 TFLOPS for NVIDIA versus 8.960 TFLOPS for AMD. FP16 is 15.21 TFLOPS versus 17.92 TFLOPS, both at a 2:1 ratio.

TDP is 185 W for the NVIDIA card and 210 W for the AMD card. The NVIDIA card is dual-slot with a 1x 8-pin power connector and a suggested PSU of 450 W. The AMD card is an IGP with no power connectors and no suggested PSU listed. The NVIDIA card uses PCIe 1.0 x4, while the AMD card uses PCIe 3.0 x16, a major difference in host interface bandwidth. The NVIDIA card has no display outputs; the AMD card has 1x HDMI 2.0b and 3x DisplayPort 1.4a. The NVIDIA card measures 229 mm in length, 111 mm in height, and 35 mm in width; the AMD card has no recorded dimensions. The NVIDIA card has a launch MSRP of 699 USD, while the AMD card has no recorded launch MSRP.

DETAILED SPECIFICATIONS

SPECIFICATION
Pro Vega 56
CMP 40HX
Core Specs
Shading Units
3,584
2,304 -35.7%
Shaders
3,584
2,304 -35.7%
TMUs
224
144 -35.7%
ROPs
64
64 0.0%
Compute Units
56
SM Count
36
Clocks
Base Clock
1138 MHz
1470 MHz
Boost Clock
1250 MHz
1650 MHz
Memory Clock
786 MHz 1572 Mbps effective
1750 MHz 14 Gbps effective
Memory
Memory Size
8 GB
8 GB
VRAM (MB)
8,192
8,192 0.0%
Memory Type
HBM2
GDDR6
Memory Bus
2048 bit
256 bit
Bandwidth
402.4 GB/s
448.0 GB/s
Cache
L1 Cache
16 KB (per CU)
64 KB (per SM)
L2 Cache
4 MB
4 MB
Performance
Pixel Rate
80.00 GPixel/s
105.6 GPixel/s
Texture Rate
280.0 GTexel/s
237.6 GTexel/s
FP32 (TFLOPS)
8.960 TFLOPS
7.603 TFLOPS
FP64 (TFLOPS)
560.0 GFLOPS (1:16)
237.6 GFLOPS (1:32)
FP16 (TFLOPS)
17.92 TFLOPS (2:1)
15.21 TFLOPS (2:1)
AI/RT
RT Cores
36
Tensor Cores
288
Power
TDP
210 W
185 W
TDP (W)
210
185 -11.9%
Suggested PSU
450 W
Power Connectors
None
1x 8-pin
Architecture
Architecture
GCN 5.0
Turing
GPU Name
Vega 10
TU106
Generation
Radeon Pro Mac (Vega Series)
Mining GPUs
Process Size
14 nm
12 nm
Transistors
12,500 million
10,800 million
Die Size
495 mm²
445 mm²
Foundry
GlobalFoundries
TSMC
Density
25.3M / mm²
24.3M / mm²
API Support
DirectX
12 (12_1)
12 Ultimate (12_2)
OpenGL
4.6
4.6
Vulkan
1.3
1.4
OpenCL
2.1
3.0
CUDA
7.5
Shader Model
6.7
6.8
Physical
Slot Width
IGP
Dual-slot
Length
229 mm 9 inches
Height
111 mm 4.4 inches
Outputs
1x HDMI 2.0b3x DisplayPort 1.4a
No outputs
Bus Interface
PCIe 3.0 x16
PCIe 1.0 x4
Other
Launch Price
699 USD
Production
End-of-life
End-of-life
View Radeon Pro Vega 56 Details View CMP 40HX Details