AMD Radeon Pro W6600X vs NVIDIA CMP 40HX Comparison

AMD
RADEON

AMD Radeon Pro W6600X

CORE STATE Navi 23
VRAM 8 GB
CLOCK SPEED 2479 MHz
TDP 120 W
BUS WIDTH 128 bit
ARCHITECTURE RDNA 2.0
nm
PROCESS 7 nm
LAUNCH DATE 2021
VS
NVIDIA
GEFORCE

CMP 40HX

CORE STATE TU106
VRAM 8 GB
CLOCK SPEED 1650 MHz
TDP 185 W
BUS WIDTH 256 bit
ARCHITECTURE Turing
nm
PROCESS 12 nm
LAUNCH DATE 2021

PERFORMANCE BENCHMARKS

geekbench_metal
107,342
N/A
geekbench_opencl
N/A
93,395
geekbench_vulkan
N/A
77,879

Analysis: AMD Radeon Pro W6600X vs NVIDIA CMP 40HX

Where Each One Wins

The benchmark data splits these two cards into very different performance profiles, and the wins are not evenly distributed. The AMD Radeon Pro W6600X is the clear leader in raw compute throughput. Its single Geekbench Metal score of 107,342 places it in the 94th percentile of all GPUs, which is a strong showing for a workstation-oriented card. That score puts it 5.4% ahead of the NVIDIA Quadro RTX 6000, which averages 101,872, and 0.6% ahead of the AMD Radeon Pro Vega II Duo at 106,750. The W6600X also sits just 2.1% behind the AMD Radeon Pro Vega II (109,617) and 3.1% behind the AMD Radeon PRO W7900 (110,725). In essence, the W6600X wins on Metal compute performance, a metric that matters for macOS-based rendering and compute workloads.

The NVIDIA CMP 40HX, by contrast, does not win on any single benchmark in the data provided. Its best result is a Geekbench OpenCL score of 93,395, which is respectable but still trails the W6600X’s Metal score. The CMP 40HX also posts a Geekbench Vulkan score of 77,879. Its average benchmark score across all tests is 85,637, which places it in the 93rd percentile of all GPUs. That is only one percentile point behind the W6600X, but the average score gap is substantial: the W6600X averages 107,342, which is 21,705 points higher, a delta of roughly 25.3%. The CMP 40HX does, however, compare favorably to its own nearest rivals. It sits 4.4% ahead of the AMD Radeon PRO W6600 (81,995) and 5.8% ahead of the AMD Radeon Pro Vega 64X (80,959). It trails the AMD Radeon PRO W7600 (87,108) by 1.7% and the NVIDIA Quadro GP100 (87,445) by 2.1%.

So, in terms of use-case wins, the W6600X wins in Metal-centric environments, which are typical of Apple ecosystem workflows. The CMP 40HX, despite lacking a win in the supplied benchmarks, still holds its own in OpenCL and Vulkan contexts, where it delivers usable performance. The data does not include any head-to-head benchmark results, so there is no direct comparison where one card beats the other in the same test. Instead, the wins are defined by API and ecosystem: the W6600X dominates Metal, while the CMP 40HX offers OpenCL and Vulkan results that are lower but not negligible.

The Verdict

The verdict is straightforward: pick the AMD Radeon Pro W6600X if your workload runs on Metal and you need maximum compute throughput. The data shows it scores 107,342 in Geekbench Metal, which is 25.3% higher than the CMP 40HX’s average of 85,637 across all its tests. The W6600X also sits in a higher performance bracket relative to its rivals, being within 3.1% of the AMD Radeon PRO W7900, a top-tier card. If you are building or upgrading a Mac Pro system, the W6600X is the obvious choice because it uses the Apple MPX bus interface, whereas the CMP 40HX uses PCIe 1.0 x4, which is an older and slower interface. The W6600X also has a lower TDP of 120 W versus the CMP 40HX’s 185 W, and it requires only a 300 W suggested PSU compared to the CMP 40HX’s 450 W suggestion.

Pick the NVIDIA CMP 40HX if your software relies on OpenCL or Vulkan and you do not need Metal support. The CMP 40HX posts an OpenCL score of 93,395 and a Vulkan score of 77,879, which are the only non-Metal results in the pack. It also has 288 tensor cores, which the W6600X lacks entirely, so if your workload uses tensor-core acceleration, the CMP 40HX is the only option here. The CMP 40HX also has a wider 256-bit memory bus and 448.0 GB/s of bandwidth, compared to the W6600X’s 128-bit bus and 256.0 GB/s. That memory advantage could matter for bandwidth-sensitive tasks, even if raw compute is lower. However, note that the CMP 40HX is a mining card with no display outputs, just like the W6600X, so neither is suitable for a typical desktop setup. The CMP 40HX is also physically smaller at 229 mm in length, 111 mm in height, and 35 mm in width, versus the W6600X which has no listed dimensions.

Head-to-Head Benchmarks

There are no direct head-to-head benchmark results in the data. The two cards were tested with different APIs: the W6600X only has a Geekbench Metal score, while the CMP 40HX has Geekbench OpenCL and Vulkan scores. This makes a direct apples-to-apples comparison impossible. However, we can still compare their average benchmark scores. The W6600X averages 107,342, while the CMP 40HX averages 85,637. That is a 21,705-point difference, which translates to the W6600X being roughly 25.3% faster on average. In terms of nearest rivals, the W6600X’s closest competitor is the AMD Radeon Pro Vega II Duo at 106,750, a delta of only 0.6%, meaning the W6600X is essentially tied with that card. The CMP 40HX’s closest rival is the AMD Radeon PRO W7600 at 87,108, a delta of -1.7%, meaning the CMP 40HX is slightly slower than that card.

The biggest win for the W6600X in the data is its Metal score of 107,342, which is 5.4% higher than the NVIDIA Quadro RTX 6000’s average of 101,872. That is a meaningful margin for a workstation card. The biggest win for the CMP 40HX, relative to its own rivals, is its 5.8% lead over the AMD Radeon Pro Vega 64X (80,959). That shows the CMP 40HX is not a slouch in its own performance tier, even if it cannot match the W6600X’s raw compute. The CMP 40HX also holds a 4.4% lead over the AMD Radeon PRO W6600 (81,995). So, while the W6600X wins the overall performance battle, the CMP 40HX is competitive within its own niche, particularly for OpenCL and Vulkan workloads.

FAQ

Q: Which card has a higher average benchmark score?

A: The AMD Radeon Pro W6600X has an average benchmark score of 107,342, which is significantly higher than the NVIDIA CMP 40HX’s average of 85,637.

Q: Does the NVIDIA CMP 40HX support Metal?

A: The data does not list a Metal benchmark for the CMP 40HX. It only shows Geekbench OpenCL (93,395) and Geekbench Vulkan (77,879) results.

Q: Which card has more memory bandwidth?

A: The NVIDIA CMP 40HX has 448.0 GB/s of bandwidth due to its 256-bit bus, while the AMD Radeon Pro W6600X has 256.0 GB/s over a 128-bit bus.

Q: Are either of these cards suitable for display output?

A: No. Both cards list "No outputs" for display connections, meaning neither can drive a monitor directly.

Q: What is the transistor density of each card?

A: The AMD Radeon Pro W6600X has a transistor density of 46.7M per mm² on a 7 nm process, while the NVIDIA CMP 40HX has 24.3M per mm² on a 12 nm process.

Q: Which card has tensor cores?

A: Only the NVIDIA CMP 40HX has tensor cores, with 288 of them. The AMD Radeon Pro W6600X lists no tensor cores.

Architecture Differences

The two cards are built on fundamentally different architectures. The AMD Radeon Pro W6600X uses the Navi 23 chip based on RDNA 2.0, manufactured on a 7 nm process at TSMC. It packs 11,060 million transistors into a 237 mm² die, yielding a transistor density of 46.7M per mm². The NVIDIA CMP 40HX uses the TU106 chip based on Turing, also manufactured at TSMC but on a 12 nm process. It has 10,800 million transistors on a much larger 445 mm² die, resulting in a lower density of 24.3M per mm². The W6600X’s smaller, denser die is a direct result of the more advanced 7 nm node.

In terms of compute units, the W6600X has 2,048 shading units, 128 TMUs, and 64 ROPs, along with 32 ray tracing cores. The CMP 40HX has more shading units at 2,304, more TMUs at 144, and the same 64 ROPs. It also has 36 ray tracing cores and 288 tensor cores. The W6600X has no tensor cores at all. The CMP 40HX’s tensor cores are a notable architectural feature that the W6600X cannot match, but the W6600X compensates with higher clock speeds: a base of 2068 MHz and a boost of 2479 MHz, versus the CMP 40HX’s 1470 MHz base and 1650 MHz boost.

Memory architecture also differs. Both use 8 GB of GDDR6, but the W6600X runs it at 2000 MHz (16 Gbps effective) over a 128-bit bus, while the CMP 40HX runs at 1750 MHz (14 Gbps effective) over a 256-bit bus. The CMP 40HX’s wider bus gives it higher bandwidth (448.0 GB/s vs 256.0 GB/s). The W6600X has higher pixel and texture rates: 158.7 GPixel/s and 317.3 GTexel/s, respectively, versus the CMP 40HX’s 105.6 GPixel/s and 237.6 GTexel/s. The W6600X also leads in FP32 performance at 10.15 TFLOPS, compared to the CMP 40HX’s 7.603 TFLOPS. FP16 performance follows the same pattern: 20.31 TFLOPS for the W6600X versus 15.21 TFLOPS for the CMP 40HX.

Specification Differences

The specification sheet shows clear divergences beyond the core architecture. The AMD Radeon Pro W6600X has a TDP of 120 W, while the NVIDIA CMP 40HX draws 185 W. The suggested PSU is also lower for the W6600X at 300 W, versus 450 W for the CMP 40HX. The W6600X uses an Apple MPX bus interface, while the CMP 40HX uses PCIe 1.0 x4, an older and slower interface. Neither card has display outputs. The CMP 40HX has a power connector specification of 1x 8-pin, whereas the W6600X lists no power connector details.

Physical dimensions differ as well. The CMP 40HX measures 229 mm in length, 111 mm in height, and 35 mm in width. The W6600X has no listed dimensions. Both are dual-slot cards. The W6600X was released on 2021-08-02, while the CMP 40HX came earlier on 2021-02-24. Both are end-of-life products. The W6600X belongs to the "Radeon Pro Mac (Navi II Series)" generation, while the CMP 40HX is part of "Mining GPUs." The W6600X supports DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4. The CMP 40HX supports the same API list. Both have a launch MSRP of 699 USD. The W6600X’s nearest rival is the AMD Radeon Pro Vega II Duo, while the CMP 40HX’s nearest rival is the AMD Radeon PRO W7600.

DETAILED SPECIFICATIONS

SPECIFICATION
Pro W6600X
CMP 40HX
Core Specs
Shading Units
2,048
2,304 +12.5%
Shaders
2,048
2,304 +12.5%
TMUs
128
144 +12.5%
ROPs
64
64 0.0%
Compute Units
32
SM Count
36
Clocks
Base Clock
2068 MHz
1470 MHz
Boost Clock
2479 MHz
1650 MHz
Memory Clock
2000 MHz 16 Gbps effective
1750 MHz 14 Gbps effective
Memory
Memory Size
8 GB
8 GB
VRAM (MB)
8,192
8,192 0.0%
Memory Type
GDDR6
GDDR6
Memory Bus
128 bit
256 bit
Bandwidth
256.0 GB/s
448.0 GB/s
Cache
L1 Cache
128 KB per Array
64 KB (per SM)
L2 Cache
2 MB
4 MB
L3 Cache
32 MB
L0 Cache
32 KB per WGP
Performance
Pixel Rate
158.7 GPixel/s
105.6 GPixel/s
Texture Rate
317.3 GTexel/s
237.6 GTexel/s
FP32 (TFLOPS)
10.15 TFLOPS
7.603 TFLOPS
FP64 (TFLOPS)
634.6 GFLOPS (1:16)
237.6 GFLOPS (1:32)
FP16 (TFLOPS)
20.31 TFLOPS (2:1)
15.21 TFLOPS (2:1)
AI/RT
RT Cores
32
36 +12.5%
Tensor Cores
288
Power
TDP
120 W
185 W
TDP (W)
120
185 +54.2%
Suggested PSU
300 W
450 W
Power Connectors
1x 8-pin
Architecture
Architecture
RDNA 2.0
Turing
GPU Name
Navi 23
TU106
Generation
Radeon Pro Mac (Navi II Series)
Mining GPUs
Process Size
7 nm
12 nm
Transistors
11,060 million
10,800 million
Die Size
237 mm²
445 mm²
Foundry
TSMC
TSMC
Density
46.7M / mm²
24.3M / mm²
API Support
DirectX
12 Ultimate (12_2)
12 Ultimate (12_2)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
2.1
3.0
CUDA
7.5
Shader Model
6.8
6.8
Physical
Slot Width
Dual-slot
Dual-slot
Length
229 mm 9 inches
Height
111 mm 4.4 inches
Outputs
No outputs
No outputs
Bus Interface
Apple MPX
PCIe 1.0 x4
Other
Launch Price
699 USD
699 USD
Production
End-of-life
End-of-life
View Radeon Pro W6600X Details View CMP 40HX Details