AMD Radeon Pro W6800X Duo vs NVIDIA CMP 40HX Comparison

AMD
RADEON

AMD Radeon Pro W6800X Duo

CORE STATE Navi 21
VRAM 32 GB
CLOCK SPEED 1967 MHz
TDP 400 W
BUS WIDTH 256 bit
ARCHITECTURE RDNA 2.0
nm
PROCESS 7 nm
LAUNCH DATE 2021
VS
NVIDIA
GEFORCE

CMP 40HX

CORE STATE TU106
VRAM 8 GB
CLOCK SPEED 1650 MHz
TDP 185 W
BUS WIDTH 256 bit
ARCHITECTURE Turing
nm
PROCESS 12 nm
LAUNCH DATE 2021

PERFORMANCE BENCHMARKS

geekbench_metal
157,365
N/A
geekbench_opencl
124,335
93,395
geekbench_vulkan
125,622
77,879

Analysis: AMD Radeon Pro W6800X Duo vs NVIDIA CMP 40HX

FAQ

Q: Which GPU has the higher average benchmark score?

A: The AMD Radeon Pro W6800X Duo records an average benchmark score of 135,774, while the NVIDIA CMP 40HX scores 85,637. That is a 58.5% gap in favor of the AMD part.

Q: How do the two compare on Geekbench OpenCL?

A: The W6800X Duo scores 124,335 versus 93,395 for the CMP 40HX, a 33.1% advantage for the AMD card in the recorded head-to-head data.

Q: Which card wins the Vulkan benchmark?

A: The AMD Radeon Pro W6800X Duo wins decisively, scoring 125,622 against 77,879 for the NVIDIA CMP 40HX, a 61.3% difference.

Q: What is the production status of both cards?

A: Both GPUs are listed as end-of-life in the database. The AMD card was released on 2021-08-02 and the NVIDIA card on 2021-02-24.

Q: Does the NVIDIA CMP 40HX have any display outputs?

A: No, the CMP 40HX has no display outputs, which reflects its mining-oriented design. The AMD W6800X Duo offers 1x HDMI 2.1 and 4x Thunderbolt outputs.

Q: What percentile rank does each GPU hold among all GPUs?

A: The AMD Radeon Pro W6800X Duo sits at the 96th percentile, while the NVIDIA CMP 40HX sits at the 93rd percentile.

Architecture Differences

The two cards come from different architectural generations and foundry processes. The AMD Radeon Pro W6800X Duo uses the Navi 21 chip built on RDNA 2.0, manufactured on a 7 nm process at TSMC. The NVIDIA CMP 40HX uses the TU106 chip based on Turing, also fabricated by TSMC but on a 12 nm process. The process gap explains part of the density difference: the AMD die packs 26,800 million transistors into 520 mm², yielding a transistor density of 51.5M per mm², while the NVIDIA die contains 10,800 million transistors across 445 mm², for a density of 24.3M per mm².

Shader resources differ substantially. The W6800X Duo carries 3,840 shading units, 240 texture mapping units, and 96 ROPs. The CMP 40HX has 2,304 shading units, 144 TMUs, and 64 ROPs. Ray tracing hardware also differs: the AMD card has 60 RT cores, while the NVIDIA card has 36 RT cores. NVIDIA, however, adds 288 tensor cores to the CMP 40HX, a feature the AMD card does not list.

Memory architecture is similar in bus width but not in capacity or speed. Both GPUs use a 256-bit memory bus with GDDR6 memory. The AMD card offers 32 GB at a memory clock of 2000 MHz with 16 Gbps effective speed, yielding 512.0 GB/s of bandwidth. The NVIDIA card has 8 GB at 1750 MHz with 14 Gbps effective speed, producing 448.0 GB/s. The AMD card therefore has four times the capacity and roughly 14% more bandwidth.

Compute throughput follows the same pattern. The W6800X Duo reaches 15.11 TFLOPS FP32 and 30.21 TFLOPS FP16 (2:1). The CMP 40HX reaches 7.603 TFLOPS FP32 and 15.21 TFLOPS FP16 (2:1). In both precision formats, the AMD card delivers approximately double the raw compute.

The physical and interface designs diverge sharply. The AMD card is a quad-slot, 267 mm long and 120 mm tall, using an Apple MPX bus interface, with a 400 W TDP and an 800 W suggested power supply. The NVIDIA card is a dual-slot, 229 mm long, 111 mm tall, and 35 mm wide, using a PCIe 1.0 x4 bus interface with a 1x 8-pin power connector, a 185 W TDP, and a 450 W suggested PSU. The W6800X Duo has display outputs; the CMP 40HX has none.

Both cards support DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4. The AMD card also has a Geekbench Metal score of 157,365, a test the NVIDIA card does not appear in.

Where Each One Wins

The AMD Radeon Pro W6800X Duo wins in every recorded head-to-head benchmark. In OpenCL, it leads by 33.1%, and in Vulkan it leads by 61.3%. The database shows two wins for the AMD card and zero for the NVIDIA card.

The AMD card is the clear choice for compute-heavy workloads in the recorded benchmarks. Its 32 GB memory capacity, 512.0 GB/s bandwidth, and FP32 throughput of 15.11 TFLOPS make it suited for large datasets and rendering tasks that require both capacity and speed. Its display outputs, including 4x Thunderbolt, also mean it can drive professional display configurations, something the CMP 40HX cannot do at all.

The NVIDIA CMP 40HX has no wins in the database, but its profile suggests a different design intent. It is a mining GPU with no display outputs, a lower 185 W TDP, and a smaller 229 mm length. Its 288 tensor cores are a distinguishing feature, though no benchmark in the database isolates tensor performance. The card also carries a PCIe 1.0 x4 interface, which limits host transfer bandwidth compared to the AMD card's Apple MPX bus.

The data indicates that for any workload represented by the Geekbench OpenCL and Vulkan suites, the AMD Radeon Pro W6800X Duo is the stronger performer. The NVIDIA card's advantages, if any, would have to lie outside the measured tests, such as in its tensor core count or lower power envelope, but the recorded benchmarks do not capture those areas.

Specification Differences

| Specification | AMD Radeon Pro W6800X Duo | NVIDIA CMP 40HX |

|---|---|---|

| Chip | Navi 21 | TU106 |

| Architecture | RDNA 2.0 | Turing |

| Process Node | 7 nm | 12 nm |

| Transistors | 26,800 million | 10,800 million |

| Die Size | 520 mm² | 445 mm² |

| Transistor Density | 51.5M / mm² | 24.3M / mm² |

| Shading Units | 3840 | 2304 |

| TMUs | 240 | 144 |

| ROPs | 96 | 64 |

| RT Cores | 60 | 36 |

| Tensor Cores | null | 288 |

| Base Clock | 1800 MHz | 1470 MHz |

| Boost Clock | 1967 MHz | 1650 MHz |

| Memory Size | 32 GB | 8 GB |

| Memory Type | GDDR6 | GDDR6 |

| Memory Bus | 256 bit | 256 bit |

| Memory Clock | 2000 MHz, 16 Gbps effective | 1750 MHz, 14 Gbps effective |

| Memory Bandwidth | 512.0 GB/s | 448.0 GB/s |

| FP32 | 15.11 TFLOPS | 7.603 TFLOPS |

| FP16 | 30.21 TFLOPS (2:1) | 15.21 TFLOPS (2:1) |

| Pixel Rate | 188.8 GPixel/s | 105.6 GPixel/s |

| Texture Rate | 472.1 GTexel/s | 237.6 GTexel/s |

| TDP | 400 W | 185 W |

| Slot Width | Quad-slot | Dual-slot |

| Power Connectors | null | 1x 8-pin |

| Suggested PSU | 800 W | 450 W |

| Bus Interface | Apple MPX | PCIe 1.0 x4 |

| Display Outputs | 1x HDMI 2.1, 4x Thunderbolt | No outputs |

| Length | 267 mm (10.5 inches) | 229 mm (9 inches) |

| Height | 120 mm (4.7 inches) | 111 mm (4.4 inches) |

| Width | null | 35 mm (1.4 inches) |

The AMD card has higher clocks in both base and boost states. Its base clock of 1800 MHz exceeds the NVIDIA card's boost clock of 1650 MHz. The AMD card also produces higher pixel and texture rates: 188.8 GPixel/s versus 105.6 GPixel/s, and 472.1 GTexel/s versus 237.6 GTexel/s.

The two GPUs differ in launch MSRP. The AMD Radeon Pro W6800X Duo had a launch MSRP of 4,999 USD. The NVIDIA CMP 40HX had a launch MSRP of 699 USD.

Head-to-Head Benchmarks

The database includes two direct comparisons between these GPUs. Both come from the Geekbench suite, and both favor the AMD Radeon Pro W6800X Duo.

In Geekbench OpenCL, the AMD card scores 124,335 against 93,395 for the NVIDIA CMP 40HX. The delta is 33.1%. This margin is consistent with the underlying specifications: the AMD card has 66.7% more shading units, 66.7% more TMUs, 50% more ROPs, and roughly double the FP32 throughput. Its memory bandwidth advantage of 512.0 GB/s versus 448.0 GB/s also helps in memory-bound OpenCL workloads, though the capacity gap of 32 GB versus 8 GB is likely the more significant factor for large working sets.

In Geekbench Vulkan, the gap widens considerably. The AMD card scores 125,622, while the NVIDIA card scores 77,879, a 61.3% difference. The larger delta in Vulkan relative to OpenCL suggests the AMD architecture handles the Vulkan workload more efficiently, or that the NVIDIA card's Turing architecture is less optimized for this particular test. The AMD card's Vulkan score is actually slightly higher than its OpenCL score, while the NVIDIA card's Vulkan score drops by 16.6% compared to its OpenCL result.

The average benchmark score tells the same story. The AMD Radeon Pro W6800X Duo averages 135,774 across its three recorded tests, including a Metal score of 157,365 that has no counterpart for the NVIDIA card. The CMP 40HX averages 85,637 across its two recorded tests. The difference in average score is 50,137 points, or 58.5%.

The percentile ranking reinforces the performance hierarchy. The AMD card sits at the 96th percentile of all GPUs, while the NVIDIA card sits at the 93rd percentile. Both are high performers in absolute terms, but the AMD card occupies a higher tier.

Rival comparisons provide additional context. The AMD card's nearest rivals include the AMD Radeon PRO W6800 with an average score of 135,396 and a delta of 0.3%, the NVIDIA A10M at 135,230 with a delta of 0.4%, the NVIDIA RTX 4000 Ada Generation at 135,218 with a delta of 0.4%, and the AMD Radeon PRO V620 at 136,472 with a delta of -0.5%. The W6800X Duo is effectively at parity with these cards, trailing the PRO V620 by only 0.5%.

The NVIDIA CMP 40HX sits in a different competitive bracket. Its nearest rivals include the AMD Radeon PRO W7600 at 87,108 with a delta of -1.7%, the NVIDIA Quadro GP100 at 87,445 with a delta of -2.1%, the AMD Radeon PRO W6600 at 81,995 with a delta of 4.4%, and the AMD Radeon Pro Vega 64X at 80,959 with a delta of 5.8%. The CMP 40HX leads the W6600 by 4.4% and the Vega 64X by 5.8%, but trails the W7600 and Quadro GP100 by small margins.

The head-to-head data shows a dominant performance advantage for the AMD Radeon Pro W6800X Duo. The NVIDIA CMP 40HX, despite being a capable GPU in its own percentile tier, does not approach the AMD card in any measured benchmark. The two wins recorded in the database belong entirely to the AMD part, with margins of 33.1% and 61.3% respectively.

DETAILED SPECIFICATIONS

SPECIFICATION
Pro W6800X Duo
CMP 40HX
Core Specs
Shading Units
3,840
2,304 -40.0%
Shaders
3,840
2,304 -40.0%
TMUs
240
144 -40.0%
ROPs
96
64 -33.3%
Compute Units
60
—
SM Count
—
36
Clocks
Base Clock
1800 MHz
1470 MHz
Boost Clock
1967 MHz
1650 MHz
Memory Clock
2000 MHz 16 Gbps effective
1750 MHz 14 Gbps effective
Memory
Memory Size
32 GB
8 GB
VRAM (MB)
32,768
8,192 -75.0%
Memory Type
GDDR6
GDDR6
Memory Bus
256 bit
256 bit
Bandwidth
512.0 GB/s
448.0 GB/s
Cache
L1 Cache
128 KB per Array
64 KB (per SM)
L2 Cache
4 MB
4 MB
L3 Cache
128 MB
—
L0 Cache
32 KB per WGP
—
Performance
Pixel Rate
188.8 GPixel/s
105.6 GPixel/s
Texture Rate
472.1 GTexel/s
237.6 GTexel/s
FP32 (TFLOPS)
15.11 TFLOPS
7.603 TFLOPS
FP64 (TFLOPS)
944.2 GFLOPS (1:16)
237.6 GFLOPS (1:32)
FP16 (TFLOPS)
30.21 TFLOPS (2:1)
15.21 TFLOPS (2:1)
AI/RT
RT Cores
60
36 -40.0%
Tensor Cores
—
288
Power
TDP
400 W
185 W
TDP (W)
400
185 -53.8%
Suggested PSU
800 W
450 W
Power Connectors
—
1x 8-pin
Architecture
Architecture
RDNA 2.0
Turing
GPU Name
Navi 21
TU106
Generation
Radeon Pro Mac (Navi II Series)
Mining GPUs
Process Size
7 nm
12 nm
Transistors
26,800 million
10,800 million
Die Size
520 mm²
445 mm²
Foundry
TSMC
TSMC
Density
51.5M / mm²
24.3M / mm²
API Support
DirectX
12 Ultimate (12_2)
12 Ultimate (12_2)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
2.1
3.0
CUDA
—
7.5
Shader Model
6.8
6.8
Physical
Slot Width
Quad-slot
Dual-slot
Length
267 mm 10.5 inches
229 mm 9 inches
Height
120 mm 4.7 inches
111 mm 4.4 inches
Outputs
1x HDMI 2.14x Thunderbolt
No outputs
Bus Interface
Apple MPX
PCIe 1.0 x4
Other
Launch Price
4,999 USD
699 USD
Production
End-of-life
End-of-life
View Radeon Pro W6800X Duo Details View CMP 40HX Details