AMD Radeon Pro W6900X vs NVIDIA CMP 40HX Comparison

AMD
RADEON

AMD Radeon Pro W6900X

CORE STATE Navi 21
VRAM 32 GB
CLOCK SPEED 2171 MHz
TDP 300 W
BUS WIDTH 256 bit
ARCHITECTURE RDNA 2.0
nm
PROCESS 7 nm
LAUNCH DATE 2021
VS
NVIDIA
GEFORCE

CMP 40HX

CORE STATE TU106
VRAM 8 GB
CLOCK SPEED 1650 MHz
TDP 185 W
BUS WIDTH 256 bit
ARCHITECTURE Turing
nm
PROCESS 12 nm
LAUNCH DATE 2021

PERFORMANCE BENCHMARKS

geekbench_metal
226,821
N/A
geekbench_opencl
130,035
93,395
geekbench_vulkan
148,865
77,879

Analysis: AMD Radeon Pro W6900X vs NVIDIA CMP 40HX

FAQ

Q: How does the AMD Radeon Pro W6900X compare to the NVIDIA CMP 40HX in overall benchmark scores?

A: The AMD Radeon Pro W6900X has an average benchmark score of 168,574, while the NVIDIA CMP 40HX averages 85,637. This places the AMD card at the 97th percentile of all GPUs, while the NVIDIA card sits at the 93rd percentile.

Q: Which GPU wins in OpenCL performance, and by how much?

A: The AMD Radeon Pro W6900X scores 130,035 in Geekbench OpenCL, compared to 93,395 for the NVIDIA CMP 40HX. That is a 39.2% advantage for the AMD card.

Q: What about Vulkan performance?

A: The AMD Radeon Pro W6900X delivers 148,865 in Geekbench Vulkan, while the NVIDIA CMP 40HX scores 77,879. The AMD card is ahead by 91.1% in this test.

Q: Are there any benchmark tests where the NVIDIA CMP 40HX wins?

A: No. Across the recorded head-to-head benchmarks (OpenCL and Vulkan), the AMD Radeon Pro W6900X wins both tests. The NVIDIA CMP 40HX has zero wins in this comparison.

Q: What is the memory configuration difference between the two?

A: The AMD Radeon Pro W6900X has 32 GB of GDDR6 memory on a 256-bit bus, yielding 512.0 GB/s bandwidth. The NVIDIA CMP 40HX has 8 GB of GDDR6 on the same 256-bit bus, providing 448.0 GB/s bandwidth.

Q: What are the launch MSRP values for these cards?

A: The AMD Radeon Pro W6900X had a launch MSRP of 5,999 USD, while the NVIDIA CMP 40HX had a launch MSRP of 699 USD.

Architecture Differences

The AMD Radeon Pro W6900X is built on the Navi 21 chip using RDNA 2.0 architecture, fabricated on a 7 nm process at TSMC. It packs 26,800 million transistors into a 520 mm² die, resulting in a transistor density of 51.5 million per mm². The NVIDIA CMP 40HX uses the TU106 chip with Turing architecture, also made by TSMC but on a 12 nm process. It contains 10,800 million transistors on a 445 mm² die, giving a density of 24.3 million per mm².

The compute resources differ dramatically. The AMD card features 5,120 shading units, 320 texture mapping units, 80 ray tracing cores, and 128 ROPs. The NVIDIA CMP 40HX has 2,304 shading units, 144 TMUs, 36 RT cores, and 64 ROPs. Notably, the NVIDIA card includes 288 tensor cores, a feature entirely absent from the AMD Radeon Pro W6900X.

Clock speeds also diverge. The AMD card runs at a base of 1825 MHz and boosts to 2171 MHz, with memory clocked at 2000 MHz (16 Gbps effective). The NVIDIA card has a lower base clock of 1470 MHz and boost of 1650 MHz, with memory at 1750 MHz (14 Gbps effective).

The bus interface sets them apart further: the AMD card uses Apple MPX, while the NVIDIA card uses PCIe 1.0 x4. Physically, the AMD card measures 267 mm in length and 120 mm in height, whereas the NVIDIA card is 229 mm long, 111 mm tall, and 35 mm wide, fitting a dual-slot form factor with a single 8-pin power connector. The AMD card draws up to 300 W with a 700 W suggested PSU, while the NVIDIA card has a 185 W TDP and 450 W suggested PSU. The AMD card provides display outputs (1x HDMI 2.1 and 4x Thunderbolt), while the NVIDIA CMP 40HX has no display outputs at all.

Head-to-Head Benchmarks

The recorded data includes two direct comparisons: Geekbench OpenCL and Geekbench Vulkan. In both, the AMD Radeon Pro W6900X dominates.

Starting with OpenCL, the AMD card scores 130,035 against the NVIDIA card's 93,395, a 39.2% lead. This gap reflects the fundamental resource differences: the AMD card has more than double the shading units (5,120 vs. 2,304), more than double the TMUs (320 vs. 144), and double the ROPs (128 vs. 64). The AMD card's FP32 throughput of 22.23 TFLOPS versus 7.603 TFLOPS for the NVIDIA card explains much of the raw compute advantage.

The Vulkan test shows an even larger margin. The AMD Radeon Pro W6900X scores 148,865, while the NVIDIA CMP 40HX manages 77,879, giving the AMD card a 91.1% advantage. This near-doubling of performance in Vulkan suggests the AMD architecture scales more efficiently in this API, possibly due to its newer RDNA 2.0 design and higher memory bandwidth (512.0 GB/s vs. 448.0 GB/s).

It is importantly the NVIDIA CMP 40HX lacks any Geekbench Metal score in the database, likely because it has no display outputs and is not intended for graphics workloads. The AMD card, by contrast, has a Metal score of 226,821, which is its strongest single benchmark result.

The average benchmark scores further contextualize the gap: the AMD card's 168,574 average is nearly double the NVIDIA card's 85,637. In relative terms, the AMD card sits 39.2% ahead in OpenCL and 91.1% ahead in Vulkan, with two wins and zero losses in the head-to-head record.

The Verdict

The data is unambiguous. The AMD Radeon Pro W6900X outperforms the NVIDIA CMP 40HX in every recorded benchmark category. Its average score of 168,574 places it in the 97th percentile of all GPUs, while the NVIDIA card's 85,637 average lands in the 93rd percentile. The AMD card also shows strong rivalry with professional workstation cards like the NVIDIA RTX 4500 Ada Generation (166,094 average, 1.5% behind) and the NVIDIA RTX A5500 (165,217 average, 2% behind). The NVIDIA CMP 40HX, meanwhile, trades blows with mid-range workstation cards: it is 1.7% behind the AMD Radeon PRO W7600 (87,108), 2.1% behind the NVIDIA Quadro GP100 (87,445), and 4.4% ahead of the AMD Radeon PRO W6600 (81,995).

Who should choose which? If the workload demands maximum compute throughput, ray tracing capability, and large memory capacity, the AMD Radeon Pro W6900X is the clear choice. It offers 32 GB of GDDR6, 80 RT cores, and FP32 performance of 22.23 TFLOPS. If the task is purely mining-oriented, the NVIDIA CMP 40HX was designed for that purpose, with no display outputs and a lower power draw of 185 W. The AMD card is a professional workstation GPU with display outputs and a 300 W TDP, suited for rendering or compute tasks where visual output is necessary.

Specification Differences

| Specification | AMD Radeon Pro W6900X | NVIDIA CMP 40HX |

|---|---|---|

| Architecture | RDNA 2.0 | Turing |

| Process Node | 7 nm | 12 nm |

| Transistors | 26,800 million | 10,800 million |

| Die Size | 520 mm² | 445 mm² |

| Transistor Density | 51.5M / mm² | 24.3M / mm² |

| Base Clock | 1825 MHz | 1470 MHz |

| Boost Clock | 2171 MHz | 1650 MHz |

| Memory Clock | 2000 MHz (16 Gbps effective) | 1750 MHz (14 Gbps effective) |

| Memory Size | 32 GB | 8 GB |

| Memory Type | GDDR6 | GDDR6 |

| Bandwidth | 512.0 GB/s | 448.0 GB/s |

| Shading Units | 5120 | 2304 |

| TMUs | 320 | 144 |

| ROPs | 128 | 64 |

| RT Cores | 80 | 36 |

| Tensor Cores | None | 288 |

| Pixel Rate | 277.9 GPixel/s | 105.6 GPixel/s |

| Texture Rate | 694.7 GTexel/s | 237.6 GTexel/s |

| FP32 | 22.23 TFLOPS | 7.603 TFLOPS |

| FP16 | 44.46 TFLOPS (2:1) | 15.21 TFLOPS (2:1) |

| TDP | 300 W | 185 W |

| Slot Width | Not specified | Dual-slot |

| Power Connectors | Not specified | 1x 8-pin |

| Suggested PSU | 700 W | 450 W |

| Bus Interface | Apple MPX | PCIe 1.0 x4 |

| Display Outputs | 1x HDMI 2.1, 4x Thunderbolt | No outputs |

| Length | 267 mm | 229 mm |

| Height | 120 mm | 111 mm |

| Width | Not specified | 35 mm |

| Release Date | 2021-08-02 | 2021-02-24 |

Where Each One Wins

The AMD Radeon Pro W6900X wins decisively in compute-heavy and graphics-oriented workloads. Its OpenCL score of 130,035 and Vulkan score of 148,865 both exceed the NVIDIA card's best results by wide margins. The card also delivers a Metal benchmark of 226,821, which is relevant for Apple ecosystem applications given its Apple MPX bus interface and Thunderbolt outputs. With 32 GB of memory, it can handle large datasets or high-resolution textures without swapping. Its higher pixel rate (277.9 GPixel/s) and texture rate (694.7 GTexel/s) further cement its advantage in rendering tasks.

The NVIDIA CMP 40HX wins only in niche scenarios. It has a lower TDP of 185 W versus 300 W, making it more power-efficient per watt for mining operations. It also includes 288 tensor cores, which the AMD card lacks entirely; this could be relevant for AI inference workloads that leverage tensor operations, though no benchmark data in the database confirms this advantage. Its smaller physical footprint (229 mm length, 35 mm width) and single 8-pin connector make it easier to install in dense mining rigs. Its launch MSRP of 699 USD is dramatically lower than the AMD card's 5,999 USD, though pricing is not the focus of this analysis. The NVIDIA card's 93rd percentile ranking shows it is still a capable performer relative to the broader GPU landscape, just not competitive with the AMD Radeon Pro W6900X in direct benchmark comparisons.

DETAILED SPECIFICATIONS

SPECIFICATION
Pro W6900X
CMP 40HX
Core Specs
Shading Units
5,120
2,304 -55.0%
Shaders
5,120
2,304 -55.0%
TMUs
320
144 -55.0%
ROPs
128
64 -50.0%
Compute Units
80
SM Count
36
Clocks
Base Clock
1825 MHz
1470 MHz
Boost Clock
2171 MHz
1650 MHz
Memory Clock
2000 MHz 16 Gbps effective
1750 MHz 14 Gbps effective
Memory
Memory Size
32 GB
8 GB
VRAM (MB)
32,768
8,192 -75.0%
Memory Type
GDDR6
GDDR6
Memory Bus
256 bit
256 bit
Bandwidth
512.0 GB/s
448.0 GB/s
Cache
L1 Cache
128 KB per Array
64 KB (per SM)
L2 Cache
4 MB
4 MB
L3 Cache
128 MB
L0 Cache
32 KB per WGP
Performance
Pixel Rate
277.9 GPixel/s
105.6 GPixel/s
Texture Rate
694.7 GTexel/s
237.6 GTexel/s
FP32 (TFLOPS)
22.23 TFLOPS
7.603 TFLOPS
FP64 (TFLOPS)
1,389.4 GFLOPS (1:16)
237.6 GFLOPS (1:32)
FP16 (TFLOPS)
44.46 TFLOPS (2:1)
15.21 TFLOPS (2:1)
AI/RT
RT Cores
80
36 -55.0%
Tensor Cores
288
Power
TDP
300 W
185 W
TDP (W)
300
185 -38.3%
Suggested PSU
700 W
450 W
Power Connectors
1x 8-pin
Architecture
Architecture
RDNA 2.0
Turing
GPU Name
Navi 21
TU106
Generation
Radeon Pro Mac (Navi II Series)
Mining GPUs
Process Size
7 nm
12 nm
Transistors
26,800 million
10,800 million
Die Size
520 mm²
445 mm²
Foundry
TSMC
TSMC
Density
51.5M / mm²
24.3M / mm²
API Support
DirectX
12 Ultimate (12_2)
12 Ultimate (12_2)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
2.1
3.0
CUDA
7.5
Shader Model
6.8
6.8
Physical
Slot Width
Dual-slot
Length
267 mm 10.5 inches
229 mm 9 inches
Height
120 mm 4.7 inches
111 mm 4.4 inches
Outputs
1x HDMI 2.14x Thunderbolt
No outputs
Bus Interface
Apple MPX
PCIe 1.0 x4
Other
Launch Price
5,999 USD
699 USD
Production
End-of-life
End-of-life
View Radeon Pro W6900X Details View CMP 40HX Details