AMD Radeon PRO W6800 vs NVIDIA CMP 40HX Comparison

AMD
RADEON

AMD Radeon PRO W6800

CORE STATE Navi 21
VRAM 32 GB
CLOCK SPEED 2322 MHz
TDP 250 W
BUS WIDTH 256 bit
ARCHITECTURE RDNA 2.0
nm
PROCESS 7 nm
LAUNCH DATE 2021
VS
NVIDIA
GEFORCE

CMP 40HX

CORE STATE TU106
VRAM 8 GB
CLOCK SPEED 1650 MHz
TDP 185 W
BUS WIDTH 256 bit
ARCHITECTURE Turing
nm
PROCESS 12 nm
LAUNCH DATE 2021

PERFORMANCE BENCHMARKS

geekbench_metal
174,420
N/A
geekbench_opencl
121,808
93,395
geekbench_vulkan
109,961
77,879

Analysis: AMD Radeon PRO W6800 vs NVIDIA CMP 40HX

Where Each One Wins

The benchmark data splits cleanly between these two cards, and the verdict is one-sided. The AMD Radeon PRO W6800 wins every recorded head-to-head test. In Geekbench OpenCL, it scores 121,808 against the NVIDIA CMP 40HX's 93,395, a 30.4% advantage. In Geekbench Vulkan, the gap widens to 41.2%, with the W6800 scoring 109,961 versus 77,879. The recorded data shows 2 wins for AMD and 0 for NVIDIA.

The W6800 also holds a commanding lead in average benchmark score. Its average across all recorded tests is 135,396, while the CMP 40HX averages 85,637. That is a difference of roughly 58% in favor of the Radeon. The percentile standings confirm the hierarchy: the W6800 sits at the 96th percentile among all GPUs, while the CMP 40HX sits at the 93rd. Both are high performers, but the W6800 is clearly in a higher tier.

Where the CMP 40HX might claim a niche is not in raw performance but in its profile. The data shows it is a smaller, lower-power card. It draws a 185 W TDP compared to 250 W for the AMD, and it requires a suggested 450 W PSU versus 600 W. It is also shorter at 229 mm (9 inches) and narrower at 35 mm (1.4 inches) versus 267 mm (10.5 inches) and 50 mm (2 inches). For a system with tight space constraints or a lower wattage budget, the data supports the NVIDIA as the lighter-duty option, despite losing every benchmark.

Architecture Differences

The two cards come from opposite generations and design philosophies. The AMD Radeon PRO W6800 uses the Navi 21 chip built on RDNA 2.0 architecture, fabricated on a 7 nm process at TSMC. It packs 26,800 million transistors on a 520 mm² die, giving a transistor density of 51.5 million per mm². The NVIDIA CMP 40HX uses the TU106 chip, which is Turing architecture, fabricated on a 12 nm process at the same foundry. It has 10,800 million transistors on a 445 mm² die, resulting in a much lower 24.3 million per mm².

The memory subsystem also differs sharply. The W6800 has 32 GB of GDDR6 on a 256-bit bus with 512.0 GB/s of bandwidth. The CMP 40HX has 8 GB of GDDR6 on the same 256-bit bus, but its bandwidth drops to 448.0 GB/s. Memory clocks are 2000 MHz (16 Gbps effective) for AMD versus 1750 MHz (14 Gbps effective) for NVIDIA.

Compute resources favor the AMD heavily. The W6800 has 3840 shading units, 240 TMUs, and 96 ROPs. The CMP 40HX has 2304 shading units, 144 TMUs, and 64 ROPs. The AMD also has 60 ray tracing cores, while the NVIDIA has 36 RT cores. Notably, the NVIDIA includes 288 tensor cores (the AMD has none listed), but the database shows no TensorFloat performance numbers for either. The result is a raw throughput advantage for the AMD: 17.83 TFLOPS FP32 versus 7.603 TFLOPS, and 35.67 TFLOPS FP16 versus 15.21 TFLOPS.

Other differences matter for deployment. The W6800 is a full PCIe 4.0 x16 card, while the CMP 40HX runs on PCIe 1.0 x4. The AMD has 6x mini-DisplayPort 1.4a outputs; the NVIDIA has no display outputs at all, a direct consequence of being a mining card. Both are dual-slot. The W6800 uses two power connectors (1x 6-pin and 1x 8-pin), while the CMP 40HX uses only one 8-pin. Both support DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4.

The Verdict

The data leads to an unambiguous conclusion. For any workload that uses OpenCL or Vulkan, the AMD Radeon PRO W6800 is the superior card by a wide margin. Its 30.4% lead in OpenCL and 41.2% lead in Vulkan are not close calls; they represent a full performance tier. The 32 GB memory capacity and 512.0 GB/s bandwidth also make it the clear choice for large datasets or memory-intensive tasks. If the task is professional rendering, compute, or any application that can leverage its ray tracing cores and higher shading unit count, the W6800 is the verdict.

The NVIDIA CMP 40HX is only defensible in a narrow set of conditions. It has no display outputs, so it cannot be used as a standard graphics card. Its 8 GB memory is a quarter of the AMD's. It loses both head-to-head tests by double-digit percentages. However, its lower 185 W TDP and shorter 229 mm length make it an option for systems with power or space limits. It also has tensor cores, which the AMD lacks, but the database has no benchmark results for tensor workloads, so no conclusion can be drawn from that feature.

For anyone needing a general-purpose GPU with strong compute and graphics, the W6800 wins on every measurable metric. The CMP 40HX is a niche product for a specific use case, and even then, its performance deficit makes it hard to recommend unless the physical constraints are absolute.

FAQ

Q: Which card has more memory bandwidth?

A: The AMD Radeon PRO W6800 has a 512.0 GB/s bandwidth, while the NVIDIA CMP 40HX has 448.0 GB/s.

Q: Can the NVIDIA CMP 40HX be used for display output?

A: No. The CMP 40HX has no display outputs. The AMD Radeon PRO W6800 includes 6x mini-DisplayPort 1.4a outputs.

Q: How much faster is the AMD in the OpenCL test?

A: The AMD Radeon PRO W6800 scores 121,808, which is 30.4% higher than the NVIDIA CMP 40HX's 93,395.

Q: Which card has more memory capacity?

A: The AMD Radeon PRO W6800 has 32 GB of GDDR6. The NVIDIA CMP 40HX has 8 GB of GDDR6.

Q: Do both cards support the same APIs?

A: Yes, both support DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4.

Q: Which card has a smaller physical footprint?

A: The NVIDIA CMP 40HX is shorter (229 mm versus 267 mm) and less wide (35 mm versus 50 mm). Its height is 111 mm versus 120 mm for the AMD.

Head-to-Head Benchmarks

The two recorded comparisons are both in the AMD's favor. The first is Geekbench OpenCL. The W6800 posts 121,808, and the CMP 40HX posts 93,395. The delta is 30.4%. To put that in perspective, the W6800's nearest rival in the overall database is the NVIDIA A10M at 135,230 (0.1% delta), while the CMP 40HX's nearest rival is the AMD Radeon PRO W7600 at 87,108 (1.7% delta). The W6800 is roughly 36% above the CMP 40HX's average rival score.

The second test is Geekbench Vulkan. The W6800 scores 109,961, and the CMP 40HX scores 77,879. The delta is 41.2%. This is the larger gap. The Vulkan test typically shows the AMD's architecture advantage in low-level API overhead and its higher shading unit count. The CMP 40HX has no Vulkan score above 78K, which places it far below the W6800's native compute pipeline.

Both tests confirm the same direction: the AMD is decisively ahead in compute workloads. The CMP 40HX's only theoretical edge is its tensor cores, but the database has no head-to-head test for tensor performance. In the two tests that are recorded, the W6800 wins by a margin that is typical of a card a tier higher.

Specification Differences

The two cards differ in nearly every major specification.

| Specification | AMD Radeon PRO W6800 | NVIDIA CMP 40HX |

| Architecture | RDNA 2.0 | Turing |

| Process node | 7 nm | 12 nm |

| Transistors | 26,800 million | 10,800 million |

| Die size | 520 mm² | 445 mm² |

| Base clock | 1575 MHz | 1470 MHz |

| Boost clock | 2322 MHz | 1650 MHz |

| Memory size | 32 GB | 8 GB |

| Memory type | GDDR6 | GDDR6 |

| Memory bus | 256 bit | 256 bit |

| Memory bandwidth | 512.0 GB/s | 448.0 GB/s |

| Shading units | 3840 | 2304 |

| TMUs | 240 | 144 |

| ROPs | 96 | 64 |

| RT cores | 60 | 36 |

| Tensor cores | None | 288 |

| FP32 performance | 17.83 TFLOPS | 7.603 TFLOPS |

| FP16 performance | 35.67 TFLOPS (2:1) | 15.21 TFLOPS (2:1) |

| TDP | 250 W | 185 W |

| Power connectors | 1x 6-pin + 1x 8-pin | 1x 8-pin |

| Suggested PSU | 600 W | 450 W |

| Bus interface | PCIe 4.0 x16 | PCIe 1.0 x4 |

| Display outputs | 6x mini-DisplayPort 1.4a | No outputs |

| Length | 267 mm (10.5 inches) | 229 mm (9 inches) |

| Height | 120 mm (4.7 inches) | 111 mm (4.4 inches) |

| Width | 50 mm (2 inches) | 35 mm (1.4 inches) |

| Launch MSRP | 2,249 USD | 699 USD |

The AMD has more of everything that affects compute and graphics performance: shading units, TMUs, ROPs, RT cores, memory capacity, bandwidth, and clock speed. The NVIDIA has tensor cores and a lower power draw. It also has a smaller physical footprint. The process node is a two-generation gap: 7 nm versus 12 nm, which explains the transistor density difference. The bus interface is another major split: PCIe 4.0 x16 versus PCIe 1.0 x4, a drastic difference for any workload that relies on data transfer to the GPU.

DETAILED SPECIFICATIONS

SPECIFICATION
PRO W6800
CMP 40HX
Core Specs
Shading Units
3,840
2,304 -40.0%
Shaders
3,840
2,304 -40.0%
TMUs
240
144 -40.0%
ROPs
96
64 -33.3%
Compute Units
60
SM Count
36
Clocks
Base Clock
1575 MHz
1470 MHz
Boost Clock
2322 MHz
1650 MHz
Memory Clock
2000 MHz 16 Gbps effective
1750 MHz 14 Gbps effective
Memory
Memory Size
32 GB
8 GB
VRAM (MB)
32,768
8,192 -75.0%
Memory Type
GDDR6
GDDR6
Memory Bus
256 bit
256 bit
Bandwidth
512.0 GB/s
448.0 GB/s
Cache
L1 Cache
128 KB per Array
64 KB (per SM)
L2 Cache
4 MB
4 MB
L3 Cache
128 MB
L0 Cache
32 KB per WGP
Performance
Pixel Rate
222.9 GPixel/s
105.6 GPixel/s
Texture Rate
557.3 GTexel/s
237.6 GTexel/s
FP32 (TFLOPS)
17.83 TFLOPS
7.603 TFLOPS
FP64 (TFLOPS)
1,114.6 GFLOPS (1:16)
237.6 GFLOPS (1:32)
FP16 (TFLOPS)
35.67 TFLOPS (2:1)
15.21 TFLOPS (2:1)
AI/RT
RT Cores
60
36 -40.0%
Tensor Cores
288
Power
TDP
250 W
185 W
TDP (W)
250
185 -26.0%
Suggested PSU
600 W
450 W
Power Connectors
1x 6-pin + 1x 8-pin
1x 8-pin
Architecture
Architecture
RDNA 2.0
Turing
GPU Name
Navi 21
TU106
Generation
Radeon Pro Navi (Navi II Series)
Mining GPUs
Process Size
7 nm
12 nm
Transistors
26,800 million
10,800 million
Die Size
520 mm²
445 mm²
Foundry
TSMC
TSMC
Density
51.5M / mm²
24.3M / mm²
API Support
DirectX
12 Ultimate (12_2)
12 Ultimate (12_2)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
2.1
3.0
CUDA
7.5
Shader Model
6.8
6.8
Physical
Slot Width
Dual-slot
Dual-slot
Length
267 mm 10.5 inches
229 mm 9 inches
Height
120 mm 4.7 inches
111 mm 4.4 inches
Outputs
6x mini-DisplayPort 1.4a
No outputs
Bus Interface
PCIe 4.0 x16
PCIe 1.0 x4
Other
Launch Price
2,249 USD
699 USD
Production
End-of-life
End-of-life
Predecessor
Radeon Pro Vega
View Radeon PRO W6800 Details View CMP 40HX Details