AMD Radeon Pro W6600M vs NVIDIA CMP 40HX Comparison
AMD Radeon Pro W6600M
CMP 40HX
PERFORMANCE BENCHMARKS
Analysis: AMD Radeon Pro W6600M vs NVIDIA CMP 40HX
The Verdict
The recorded data clearly separates these two GPUs by role and workload. The NVIDIA CMP 40HX is the stronger compute part in raw benchmark terms, with an average benchmark score of 85637 against 61896 for the AMD Radeon Pro W6600M. That places the NVIDIA part in the 93rd percentile of all GPUs in the database, while the AMD mobile part sits in the 89th percentile. The CMP 40HX wins both recorded head-to-head tests, by 66.4% in Geekbench OpenCL and by 15.1% in Geekbench Vulkan.
However, the CMP 40HX has no display outputs, a PCIe 1.0 x4 bus interface, and is explicitly part of a Mining GPUs generation. It was built for a single purpose, and the data reflects that. The AMD Radeon Pro W6600M is a mobile professional part with a 90 W TDP and PCIe 4.0 x16 connectivity, intended for portable workstations. The CMP 40HX launched on 2021-02-24 with a launch MSRP of 699 USD, while the W6600M arrived later on 2021-06-07 with no recorded launch MSRP. Neither part is in production; both are end-of-life.
For a user seeking maximum compute throughput in a fixed, non-display mining or compute rig, the CMP 40HX is the data-backed choice. For a mobile workstation needing professional Radeon Pro drivers, portable power envelopes, and standard PCIe integration, the W6600M is the only viable option between these two. There is no overlap in intended use cases.
Architecture Differences
The two chips diverge sharply at the process and microarchitecture level. The NVIDIA CMP 40HX uses the TU106 chip on a 12 nm TSMC process, with 10,800 million transistors on a 445 mm² die. The AMD Radeon Pro W6600M uses the Navi 23 chip on a 7 nm TSMC process, with 11,060 million transistors on a 237 mm² die. The transistor density tells the story: 24.3M / mm² for the NVIDIA part versus 46.7M / mm² for the AMD part. The AMD chip packs slightly more transistors into roughly half the silicon area.
Architecturally, the CMP 40HX is Turing, while the W6600M is RDNA 2.0. The NVIDIA part carries 2304 shading units, 144 TMUs, 64 ROPs, 36 RT cores, and 288 tensor cores. The AMD part has 1792 shading units, 112 TMUs, 64 ROPs, and 28 RT cores, with no tensor cores recorded. The compute rates are close: FP32 is 7.603 TFLOPS for the CMP 40HX versus 7.290 TFLOPS for the W6600M, and FP16 is 15.21 TFLOPS (2:1) versus 14.58 TFLOPS (2:1). Pixel rate favors the AMD part at 130.2 GPixel/s versus 105.6 GPixel/s, while texture rate is nearly identical at 237.6 GTexel/s versus 227.8 GTexel/s.
Memory architecture differs substantially. Both use 8 GB of GDDR6, but the CMP 40HX runs a 256 bit bus with 448.0 GB/s bandwidth, while the W6600M uses a 128 bit bus with 224.0 GB/s bandwidth. Memory clock is the same at 1750 MHz, or 14 Gbps effective. The CMP 40HX therefore has exactly twice the memory bandwidth of the AMD part, a critical factor in compute workloads.
Power and physical design are opposites. The CMP 40HX consumes 185 W, requires a 1x 8-pin power connector, a suggested 450 W PSU, and occupies a dual-slot, 229 mm (9 inches) long, 111 mm (4.4 inches) high, 35 mm (1.4 inches) wide card. The W6600M is an IGP (integrated graphics processor) with a 90 W TDP, no power connectors, no recorded dimensions, and display outputs described as portable device dependent. The CMP 40HX has no display outputs at all.
The bus interface is another major split: PCIe 1.0 x4 for the CMP 40HX versus PCIe 4.0 x16 for the W6600M. That PCIe 1.0 x4 link is a severe bottleneck for any workload that requires host interaction, which reinforces the mining-specific design of the CMP 40HX. Both parts support DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4, so API compatibility is identical.
FAQ
Q: Which GPU has higher memory bandwidth?
A: The NVIDIA CMP 40HX. Its 256 bit memory bus delivers 448.0 GB/s, exactly double the 224.0 GB/s of the AMD Radeon Pro W6600M, which uses a 128 bit bus.
Q: Which GPU is better in Vulkan compute?
A: The NVIDIA CMP 40HX. In Geekbench Vulkan, it scores 77879 versus 67652 for the AMD part, a 15.1% advantage.
Q: Can the NVIDIA CMP 40HX be used for display output?
A: No. The CMP 40HX has no display outputs. The AMD Radeon Pro W6600M has outputs described as portable device dependent, meaning they are determined by the host laptop.
Q: Which GPU has the higher transistor density?
A: The AMD Radeon Pro W6600M. Its Navi 23 chip on 7 nm packs 46.7M transistors per mm², versus 24.3M / mm² for the TU106 in the CMP 40HX on 12 nm.
Q: Are these GPUs still in production?
A: No. Both are end-of-life. The NVIDIA CMP 40HX was released on 2021-02-24, and the AMD Radeon Pro W6600M was released on 2021-06-07.
Q: Which GPU has tensor cores?
A: The NVIDIA CMP 40HX has 288 tensor cores. The AMD Radeon Pro W6600M has no recorded tensor cores.
Specification Differences
The following fields differ between the NVIDIA CMP 40HX and AMD Radeon Pro W6600M:
- Chip: TU106 versus Navi 23
- Architecture: Turing versus RDNA 2.0
- Generation: Mining GPUs versus Radeon Pro Mobile (W6x00M)
- Process node: 12 nm versus 7 nm
- Transistors: 10,800 million versus 11,060 million
- Die size: 445 mm² versus 237 mm²
- Transistor density: 24.3M / mm² versus 46.7M / mm²
- Base clock: 1470 MHz versus 1224 MHz
- Boost clock: 1650 MHz versus 2034 MHz
- Memory bus width: 256 bit versus 128 bit
- Memory bandwidth: 448.0 GB/s versus 224.0 GB/s
- Shading units: 2304 versus 1792
- TMUs: 144 versus 112
- RT cores: 36 versus 28
- Tensor cores: 288 versus null
- Pixel rate: 105.6 GPixel/s versus 130.2 GPixel/s
- Texture rate: 237.6 GTexel/s versus 227.8 GTexel/s
- FP32: 7.603 TFLOPS versus 7.290 TFLOPS
- FP16: 15.21 TFLOPS (2:1) versus 14.58 TFLOPS (2:1)
- TDP: 185 W versus 90 W
- Slot width: Dual-slot versus IGP
- Power connectors: 1x 8-pin versus None
- Suggested PSU: 450 W versus null
- Bus interface: PCIe 1.0 x4 versus PCIe 4.0 x16
- Display outputs: No outputs versus Portable Device Dependent
- Dimensions: 229 mm (9 inches) by 111 mm (4.4 inches) by 35 mm (1.4 inches) versus null
- Release date: 2021-02-24 versus 2021-06-07
- Predecessor: null versus FirePro Mobile
- Launch MSRP: 699 USD versus null
Shared specifications include 8 GB GDDR6 memory, 64 ROPs, 1750 MHz memory clock (14 Gbps effective), DirectX 12 Ultimate (12_2), OpenGL 4.6, Vulkan 1.4, and end-of-life production status.
Head-to-Head Benchmarks
The database records two benchmark comparisons between these GPUs, and the NVIDIA CMP 40HX wins both. The first is Geekbench OpenCL, where the CMP 40HX scores 93395 against 56140 for the Radeon Pro W6600M. That is a 66.4% delta, the largest gap in either test. The CMP 40HX also has a higher average benchmark score of 85637, which sits 4.4% above the AMD Radeon PRO W6600 desktop part (81995) and 5.8% above the AMD Radeon Pro Vega 64X (80959). Its nearest rival overall is the AMD Radeon PRO W7600 at 87108, which beats it by 1.7%, and the NVIDIA Quadro GP100 at 87445, which beats it by 2.1%.
The second test is Geekbench Vulkan, where the CMP 40HX scores 77879 against 67652 for the W6600M, a 15.1% advantage. That margin is smaller than in OpenCL, but still decisive. The W6600M's average benchmark score of 61896 places it just 0.3% below the AMD Radeon 8050S (62108), and 2.6% above both the NVIDIA GeForce RTX 4090 (60347) and the Intel Arc Pro A60 (60326), with the AMD Radeon Pro Vega 48 (60140) trailing by 2.9%. The data shows the W6600M sits in a tight mid-range cluster, while the CMP 40HX competes at a higher performance tier.
The OpenCL result is worth emphasizing. A 66.4% lead in a compute API test is not incremental; it reflects the combined effect of double the memory bandwidth (448.0 GB/s versus 224.0 GB/s), a wider 256 bit bus, more shading units (2304 versus 1792), and a higher FP32 throughput (7.603 TFLOPS versus 7.290 TFLOPS). The Vulkan gap of 15.1% is smaller, suggesting that the RDNA 2.0 architecture narrows the distance when using Vulkan's lower-level execution model, but the CMP 40HX still holds a comfortable lead.
The pixel rate is the one metric where the W6600M clearly wins: 130.2 GPixel/s versus 105.6 GPixel/s, a 23.3% advantage, driven by its higher 2034 MHz boost clock against 1650 MHz for the NVIDIA part. Texture rate is almost a tie, with 237.6 GTexel/s for the CMP 40HX and 227.8 GTexel/s for the W6600M. The CMP 40HX also leads in FP16 compute at 15.21 TFLOPS versus 14.58 TFLOPS, a 4.3% margin.
In absolute terms, the CMP 40HX is the faster GPU in the database's recorded tests, but the W6600M is the more flexible and efficient component. The 90 W TDP of the AMD part is less than half the 185 W of the NVIDIA card, and the mobile form factor eliminates the need for external power connectors. The CMP 40HX cannot drive a display, and its PCIe 1.0 x4 interface would throttle data-intensive host communication. The W6600M's PCIe 4.0 x16 link is the modern standard. The benchmark scores favor the CMP 40HX, but the specification sheet favors the W6600M for any role involving graphics output, portability, or system integration.