AMD Radeon Pro VII vs NVIDIA CMP 40HX Comparison
AMD Radeon Pro VII
CMP 40HX
PERFORMANCE BENCHMARKS
Analysis: AMD Radeon Pro VII vs NVIDIA CMP 40HX
The AMD Radeon Pro VII and NVIDIA CMP 40HX occupy the same performance percentile (93rd) but are built for entirely different purposes. The Pro VII is a professional workstation card with 16 GB of HBM2 memory, a 4096-bit bus, and six display outputs, while the CMP 40HX is a mining-focused GPU with no display outputs, 8 GB of GDDR6, and a PCIe 1.0 x4 interface. Benchmark data shows a split decision: the NVIDIA card wins OpenCL by 3.5%, while the AMD card wins Vulkan by a substantial 19.2%. This guide walks through the data to explain where each card excels, what architectural choices drive those results, and which type of user should pick which.
Where Each One Wins
The benchmark results divide cleanly by API. In Geekbench OpenCL, the NVIDIA CMP 40HX scores 93,395 against the AMD Radeon Pro VII’s 90,148, a 3.5% advantage. This is a modest lead, but it is consistent with the CMP 40HX’s higher base clock (1470 MHz vs. 1400 MHz) and its Turing architecture’s strength in compute workloads that scale well across its 2304 shading units. The OpenCL result also places the CMP 40HX within 1.7% of the AMD Radeon PRO W7600 (87,108) and 2.1% of the NVIDIA Quadro GP100 (87,445), confirming it sits in a competitive compute tier despite its mining-oriented design.
The AMD Radeon Pro VII wins decisively in Geekbench Vulkan, scoring 92,862 versus 77,879 for the CMP 40HX. That 19.2% margin is the largest performance gap in this comparison. The Pro VII’s Vulkan advantage likely stems from its GCN 5.1 architecture’s mature Vulkan driver stack and its massive 4096-bit memory bus, which delivers 1.02 TB/s of bandwidth—more than double the CMP 40HX’s 448.0 GB/s. For workloads that are bandwidth-sensitive, such as large data transfers or high-resolution rendering, the Pro VII’s memory subsystem provides a clear edge.
Beyond these two head-to-head tests, the Pro VII’s average benchmark score of 97,131 outpaces the CMP 40HX’s 85,637 by roughly 13.4%. This aggregate figure reflects the Pro VII’s additional benchmark result in Geekbench Metal (108,383), a test the CMP 40HX does not participate in. The Pro VII’s nearest rival, the AMD Radeon RX 7900M, scores 97,487 (just 0.4% higher), while the NVIDIA Quadro RTX 6000 sits 4.7% above at 101,872. The CMP 40HX, by contrast, trails its closest competitor, the AMD Radeon PRO W7600, by 1.7% and beats the AMD Radeon Pro Vega 64X by 5.8% (80,959).
Architecture Differences
The two cards are built on fundamentally different architectures from different eras. The AMD Radeon Pro VII uses the Vega 20 chip on the GCN 5.1 architecture, fabricated on a 7 nm TSMC process. This node packs 13,230 million transistors into a 331 mm² die, yielding a transistor density of 40.0 million per mm². The NVIDIA CMP 40HX uses the TU106 chip on the Turing architecture, fabricated on a 12 nm TSMC process. It contains 10,800 million transistors on a much larger 445 mm² die, resulting in a lower density of 24.3 million per mm². The smaller process node gives AMD a clear manufacturing advantage, allowing more transistors in less space.
Core configuration differs sharply. The Pro VII fields 3840 shading units, 240 texture mapping units, and 64 ROPs. The CMP 40HX has 2304 shading units, 144 TMUs, and 64 ROPs. The Pro VII’s 66.7% more shading units translate directly to higher compute throughput: 13.06 TFLOPS FP32 versus 7.603 TFLOPS for the CMP 40HX. The Pro VII also leads in texture rate (408.0 GTexel/s vs. 237.6 GTexel/s) and pixel rate (108.8 GPixel/s vs. 105.6 GPixel/s). However, the CMP 40HX includes 36 RT cores and 288 tensor cores, which the Pro VII lacks entirely—a nod to Turing’s real-time ray tracing and AI acceleration capabilities, even if those features are unused in mining workloads.
Memory architecture is where the cards diverge most dramatically. The Pro VII uses 16 GB of HBM2 on a 4096-bit bus, achieving 1.02 TB/s of bandwidth. The CMP 40HX uses 8 GB of GDDR6 on a 256-bit bus, delivering 448.0 GB/s. This 2.28x bandwidth advantage for AMD is the single largest specification gap between the two cards. Clock speeds are closer: the CMP 40HX has a higher base clock (1470 MHz vs. 1400 MHz), but the Pro VII boosts slightly higher (1700 MHz vs. 1650 MHz). Power consumption differs as well, with the Pro VII rated at 250 W TDP and the CMP 40HX at 185 W TDP.
Head-to-Head Benchmarks
The Geekbench OpenCL test shows the CMP 40HX winning with 93,395 points against the Pro VII’s 90,148, a 3.5% margin. This is the only test where the NVIDIA card comes out ahead. The result is notable because the Pro VII has significantly more shading units and over twice the memory bandwidth, yet the CMP 40HX still manages to edge it out. This suggests that OpenCL performance on this workload is more sensitive to clock speed and driver optimization than to raw core count or memory throughput. The CMP 40HX’s 1470 MHz base clock, 180 MHz higher than the Pro VII’s, likely contributes to this result.
The Geekbench Vulkan test flips the outcome decisively. The Pro VII scores 92,862, while the CMP 40HX manages only 77,879—a 19.2% gap in AMD’s favor. This is the largest delta in any benchmark between the two cards. Vulkan workloads often benefit from efficient memory access patterns, and the Pro VII’s 1.02 TB/s bandwidth provides a substantial advantage over the CMP 40HX’s 448.0 GB/s. Additionally, the Pro VII’s GCN architecture has been in the market since 2017, giving AMD time to mature its Vulkan drivers. The CMP 40HX, despite supporting Vulkan 1.4 (a newer version than the Pro VII’s 1.3), cannot overcome the bandwidth deficit in this test.
Looking at the broader benchmark context, the Pro VII’s Metal score of 108,383 is its strongest result, exceeding both its OpenCL and Vulkan scores. The CMP 40HX does not have a Metal score, which is expected given its lack of display outputs and macOS compatibility. Averaging across all available benchmarks, the Pro VII achieves 97,131 versus 85,637 for the CMP 40HX, a 13.4% overall advantage. The Pro VII’s nearest rival by average score is the AMD Radeon RX 7900M at 97,487 (0.4% above), while the CMP 40HX’s closest competitor is the AMD Radeon PRO W7600 at 87,108 (1.7% above the CMP).
FAQ
Q: Which card has higher raw compute throughput?
A: The AMD Radeon Pro VII delivers 13.06 TFLOPS FP32, while the NVIDIA CMP 40HX delivers 7.603 TFLOPS FP32. The Pro VII also leads in texture rate (408.0 GTexel/s vs. 237.6 GTexel/s) and pixel rate (108.8 GPixel/s vs. 105.6 GPixel/s).
Q: Why does the CMP 40HX win OpenCL despite having fewer cores?
A: The CMP 40HX has a higher base clock (1470 MHz vs. 1400 MHz) and scores 93,395 in OpenCL versus 90,148 for the Pro VII, a 3.5% margin. The result suggests OpenCL performance here favors clock speed and driver efficiency over core count or memory bandwidth.
Q: What is the memory bandwidth difference?
A: The Pro VII has 1.02 TB/s of bandwidth from 16 GB of HBM2 on a 4096-bit bus. The CMP 40HX has 448.0 GB/s from 8 GB of GDDR6 on a 256-bit bus. The Pro VII’s bandwidth is 2.28 times higher.
Q: Does the CMP 40HX support ray tracing?
A: Yes, the CMP 40HX includes 36 RT cores and 288 tensor cores, which the Pro VII does not have. However, the CMP 40HX has no display outputs, so these features cannot be used for visual rendering on a connected monitor.
Q: Which card has a better overall benchmark average?
A: The AMD Radeon Pro VII has an average benchmark score of 97,131, which is 13.4% higher than the CMP 40HX’s 85,637. The Pro VII’s additional Metal benchmark (108,383) contributes to this higher average.
Q: What are the power requirements for each card?
A: The Pro VII has a 250 W TDP and requires a 600 W power supply with a 1x 6-pin plus 1x 8-pin connector setup. The CMP 40HX has a 185 W TDP, requires a 450 W power supply, and uses a single 8-pin connector.
Specification Differences
| Specification | AMD Radeon Pro VII | NVIDIA CMP 40HX |
|---|---|---|
| Architecture | GCN 5.1 | Turing |
| Process Node | 7 nm | 12 nm |
| Transistors | 13,230 million | 10,800 million |
| Die Size | 331 mm² | 445 mm² |
| Transistor Density | 40.0M / mm² | 24.3M / mm² |
| Base Clock | 1400 MHz | 1470 MHz |
| Boost Clock | 1700 MHz | 1650 MHz |
| Memory Size | 16 GB | 8 GB |
| Memory Type | HBM2 | GDDR6 |
| Memory Bus Width | 4096 bit | 256 bit |
| Memory Bandwidth | 1.02 TB/s | 448.0 GB/s |
| Shading Units | 3840 | 2304 |
| TMUs | 240 | 144 |
| RT Cores | None | 36 |
| Tensor Cores | None | 288 |
| FP32 Performance | 13.06 TFLOPS | 7.603 TFLOPS |
| FP16 Performance | 26.11 TFLOPS (2:1) | 15.21 TFLOPS (2:1) |
| TDP | 250 W | 185 W |
| Power Connectors | 1x 6-pin + 1x 8-pin | 1x 8-pin |
| Suggested PSU | 600 W | 450 W |
| Bus Interface | PCIe 4.0 x16 | PCIe 1.0 x4 |
| Display Outputs | 6x mini-DisplayPort 1.4a | No outputs |
| DirectX Support | 12 (12_1) | 12 Ultimate (12_2) |
| Vulkan Support | 1.3 | 1.4 |
| Length | 305 mm (12 inches) | 229 mm (9 inches) |
| Width | Not specified | 35 mm (1.4 inches) |
| Release Date | 2020-05-12 | 2021-02-24 |
| Launch MSRP | 1,899 USD | 699 USD |
The Verdict
The benchmark data supports a clear split: the AMD Radeon Pro VII is the better choice for general-purpose compute and any workload that leverages Vulkan or Metal, while the NVIDIA CMP 40HX has a narrow edge in OpenCL but otherwise lacks the feature set for professional use. The Pro VII’s 19.2% Vulkan win and its 13.4% higher average benchmark score make it the stronger overall performer. Its 16 GB of HBM2 memory with 1.02 TB/s bandwidth is a massive advantage for memory-bound tasks, and its six display outputs enable multi-monitor professional workflows. The CMP 40HX, with no display outputs and a PCIe 1.0 x4 interface, is purpose-built for mining, which its 3.5% OpenCL win does little to offset.
For a professional user needing a workstation card with high-bandwidth memory, multiple display outputs, and strong Vulkan/Metal performance, the Radeon Pro VII is the data-backed choice—even at its higher launch MSRP of 1,899 USD versus 699 USD for the CMP 40HX. For a mining operation that prioritizes OpenCL compute and lower power draw (185 W vs. 250 W), the CMP 40HX offers a modest performance win in that specific API. However, the Pro VII’s superior average score and broader feature set make it the more versatile card for anything beyond dedicated mining. The CMP 40HX’s RT and tensor cores are present but irrelevant without display outputs, reinforcing its narrow mining focus. The data shows the Pro VII wins on versatility and overall performance; the CMP 40HX wins only on OpenCL efficiency and price.