AMD Radeon Pro Vega II vs NVIDIA CMP 40HX Comparison
AMD Radeon Pro Vega II
CMP 40HX
PERFORMANCE BENCHMARKS
Analysis: AMD Radeon Pro Vega II vs NVIDIA CMP 40HX
# Head-to-Head Benchmarks
The benchmark data presents a clear picture: AMD Radeon Pro Vega II wins both head-to-head tests, but the margin varies dramatically depending on the API. In Geekbench OpenCL, the AMD card scores 99,048 against the NVIDIA CMP 40HX's 93,395, a modest 6.1% advantage. That is a meaningful lead, yet close enough to suggest that raw compute workloads in OpenCL do not fully separate these two very different designs.
The Geekbench Vulkan test tells a far more decisive story. The Radeon Pro Vega II posts 99,621, while the CMP 40HX manages only 77,879. That is a 27.9% gap — a crushing margin that indicates the AMD card's architectural advantages compound heavily in Vulkan workloads. Interestingly, the Radeon Pro Vega II's Vulkan score is nearly identical to its OpenCL result (99,621 vs 99,048), suggesting consistent performance across APIs. The CMP 40HX, by contrast, drops from 93,395 in OpenCL to 77,879 in Vulkan, a 16.6% regression within the same card. This inconsistency hints that the NVIDIA part's driver or hardware scheduler is significantly less efficient at translating Vulkan commands into execution.
Looking at the overall average benchmark scores, the AMD card achieves 109,617 across all tests, placing it in the 94th percentile of all GPUs. The CMP 40HX averages 85,637, sitting just one percentile lower at 93. That single-percentile difference underscores how competitive the field is around these two cards — the absolute scores differ by 28%, yet both rank among the top 7% of all GPUs ever tested.
The nearest-rival data adds context. The Radeon Pro Vega II sits within 3.8% of the NVIDIA RTX A5500 Mobile (113,944 average), and within 2.7% of its own dual-chip sibling, the Radeon Pro Vega II Duo (106,750). The CMP 40HX trails the AMD Radeon PRO W7600 by just 1.7% (87,108) and the NVIDIA Quadro GP100 by 2.1% (87,445). These proximity scores suggest both cards are well-positioned in their respective performance tiers, with the AMD card occupying a higher absolute tier.
# Architecture Differences
The two GPUs could hardly be more different in their underlying design philosophies. The AMD Radeon Pro Vega II uses the Vega 20 chip built on GCN 5.1 architecture, manufactured on TSMC's 7 nm process. The NVIDIA CMP 40HX employs the TU106 chip with Turing architecture, also from TSMC but on a mature 12 nm node. This process gap is substantial: the AMD die measures 331 mm² and packs 13,230 million transistors, yielding a density of 40.0 million per mm². The NVIDIA chip is physically larger at 445 mm² but contains fewer transistors — 10,800 million — giving it a lower density of 24.3 million per mm². The 7 nm node clearly enables the AMD design to cram nearly 23% more transistors into a 26% smaller die.
Memory architecture diverges completely. The Radeon Pro Vega II carries 32 GB of HBM2 on a 4096-bit bus, delivering 825.3 GB/s of bandwidth. The CMP 40HX settles for 8 GB of GDDR6 on a 256-bit bus, achieving 448.0 GB/s. That is an 84% bandwidth advantage for the AMD card and four times the capacity. The memory clock rates reflect the different technologies: the HBM2 runs at 806 MHz (1612 Mbps effective), while the GDDR6 operates at 1750 MHz (14 Gbps effective). Higher raw clock does not compensate for the narrower bus on the NVIDIA side.
Compute resources tell a similar tale. The Radeon Pro Vega II fields 4096 shading units, 256 texture mapping units, and 64 ROPs. The CMP 40HX counters with 2304 shading units, 144 TMUs, and 64 ROPs. The AMD card has 78% more shaders and 78% more TMUs, though ROP counts match. Peak FP32 throughput stands at 14.09 TFLOPS for AMD versus 7.603 TFLOPS for NVIDIA — an 85% difference. FP16 rates follow proportionally: 28.18 TFLOPS versus 15.21 TFLOPS, both at 2:1 ratios. However, the NVIDIA card brings dedicated hardware the AMD part lacks: 36 RT cores for ray tracing and 288 tensor cores for AI acceleration. These features are entirely absent from the Radeon Pro Vega II's spec sheet.
Pixel and texture fill rates align with these disparities. The AMD card achieves 110.1 GPixel/s and 440.3 GTexel/s, compared to 105.6 GPixel/s and 237.6 GTexel/s for NVIDIA. The pixel rates are nearly identical, which makes sense given the equal ROP counts and similar boost clocks (1720 MHz AMD vs 1650 MHz NVIDIA). Texture rate, however, heavily favors AMD due to the extra TMUs.
Power and physical design diverge sharply. The Radeon Pro Vega II draws up to 475 W and requires a quad-slot cooler, with an 850 W suggested PSU. The CMP 40HX sips 185 W, fits in a dual-slot, uses a single 8-pin connector, and suggests a 450 W PSU. The NVIDIA card's dimensions are specified at 229 mm length, 111 mm height, and 35 mm width; the AMD card's dimensions are not provided. The bus interfaces also differ: AMD uses Apple MPX, while NVIDIA uses PCIe 1.0 x4 — an oddly old and narrow interface for a mining card, though mining workloads rarely stress host bandwidth.
# FAQ
Q: Which card is faster in OpenCL compute?
A: The AMD Radeon Pro Vega II wins the Geekbench OpenCL test with a score of 99,048 versus 93,395 for the NVIDIA CMP 40HX, a 6.1% advantage.
Q: How much does the AMD card win in Vulkan?
A: The margin is 27.9% — the Radeon Pro Vega II scores 99,621 while the CMP 40HX scores 77,879 in Geekbench Vulkan.
Q: Do these cards support ray tracing?
A: The NVIDIA CMP 40HX includes 36 RT cores for ray tracing acceleration. The AMD Radeon Pro Vega II has no RT cores listed, indicating no dedicated ray tracing hardware.
Q: What are the memory capacity and bandwidth differences?
A: The AMD card offers 32 GB of HBM2 with 825.3 GB/s bandwidth on a 4096-bit bus. The NVIDIA card provides 8 GB of GDDR6 with 448.0 GB/s on a 256-bit bus.
Q: Which card has the higher transistor density?
A: The AMD Radeon Pro Vega II achieves 40.0 million transistors per mm² on TSMC's 7 nm node, versus 24.3 million per mm² for the CMP 40HX on 12 nm.
Q: Are both cards still in production?
A: No. Both are marked as end-of-life. The AMD card launched on 2019-06-02, while the NVIDIA card launched later on 2021-02-24.
# Specification Differences
| Specification | AMD Radeon Pro Vega II | NVIDIA CMP 40HX |
|---|---|---|
| Architecture | GCN 5.1 | Turing |
| Process Node | 7 nm | 12 nm |
| Transistors | 13,230 million | 10,800 million |
| Die Size | 331 mm² | 445 mm² |
| Transistor Density | 40.0M / mm² | 24.3M / mm² |
| Base Clock | 1574 MHz | 1470 MHz |
| Boost Clock | 1720 MHz | 1650 MHz |
| Memory Clock | 806 MHz / 1612 Mbps effective | 1750 MHz / 14 Gbps effective |
| Memory Size | 32 GB | 8 GB |
| Memory Type | HBM2 | GDDR6 |
| Memory Bus Width | 4096 bit | 256 bit |
| Memory Bandwidth | 825.3 GB/s | 448.0 GB/s |
| Shading Units | 4096 | 2304 |
| TMUs | 256 | 144 |
| ROPs | 64 | 64 |
| RT Cores | None | 36 |
| Tensor Cores | None | 288 |
| Pixel Rate | 110.1 GPixel/s | 105.6 GPixel/s |
| Texture Rate | 440.3 GTexel/s | 237.6 GTexel/s |
| FP32 | 14.09 TFLOPS | 7.603 TFLOPS |
| FP16 | 28.18 TFLOPS (2:1) | 15.21 TFLOPS (2:1) |
| TDP | 475 W | 185 W |
| Slot Width | Quad-slot | Dual-slot |
| Power Connectors | Not specified | 1x 8-pin |
| Suggested PSU | 850 W | 450 W |
| Bus Interface | Apple MPX | PCIe 1.0 x4 |
| Display Outputs | 1x HDMI 2.0b, 4x Thunderbolt | No outputs |
| DirectX Support | 12 (12_1) | 12 Ultimate (12_2) |
| Vulkan Support | 1.3 | 1.4 |
| Release Date | 2019-06-02 | 2021-02-24 |
| Launch MSRP | 2,199 USD | 699 USD |
# The Verdict
The data points to two GPUs built for entirely different purposes, despite both being end-of-life products. The AMD Radeon Pro Vega II wins every benchmark where both are measured, and it wins decisively in Vulkan. Its 27.9% Vulkan advantage and 6.1% OpenCL lead are backed by a 28% higher average benchmark score (109,617 vs 85,637). The 94th percentile ranking versus 93rd confirms the AMD card sits slightly higher in the global performance distribution.
Yet the CMP 40HX is not without merit. It delivers 78% of the AMD card's average score while consuming only 39% of the power (185 W vs 475 W). For workloads that favor its Turing architecture — particularly those leveraging its 36 RT cores or 288 tensor cores — the NVIDIA part could close or even reverse the gap in scenarios not captured by Geekbench's compute tests. The DirectX 12 Ultimate support and Vulkan 1.4 compliance also give it a more modern API surface.
The AMD card's strengths are raw compute throughput, massive memory bandwidth, and API consistency. Its 32 GB HBM2 pool at 825.3 GB/s is a professional-grade specification that dwarfs the CMP 40HX's 8 GB GDDR6. But that capability comes at a steep cost: triple the power draw, a quad-slot cooler, and an Apple MPX bus interface that restricts system compatibility.
# Where Each One Wins
AMD Radeon Pro Vega II wins in every measured benchmark category. Its 14.09 TFLOPS FP32 throughput, 4096 shading units, and 825.3 GB/s bandwidth make it the clear choice for compute-heavy workloads like rendering, scientific simulation, or large dataset processing. The 32 GB memory capacity allows working with models and scenes that would exhaust the CMP 40HX's 8 GB frame buffer. The consistent performance across OpenCL and Vulkan (differing by only 0.6%) suggests predictable behavior across application types. The pixel rate advantage (110.1 vs 105.6 GPixel/s) gives it a slight edge in fill-rate-bound scenarios.
NVIDIA CMP 40HX wins in power efficiency and specialized features, despite losing every benchmark. At 185 W versus 475 W, it delivers competitive OpenCL performance (within 6.1%) while using far less power, making it suitable for multi-GPU mining rigs where thermal and power budgets are constrained. The 36 RT cores and 288 tensor cores enable ray tracing and AI acceleration that the AMD card simply cannot perform — workloads like real-time ray-traced rendering or neural network inference would run on the NVIDIA card but not at all on the AMD part. The dual-slot form factor and single 8-pin connector also simplify installation compared to the AMD card's quad-slot requirement. Its Vulkan 1.4 support is newer than the AMD card's 1.3, and DirectX 12 Ultimate offers more advanced DX12 features. For mining-specific tasks — the card's stated generation purpose — the lower power draw and no-display-output design are deliberate advantages. The 8 GB memory, while smaller, is ample for many cryptographic workloads that do not require large frame buffers.