NVIDIA CMP 30HX vs NVIDIA CMP 40HX Comparison
NVIDIA CMP 30HX
CMP 40HX
PERFORMANCE BENCHMARKS
Analysis: NVIDIA CMP 30HX vs NVIDIA CMP 40HX
Where Each One Wins
The NVIDIA CMP 40HX is the clear performance leader in every recorded benchmark category. The data shows the CMP 40HX wins both head-to-head tests, with its largest advantage appearing in OpenCL compute workloads. The CMP 30HX does not win any benchmark in the database; its role is that of a lower-power, lower-throughput alternative within the same Turing-based mining GPU family.
For compute-focused tasks measured by Geekbench OpenCL, the CMP 40HX dominates. Its score of 93,395 versus the CMP 30HX's 65,199 represents a 43.2% advantage. This gap is substantial and indicates the CMP 40HX is the better choice for raw compute throughput, which aligns with its larger chip and wider memory interface.
In Vulkan workloads, the CMP 40HX again comes out ahead, scoring 77,879 against the CMP 30HX's 62,484, a 24.6% lead. While the Vulkan gap is smaller than the OpenCL gap, it is still decisive. The CMP 30HX's higher boost clock of 1785 MHz compared to the CMP 40HX's 1650 MHz does not compensate for the latter's greater shading unit count and memory bandwidth.
The CMP 40HX also holds a higher percentile ranking among all GPUs. It sits at the 93rd percentile, while the CMP 30HX sits at the 89th. That four-point gap in the database's percentile distribution reinforces the CMP 40HX's position as the stronger part, even though both are end-of-life mining products with no display outputs.
Architecture Differences
Both cards use NVIDIA's Turing architecture and are fabricated on TSMC's 12 nm process, but they are built around different chips. The CMP 40HX uses the TU106 die, which contains 10,800 million transistors on a 445 mm² die, yielding a transistor density of 24.3 million per mm². The CMP 30HX uses the TU116 die, with 6,600 million transistors on a 284 mm² die, for a density of 23.2 million per mm².
The CMP 40HX has considerably more execution resources. It features 2,304 shading units, 144 texture mapping units, and 64 raster operations pipelines. The CMP 30HX is cut down to 1,408 shading units, 88 TMUs, and 48 ROPs. These differences directly explain the CMP 40HX's higher pixel rate of 105.6 GPixel/s versus 85.68 GPixel/s, and its texture rate of 237.6 GTexel/s versus 157.1 GTexel/s.
A notable architectural split is in ray tracing and tensor cores. The CMP 40HX includes 36 RT cores and 288 tensor cores, while the CMP 30HX has none listed for either. This makes the CMP 40HX a more complete Turing implementation, though the practical impact is limited because these are mining-specific cards with no display outputs. The API support reflects the hardware: the CMP 40HX supports DirectX 12 Ultimate (12_2), while the CMP 30HX only reaches DirectX 12 (12_1). Both cards support OpenGL 4.6 and Vulkan 1.4.
Memory configurations also differ. The CMP 40HX has 8 GB of GDDR6 on a 256-bit bus, delivering 448.0 GB/s of bandwidth. The CMP 30HX has 6 GB of GDDR6 on a 192-bit bus, providing 336.0 GB/s. Both run memory at 1750 MHz with 14 Gbps effective speed, so the bandwidth difference comes entirely from the bus width.
Power and physical specs differ as well. The CMP 40HX has a TDP of 185 W and a suggested PSU of 450 W, while the CMP 30HX draws 125 W and recommends a 300 W PSU. Both are dual-slot cards, use a single 8-pin power connector, and share identical dimensions: 229 mm long, 111 mm high, and 35 mm wide. Both use a PCIe 1.0 x4 bus interface and have no display outputs.
Head-to-Head Benchmarks
The database records two head-to-head benchmarks between these cards, and the CMP 40HX wins both. The first is Geekbench OpenCL. Here the CMP 40HX scores 93,395, while the CMP 30HX scores 65,199. The delta is 43.2% in favor of the CMP 40HX. This is the largest margin between the two cards and reflects the CMP 40HX's superior compute throughput, driven by 2,304 shading units and nearly double the memory bandwidth (448.0 GB/s versus 336.0 GB/s).
The second benchmark is Geekbench Vulkan. The CMP 40HX scores 77,879, and the CMP 30HX scores 62,484. The CMP 40HX leads by 24.6%. The smaller Vulkan gap compared to OpenCL suggests that the CMP 30HX's higher boost clock of 1785 MHz helps somewhat in API-bound workloads, but the CMP 40HX still holds a commanding lead.
Looking at the average benchmark scores, the CMP 40HX averages 85,637 across all recorded tests, while the CMP 30HX averages 63,842. That is a difference of roughly 34%. The CMP 40HX's nearest rivals in the database are the AMD Radeon PRO W7600 at 87,108 (1.7% ahead of the CMP 40HX) and the NVIDIA Quadro GP100 at 87,445 (2.1% ahead). The CMP 40HX is 4.4% ahead of the AMD Radeon PRO W6600 (81,995) and 5.8% ahead of the AMD Radeon Pro Vega 64X (80,959).
The CMP 30HX's nearest rivals are clustered closely around its average score. The AMD Radeon RX 9060 XT LP matches it exactly at 63,830 (0% delta), the AMD Radeon RX 7600M is 0.1% ahead at 63,775, and the AMD Radeon Pro Vega 56 is 0.2% ahead at 63,693. The AMD Radeon Pro WX 9100 is 0.6% behind at 64,212. This tight grouping shows the CMP 30HX is positioned at a performance tier where small differences matter, while the CMP 40HX sits in a more comfortable position above its nearest competitors.
The Verdict
The data is unambiguous: the NVIDIA CMP 40HX is the stronger card in every measured category. It wins both head-to-head benchmarks, delivers a 43.2% advantage in OpenCL and a 24.6% advantage in Vulkan, and holds a higher average benchmark score and percentile ranking. For any workload that relies on raw compute, the CMP 40HX is the correct pick.
The CMP 30HX should only be chosen when its lower power draw is the deciding factor. It uses 125 W versus the CMP 40HX's 185 W, and its suggested PSU is 300 W compared to 450 W. In a mining or compute environment where power density or thermal limits are tight, the CMP 30HX offers a usable performance level at substantially lower consumption. However, that is a narrow use case. The performance penalty is steep: 43.2% in OpenCL and 24.6% in Vulkan.
The CMP 40HX also benefits from a more complete feature set, including 36 RT cores and 288 tensor cores, which the CMP 30HX lacks entirely. While these features are largely irrelevant for mining, they make the CMP 40HX a more flexible compute device for any task that can leverage them.
Both cards are end-of-life products with no display outputs, so neither is suited for gaming or workstation display use. They are purely compute accelerators. Between the two, the CMP 40HX is the one that delivers the higher throughput and broader capability. The CMP 30HX is the efficiency choice, but the recorded data shows it sacrifices a large share of performance to achieve that efficiency.
FAQ
Q: Which card has the higher OpenCL benchmark score?
A: The NVIDIA CMP 40HX scores 93,395 in Geekbench OpenCL, while the CMP 30HX scores 65,199. The CMP 40HX leads by 43.2%.
Q: Does the CMP 30HX win any benchmark in the database?
A: No. The CMP 40HX wins both recorded head-to-head benchmarks: Geekbench OpenCL and Geekbench Vulkan.
Q: What is the memory bandwidth difference between the two cards?
A: The CMP 40HX has 448.0 GB/s of bandwidth from 8 GB of GDDR6 on a 256-bit bus. The CMP 30HX has 336.0 GB/s from 6 GB of GDDR6 on a 192-bit bus.
Q: Do these cards support ray tracing?
A: The CMP 40HX includes 36 RT cores and 288 tensor cores. The CMP 30HX has no RT cores or tensor cores listed.
Q: What is the power consumption difference?
A: The CMP 40HX has a TDP of 185 W with a suggested PSU of 450 W. The CMP 30HX has a TDP of 125 W with a suggested PSU of 300 W.
Q: Are these cards usable for display output?
A: No. Both cards have no display outputs and are designed for compute or mining workloads.
Specification Differences
| Specification | NVIDIA CMP 40HX | NVIDIA CMP 30HX |
|---|---|---|
| Chip | TU106 | TU116 |
| Transistors | 10,800 million | 6,600 million |
| Die Size | 445 mm² | 284 mm² |
| Transistor Density | 24.3M / mm² | 23.2M / mm² |
| Base Clock | 1470 MHz | 1530 MHz |
| Boost Clock | 1650 MHz | 1785 MHz |
| Memory Size | 8 GB | 6 GB |
| Memory Bus Width | 256 bit | 192 bit |
| Memory Bandwidth | 448.0 GB/s | 336.0 GB/s |
| Shading Units | 2304 | 1408 |
| TMUs | 144 | 88 |
| ROPs | 64 | 48 |
| RT Cores | 36 | None |
| Tensor Cores | 288 | None |
| Pixel Rate | 105.6 GPixel/s | 85.68 GPixel/s |
| Texture Rate | 237.6 GTexel/s | 157.1 GTexel/s |
| FP32 Performance | 7.603 TFLOPS | 5.027 TFLOPS |
| FP16 Performance | 15.21 TFLOPS (2:1) | 10.05 TFLOPS (2:1) |
| TDP | 185 W | 125 W |
| Suggested PSU | 450 W | 300 W |
| DirectX Support | 12 Ultimate (12_2) | 12 (12_1) |
| Launch MSRP | 699 USD | 799 USD |
| Average Benchmark Score | 85,637 | 63,842 |
| Percentile vs All GPUs | 93 | 89 |