NVIDIA CMP 40HX vs NVIDIA Quadro P6000 Comparison

NVIDIA
GEFORCE

NVIDIA CMP 40HX

CORE STATE TU106
VRAM 8 GB
CLOCK SPEED 1650 MHz
TDP 185 W
BUS WIDTH 256 bit
ARCHITECTURE Turing
nm
PROCESS 12 nm
LAUNCH DATE 2021
VS
NVIDIA
GEFORCE

Quadro P6000

CORE STATE GP102
VRAM 24 GB
CLOCK SPEED 1645 MHz
TDP 250 W
BUS WIDTH 384 bit
ARCHITECTURE Pascal
nm
PROCESS 16 nm
LAUNCH DATE 2016

PERFORMANCE BENCHMARKS

geekbench_opencl
93,395
66,382
geekbench_vulkan
77,879
73,590

Analysis: NVIDIA CMP 40HX vs NVIDIA Quadro P6000

The NVIDIA CMP 40HX and NVIDIA Quadro P6000 represent two very different design philosophies from the same manufacturer, separated by roughly four and a half years of architectural evolution. The data shows a clear, if nuanced, picture: the newer mining-focused card dominates in the available compute benchmarks, while the older professional workstation card counters with massive memory capacity and a vastly different feature set. The benchmark results, however, tell only part of the story, as the architectural and specification gaps are equally revealing.

Head-to-Head Benchmarks

The head-to-head data is unambiguous in favor of the CMP 40HX. In the Geekbench OpenCL test, the CMP 40HX scores 93,395, while the Quadro P6000 manages 66,382. This represents a 40.7% advantage for the Turing-based card, a substantial lead that underscores the generational leap in compute efficiency. The gap narrows considerably in the Geekbench Vulkan test, where the CMP 40HX scores 77,879 against the P6000’s 73,590, a 5.8% margin. This smaller delta in Vulkan suggests that the Pascal architecture’s raw throughput can still compete when the workload aligns with its strengths, though it still loses the round.

Looking at the aggregate data, the CMP 40HX holds an average benchmark score of 85,637, placing it in the 93rd percentile of all GPUs. The Quadro P6000, by contrast, averages 69,986, which lands it in the 90th percentile. The CMP 40HX’s nearest rivals include the AMD Radeon PRO W7600, which scores 87,108 and trails by -1.7%, and the NVIDIA Quadro GP100 at 87,445, which is -2.1% behind. The Quadro P6000’s competitive set is different; it sits just 0.2% ahead of the AMD Radeon Pro WX 8200 (69,870) and -0.2% behind the NVIDIA RTX A3000 Mobile (70,140). These percentile and rival comparisons suggest that while the CMP 40HX punches well above its weight, the P6000 is firmly a mid-pack professional card in raw compute terms, even in its own generation.

The two benchmark wins for the CMP 40HX are decisive in OpenCL but closer in Vulkan. The OpenCL result is particularly striking because it suggests that the CMP 40HX’s shader count, while lower than the P6000’s, is far better utilized per unit. The data shows a 40.7% score improvement, yet the CMP 40HX has fewer shading units (2,304 vs 3,840). This implies that the Turing architecture’s instruction efficiency and memory subsystem more than compensate for the raw unit deficit.

Architecture Differences

The architectural gap between these two cards is generational. The CMP 40HX is built on the Turing architecture using a 12 nm process at TSMC, with the TU106 chip containing 10,800 million transistors on a 445 mm² die, yielding a transistor density of 24.3M / mm². The Quadro P6000 uses the older Pascal architecture, fabricated on a 16 nm TSMC node, with the GP102 chip housing 11,800 million transistors across 471 mm², for a density of 25.1M / mm². Interestingly, the P6000 has slightly higher raw transistor count and die size, but the newer process node on the CMP 40HX allows for higher clock speeds and new hardware features.

Clock behavior differs notably. The CMP 40HX has a base clock of 1470 MHz and a boost clock of 1650 MHz, while the P6000 runs at 1506 MHz base and 1645 MHz boost. The CMP 40HX’s boost advantage of 5 MHz is negligible, but its base clock is 36 MHz lower, meaning the architectural efficiency must come from elsewhere. The memory configuration is another major split: the CMP 40HX uses 8 GB of GDDR6 on a 256-bit bus, delivering 448.0 GB/s of bandwidth, whereas the P6000 offers 24 GB of GDDR5X on a 384-bit bus at 432.8 GB/s. The CMP 40HX’s bandwidth advantage of 15.2 GB/s is modest, but the P6000’s triple memory capacity is a massive differentiator for large datasets.

The most profound differences lie in the compute and feature blocks. The CMP 40HX includes 36 RT cores and 288 tensor cores, enabling hardware ray tracing and AI acceleration, while the P6000 has none of these, with values marked as null. The FP32 throughput tells a similar tale: the CMP 40HX delivers 7.603 TFLOPS vs the P6000’s 12.63 TFLOPS, a 66% advantage for the older card. However, in FP16, the CMP 40HX achieves 15.21 TFLOPS (2:1) while the P6000 is limited to 197.4 GFLOPS (1:64), meaning the Turing card is orders of magnitude faster in half-precision workloads. The API support also diverges: the CMP 40HX supports DirectX 12 Ultimate (12_2) while the P6000 only reaches DirectX 12 (12_1), though both support OpenGL 4.6 and Vulkan 1.4.

The Verdict

The data points to a clear split based on workload type. For compute-centric tasks, particularly those leveraging FP16, ray tracing, or tensor operations, the NVIDIA CMP 40HX is the superior choice. Its 40.7% OpenCL win and 5.8% Vulkan win, combined with RT and tensor core support, make it a more future-proof compute engine. The CMP 40HX’s average score of 85,637 versus the P6000’s 69,986 reinforces this, as does its higher percentile rank (93rd vs 90th).

However, the Quadro P6000 is not without merit. Its 24 GB of memory is three times larger than the CMP 40HX’s 8 GB, and its 384-bit bus, while slightly slower in bandwidth, offers higher capacity for massive texture sets or large model datasets. For tasks that require heavy data residency without frequent transfers, the P6000’s memory pool is a decisive factor. Additionally, the P6000 has display outputs (1x DVI, 4x DisplayPort 1.4a) while the CMP 40HX has none, making the former the only viable option for any visual output. The P6000’s FP32 throughput of 12.63 TFLOPS also exceeds the CMP 40HX’s 7.603 TFLOPS, so single-precision float-heavy workloads will favor the older card.

FAQ

Q: Which card scores higher in Geekbench OpenCL?

A: The NVIDIA CMP 40HX scores 93,395 versus the Quadro P6000’s 66,382, a 40.7% advantage.

Q: Does the Quadro P6000 have more memory than the CMP 40HX?

A: Yes, the P6000 has 24 GB of GDDR5X, while the CMP 40HX has 8 GB of GDDR6.

Q: What is the difference in FP16 performance?

A: The CMP 40HX delivers 15.21 TFLOPS in FP16, while the P6000 manages only 197.4 GFLOPS, a massive gap favoring the Turing card.

Q: Which card has a higher average benchmark score?

A: The CMP 40HX averages 85,637 across benchmarks, compared to the P6000’s 69,986.

Q: Do both cards support ray tracing?

A: No, only the CMP 40HX has 36 RT cores; the P6000 has no RT core support.

Q: Can either card output to a display?

A: Only the Quadro P6000 has display outputs (1x DVI, 4x DisplayPort 1.4a); the CMP 40HX has no outputs.

Where Each One Wins

The NVIDIA CMP 40HX wins decisively in compute-heavy synthetic benchmarks. It leads in both OpenCL and Vulkan tests, with the OpenCL gap being particularly large at 40.7%. Its architecture includes RT cores and tensor cores, which are absent from the P6000, making it the only choice for ray-traced rendering or AI inference workloads. The 12 nm Turing process also allows for more efficient FP16 processing, with 15.21 TFLOPS versus the P6000’s meager 197.4 GFLOPS, so any half-precision workload will heavily favor the CMP 40HX. Its higher bandwidth (448.0 GB/s vs 432.8 GB/s) also gives it a slight edge in memory-throughput-bound tasks.

The NVIDIA Quadro P6000 wins on memory capacity and raw single-precision compute. Its 24 GB frame buffer is triple the CMP 40HX’s 8 GB, enabling workloads that require massive in-memory datasets without spillover. Its FP32 performance of 12.63 TFLOPS exceeds the CMP 40HX’s 7.603 TFLOPS, meaning traditional graphics and compute tasks that rely on FP32 will run faster on the P6000. It also provides display connectivity, which the CMP 40HX lacks entirely, making it the only option for workstation use with monitors. The P6000’s PCIe 3.0 x16 interface is also far more standard than the CMP 40HX’s PCIe 1.0 x4, which could bottleneck data transfers in some systems.

Specification Differences

The two cards differ across nearly every core specification. The CMP 40HX is built on Turing at 12 nm, while the P6000 uses Pascal at 16 nm. The CMP 40HX has 2,304 shading units, 144 TMUs, and 64 ROPs; the P6000 has 3,840 shading units, 240 TMUs, and 96 ROPs. The CMP 40HX includes 36 RT cores and 288 tensor cores, while the P6000 has none. Clock speeds are close, with the CMP 40HX at 1470 MHz base and 1650 MHz boost, versus the P6000’s 1506 MHz base and 1645 MHz boost. Memory differs substantially: 8 GB GDDR6 on a 256-bit bus with 448.0 GB/s bandwidth for the CMP 40HX, versus 24 GB GDDR5X on a 384-bit bus with 432.8 GB/s for the P6000. Pixel rate is 105.6 GPixel/s for the CMP 40HX and 157.9 GPixel/s for the P6000; texture rate is 237.6 GTexel/s vs 394.8 GTexel/s. FP32 is 7.603 TFLOPS vs 12.63 TFLOPS, and FP16 is 15.21 TFLOPS vs 197.4 GFLOPS. Power draw is 185 W for the CMP 40HX and 250 W for the P6000, with the former suggesting a 450 W PSU and the latter a 600 W PSU. The CMP 40HX uses a PCIe 1.0 x4 interface and has no display outputs, while the P6000 uses PCIe 3.0 x16 and offers 1x DVI, 4x DisplayPort 1.4a. The CMP 40HX supports DirectX 12 Ultimate, but the P6000 only supports DirectX 12 (12_1); both support OpenGL 4.6 and Vulkan 1.4. The physical dimensions also differ, with the CMP 40HX at 229 mm length and the P6000 at 267 mm.

DETAILED SPECIFICATIONS

SPECIFICATION
CMP 40HX
Quadro P6000
Core Specs
Shading Units
2,304
3,840 +66.7%
Shaders
2,304
3,840 +66.7%
TMUs
144
240 +66.7%
ROPs
64
96 +50.0%
SM Count
36
30 -16.7%
Clocks
Base Clock
1470 MHz
1506 MHz
Boost Clock
1650 MHz
1645 MHz
Memory Clock
1750 MHz 14 Gbps effective
1127 MHz 9 Gbps effective
Memory
Memory Size
8 GB
24 GB
VRAM (MB)
8,192
24,576 +200.0%
Memory Type
GDDR6
GDDR5X
Memory Bus
256 bit
384 bit
Bandwidth
448.0 GB/s
432.8 GB/s
Cache
L1 Cache
64 KB (per SM)
48 KB (per SM)
L2 Cache
4 MB
3 MB
Performance
Pixel Rate
105.6 GPixel/s
157.9 GPixel/s
Texture Rate
237.6 GTexel/s
394.8 GTexel/s
FP32 (TFLOPS)
7.603 TFLOPS
12.63 TFLOPS
FP64 (TFLOPS)
237.6 GFLOPS (1:32)
394.8 GFLOPS (1:32)
FP16 (TFLOPS)
15.21 TFLOPS (2:1)
197.4 GFLOPS (1:64)
AI/RT
RT Cores
36
Tensor Cores
288
Power
TDP
185 W
250 W
TDP (W)
185
250 +35.1%
Suggested PSU
450 W
600 W
Power Connectors
1x 8-pin
1x 8-pin
Architecture
Architecture
Turing
Pascal
GPU Name
TU106
GP102
Generation
Mining GPUs
Quadro Pascal (Px000)
Process Size
12 nm
16 nm
Transistors
10,800 million
11,800 million
Die Size
445 mm²
471 mm²
Foundry
TSMC
TSMC
Density
24.3M / mm²
25.1M / mm²
API Support
DirectX
12 Ultimate (12_2)
12 (12_1)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
3.0
3.0
CUDA
7.5
6.1
Shader Model
6.8
6.8
Physical
Slot Width
Dual-slot
Dual-slot
Length
229 mm 9 inches
267 mm 10.5 inches
Height
111 mm 4.4 inches
111 mm 4.4 inches
Outputs
No outputs
1x DVI4x DisplayPort 1.4a
Bus Interface
PCIe 1.0 x4
PCIe 3.0 x16
Other
Launch Price
699 USD
5,999 USD
Production
End-of-life
End-of-life
Predecessor
Quadro Maxwell
Successor
Quadro Volta
View CMP 40HX Details View Quadro P6000 Details