NVIDIA CMP 40HX vs NVIDIA P102-100 Comparison

NVIDIA
GEFORCE

NVIDIA CMP 40HX

CORE STATE TU106
VRAM 8 GB
CLOCK SPEED 1650 MHz
TDP 185 W
BUS WIDTH 256 bit
ARCHITECTURE Turing
nm
PROCESS 12 nm
LAUNCH DATE 2021
VS
NVIDIA
GEFORCE

P102-100

CORE STATE GP102
VRAM 5 GB
CLOCK SPEED 1683 MHz
TDP 250 W
BUS WIDTH 320 bit
ARCHITECTURE Pascal
nm
PROCESS 16 nm
LAUNCH DATE 2018

PERFORMANCE BENCHMARKS

geekbench_opencl
93,395
49,602
geekbench_vulkan
77,879
67,454

Analysis: NVIDIA CMP 40HX vs NVIDIA P102-100

Head-to-Head Benchmarks

The recorded data shows a decisive overall victory for the NVIDIA CMP 40HX, which takes both head-to-head benchmark wins. The largest margin appears in the Geekbench OpenCL test, where the CMP 40HX scores 93,395 against the P102-100’s 49,602. That represents an 88.3% delta, a substantial lead that places the CMP 40HX in a different performance tier for compute-heavy workloads. In the Geekbench Vulkan test, the gap narrows considerably but still favors the CMP 40HX: it scores 77,879 versus 67,454 for the P102-100, a 15.5% advantage. The Vulkan result is more competitive, suggesting the P102-100’s Pascal architecture handles the API’s workload characteristics relatively better than it does OpenCL, but it still cannot overcome the CMP 40HX’s overall compute throughput.

The average benchmark score reinforces this picture. The CMP 40HX posts an average of 85,637, while the P102-100 averages 58,528. That difference translates into a percentile ranking gap as well: the CMP 40HX sits at the 93rd percentile among all GPUs in the database, while the P102-100 ranks at the 88th percentile. The CMP 40HX’s nearest rivals include the AMD Radeon PRO W7600 at an average score of 87,108 (a delta of -1.7% relative to the CMP 40HX) and the NVIDIA Quadro GP100 at 87,445 (delta -2.1%). The P102-100, by contrast, sits just below the AMD Radeon PRO V710, which averages 58,657 (delta -0.2%), and just above the AMD Radeon RX 6950 XT at 58,392 (delta 0.2%). These rival positions indicate that the CMP 40HX competes in a much higher performance bracket, while the P102-100 is clustered with mid-range cards from a later generation.

The head-to-head delta figures are worth parsing carefully. An 88.3% lead in OpenCL is not a marginal difference; it suggests the CMP 40HX delivers nearly double the raw compute performance in that specific API. The Vulkan delta of 15.5% is more modest but still a clear win. Across the two tests, the CMP 40HX wins both, giving it a 2-0 record in the head-to-head comparison. No benchmark in the database shows the P102-100 ahead.

Where Each One Wins

The CMP 40HX wins in every measured category, but the degree of advantage varies by workload type. In OpenCL, which often stresses raw FP32 throughput and memory bandwidth for general-purpose compute, the CMP 40HX’s Turing architecture with 2,304 shading units and 36 RT cores delivers a massive lead. The 88.3% delta in OpenCL is the clearest indicator that compute-heavy tasks, such as hashing algorithms or scientific workloads, will strongly favor the CMP 40HX. Its FP32 rating of 7.603 TFLOPS, while lower than the P102-100’s 10.77 TFLOPS on paper, does not translate into a loss in the actual benchmark; the CMP 40HX still wins OpenCL by a wide margin, suggesting that driver optimization and architectural efficiency matter more than raw theoretical throughput in this test.

The Vulkan test, which often exercises graphics pipeline features and driver overhead, shows a closer contest. The CMP 40HX’s 15.5% lead here indicates that the P102-100’s Pascal architecture, despite being older, is not entirely outclassed in graphics-adjacent workloads. However, the CMP 40HX still holds the advantage, likely due to its support for DirectX 12 Ultimate (12_2) and Vulkan 1.4, whereas the P102-100 only reaches DirectX 12 (12_1) with the same Vulkan 1.4 version. The CMP 40HX also has a higher memory bandwidth at 448.0 GB/s versus 440.3 GB/s, which may contribute to its Vulkan lead.

For use-case segmentation, the data suggests the CMP 40HX is the better choice for any application that relies on OpenCL compute, where its lead is overwhelming. The P102-100, despite its lower average score, does have a higher FP32 rating (10.77 TFLOPS) and more shading units (3,200 versus 2,304), which could theoretically benefit workloads that scale well with raw shader count but are not well represented in the Geekbench suite. The database does not include such tests, so this remains a qualitative observation. In Vulkan-based scenarios, the CMP 40HX is still ahead, but the smaller delta means the P102-100 is closer to competitive, and users with Pascal-optimized code might see less of a penalty. Overall, the CMP 40HX is the superior performer in every recorded benchmark, with the OpenCL result being the standout differentiator.

Architecture Differences

The two GPUs come from different NVIDIA architectures and process nodes. The CMP 40HX uses the TU106 chip built on Turing architecture at a 12 nm TSMC process, while the P102-100 uses the GP102 chip on Pascal architecture at a 16 nm TSMC process. This generational gap explains several key differences. The CMP 40HX packs 10,800 million transistors on a 445 mm² die, giving a transistor density of 24.3 million per mm². The P102-100 has 11,800 million transistors on a slightly larger 471 mm² die, resulting in a density of 25.1 million per mm². Despite having fewer transistors, the CMP 40HX achieves higher benchmark scores, indicating that Turing’s architectural improvements over Pascal outweigh the raw transistor count advantage of the P102-100.

The CMP 40HX includes 36 RT cores and 288 tensor cores, features that are entirely absent from the P102-100, which has no RT cores and no tensor cores. These specialized units enable hardware-accelerated ray tracing and AI workloads on the CMP 40HX, though the mining-focused nature of both cards means these features are secondary. The shading unit count differs significantly: the CMP 40HX has 2,304 shading units, 144 texture mapping units, and 64 ROPs, while the P102-100 has 3,200 shading units, 200 TMUs, and 80 ROPs. The P102-100’s higher counts in these traditional rasterization metrics give it higher theoretical pixel and texture rates: 134.6 GPixel/s and 336.6 GTexel/s versus 105.6 GPixel/s and 237.6 GTexel/s for the CMP 40HX. Yet these theoretical advantages do not translate into benchmark wins, as the CMP 40HX’s newer architecture and driver support prove more effective in practice.

Memory configurations also diverge. The CMP 40HX uses 8 GB of GDDR6 on a 256-bit bus, running at 1750 MHz with 14 Gbps effective speed, yielding 448.0 GB/s of bandwidth. The P102-100 uses 5 GB of GDDR5X on a 320-bit bus, running at 1376 MHz with 11 Gbps effective speed, yielding 440.3 GB/s of bandwidth. The CMP 40HX has more capacity and slightly more bandwidth, despite a narrower bus, due to the higher memory clock. The FP16 capabilities are starkly different: the CMP 40HX delivers 15.21 TFLOPS FP16 at a 2:1 ratio, while the P102-100 delivers only 168.3 GFLOPS FP16 at a 1:64 ratio. This makes the CMP 40HX dramatically better for mixed-precision compute, a common requirement in AI inference and some mining algorithms.

Power and physical specifications also differ. The CMP 40HX has a TDP of 185 W with a single 8-pin power connector and a suggested PSU of 450 W, while the P102-100 has a TDP of 250 W with two 8-pin connectors and a suggested PSU of 600 W. The CMP 40HX is shorter at 229 mm (9 inches) versus 267 mm (10.5 inches) for the P102-100, though both are dual-slot cards. The P102-100 has no recorded height or width in the database, while the CMP 40HX is 111 mm (4.4 inches) tall and 35 mm (1.4 inches) wide. Both cards have no display outputs and use a PCIe 1.0 x4 bus interface, which is unusual and reflects their mining-oriented design. The CMP 40HX’s API support includes DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4, while the P102-100 only reaches DirectX 12 (12_1) with the same OpenGL and Vulkan versions.

FAQ

Q: Which GPU has the higher average benchmark score?

A: The NVIDIA CMP 40HX has an average benchmark score of 85,637, while the NVIDIA P102-100 averages 58,528. The CMP 40HX also ranks at the 93rd percentile among all GPUs, compared to the P102-100’s 88th percentile.

Q: How much faster is the CMP 40HX in OpenCL?

A: In the Geekbench OpenCL test, the CMP 40HX scores 93,395 versus the P102-100’s 49,602, a delta of 88.3%. This is the largest performance gap between the two cards in any recorded benchmark.

Q: Does the P102-100 win any benchmark?

A: No. The head-to-head data shows the CMP 40HX winning both the OpenCL and Vulkan tests, giving it a 2-0 record. The P102-100 has zero wins in the recorded benchmarks.

Q: What are the memory specifications of each card?

A: The CMP 40HX has 8 GB of GDDR6 on a 256-bit bus with a bandwidth of 448.0 GB/s. The P102-100 has 5 GB of GDDR5X on a 320-bit bus with a bandwidth of 440.3 GB/s.

Q: Which card has more shading units?

A: The P102-100 has 3,200 shading units, while the CMP 40HX has 2,304. However, the CMP 40HX still achieves higher benchmark scores despite the lower count, thanks to its Turing architecture and features like RT cores and tensor cores.

Q: What is the TDP of each card?

A: The CMP 40HX has a TDP of 185 W with a single 8-pin power connector, while the P102-100 has a TDP of 250 W with two 8-pin connectors. The suggested PSU is 450 W for the CMP 40HX and 600 W for the P102-100.

Specification Differences

The two cards differ in nearly every core specification category. The CMP 40HX uses a TU106 chip on a 12 nm TSMC process with 10,800 million transistors on a 445 mm² die, while the P102-100 uses a GP102 chip on a 16 nm TSMC process with 11,800 million transistors on a 471 mm² die. The CMP 40HX has a base clock of 1470 MHz and a boost clock of 1650 MHz, whereas the P102-100 has a base clock of 1582 MHz and a boost clock of 1683 MHz. Memory differs by size, type, bus width, and clock: the CMP 40HX has 8 GB GDDR6 at 1750 MHz (14 Gbps effective) on a 256-bit bus, while the P102-100 has 5 GB GDDR5X at 1376 MHz (11 Gbps effective) on a 320-bit bus. The CMP 40HX’s bandwidth is 448.0 GB/s versus 440.3 GB/s for the P102-100.

Compute resources differ markedly: the CMP 40HX has 2,304 shading units, 144 TMUs, 64 ROPs, 36 RT cores, and 288 tensor cores, while the P102-100 has 3,200 shading units, 200 TMUs, and 80 ROPs with no RT or tensor cores. The pixel rate is 105.6 GPixel/s for the CMP 40HX versus 134.6 GPixel/s for the P102-100, and the texture rate is 237.6 GTexel/s versus 336.6 GTexel/s. FP32 performance is 7.603 TFLOPS for the CMP 40HX and 10.77 TFLOPS for the P102-100, but FP16 performance is drastically different: 15.21 TFLOPS (2:1) for the CMP 40HX versus 168.3 GFLOPS (1:64) for the P102-100.

Power and physical specifications also diverge: the CMP 40HX has a TDP of 185 W, one 8-pin connector, and a suggested PSU of 450 W, while the P102-100 has a TDP of 250 W, two 8-pin connectors, and a suggested PSU of 600 W. The CMP 40HX is 229 mm long (9 inches), 111 mm tall (4.4 inches), and 35 mm wide (1.4 inches), while the P102-100 is 267 mm long (10.5 inches) with no recorded height or width. API support differs for DirectX: the CMP 40HX supports 12 Ultimate (12_2), while the P102-100 only supports 12 (12_1); both support OpenGL 4.6 and Vulkan 1.4. The release dates are also different, with the P102-100 launching in February 2018 and the CMP 40HX in February 2021, both now end-of-life. The CMP 40HX has a launch MSRP of 699 USD, while the P102-100 has no recorded launch MSRP.

DETAILED SPECIFICATIONS

SPECIFICATION
CMP 40HX
P102-100
Core Specs
Shading Units
2,304
3,200 +38.9%
Shaders
2,304
3,200 +38.9%
TMUs
144
200 +38.9%
ROPs
64
80 +25.0%
SM Count
36
25 -30.6%
Clocks
Base Clock
1470 MHz
1582 MHz
Boost Clock
1650 MHz
1683 MHz
Memory Clock
1750 MHz 14 Gbps effective
1376 MHz 11 Gbps effective
Memory
Memory Size
8 GB
5 GB
VRAM (MB)
8,192
5,120 -37.5%
Memory Type
GDDR6
GDDR5X
Memory Bus
256 bit
320 bit
Bandwidth
448.0 GB/s
440.3 GB/s
Cache
L1 Cache
64 KB (per SM)
48 KB (per SM)
L2 Cache
4 MB
2.5 MB
Performance
Pixel Rate
105.6 GPixel/s
134.6 GPixel/s
Texture Rate
237.6 GTexel/s
336.6 GTexel/s
FP32 (TFLOPS)
7.603 TFLOPS
10.77 TFLOPS
FP64 (TFLOPS)
237.6 GFLOPS (1:32)
336.6 GFLOPS (1:32)
FP16 (TFLOPS)
15.21 TFLOPS (2:1)
168.3 GFLOPS (1:64)
AI/RT
RT Cores
36
Tensor Cores
288
Power
TDP
185 W
250 W
TDP (W)
185
250 +35.1%
Suggested PSU
450 W
600 W
Power Connectors
1x 8-pin
2x 8-pin
Architecture
Architecture
Turing
Pascal
GPU Name
TU106
GP102
Generation
Mining GPUs
Mining GPUs
Process Size
12 nm
16 nm
Transistors
10,800 million
11,800 million
Die Size
445 mm²
471 mm²
Foundry
TSMC
TSMC
Density
24.3M / mm²
25.1M / mm²
API Support
DirectX
12 Ultimate (12_2)
12 (12_1)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
3.0
3.0
CUDA
7.5
6.1
Shader Model
6.8
6.8
Physical
Slot Width
Dual-slot
Dual-slot
Length
229 mm 9 inches
267 mm 10.5 inches
Height
111 mm 4.4 inches
Outputs
No outputs
No outputs
Bus Interface
PCIe 1.0 x4
PCIe 1.0 x4
Other
Launch Price
699 USD
Production
End-of-life
End-of-life
View CMP 40HX Details View P102-100 Details