NVIDIA CMP 70HX vs NVIDIA P104-100 Comparison

NVIDIA
GEFORCE

NVIDIA CMP 70HX

CORE STATE GA104
VRAM 8 GB
CLOCK SPEED 1395 MHz
TDP
BUS WIDTH 256 bit
ARCHITECTURE Ampere
nm
PROCESS 8 nm
LAUNCH DATE
VS
NVIDIA
GEFORCE

P104-100

CORE STATE GP104
VRAM 4 GB
CLOCK SPEED 1733 MHz
TDP
BUS WIDTH 256 bit
ARCHITECTURE Pascal
nm
PROCESS 16 nm
LAUNCH DATE 2017

PERFORMANCE BENCHMARKS

geekbench_opencl
25,135
52,368
geekbench_vulkan
35,817
45,165
3dmark_3dmark_steel_nomad_dx12
N/A
1,413

Analysis: NVIDIA CMP 70HX vs NVIDIA P104-100

# NVIDIA P104-100 vs NVIDIA CMP 70HX

Both of these cards come from NVIDIA’s dedicated mining lineup, share a dual-slot, 267 mm form factor, and ship with no display outputs. They even require the same 200 W suggested PSU. But that is where the similarity ends. The P104-100, built on Pascal, and the CMP 70HX, built on Ampere, occupy very different positions in the benchmark hierarchy, and the data tells a surprisingly one-sided story despite the newer architecture on the CMP 70HX’s side.

Head-to-Head Benchmarks

The P104-100 wins both head-to-head tests, and in one case, it wins by a staggering margin. In Geekbench OpenCL, the P104-100 scores 52,368 against the CMP 70HX’s 25,135. That is a 108.3% advantage — meaning the older Pascal card more than doubles the OpenCL throughput of the newer Ampere card. This is not a small edge or a driver quirk; it is a decisive generational gap in compute performance for this specific workload.

The Vulkan results tell a similar story, though with a narrower margin. The P104-100 posts 45,165 in Geekbench Vulkan, while the CMP 70HX scores 35,817. The delta here is 26.1% in favor of the P104-100. Still a clear win, but the gap closes considerably compared to the OpenCL test. This suggests that the CMP 70HX’s architecture handles Vulkan’s compute paths relatively better than OpenCL, though it still cannot catch the Pascal card.

The average benchmark scores reinforce this trend. The P104-100 carries an average score of 32,982 across all its benchmark entries, while the CMP 70HX averages 30,476. That is roughly an 8% overall lead for the P104-100. Interestingly, the CMP 70HX’s nearest rival list shows it trading blows with the AMD Radeon RX 6700 (0.1% advantage), the AMD Radeon RX 6800 (1.3% ahead), and the NVIDIA GeForce RTX 3070 Ti (1.8% ahead). Meanwhile, the P104-100 sits within 0.7% of the NVIDIA T600 Mobile, NVIDIA T550 Mobile, and GeForce RTX 3050 Mobile. Both cards are clustered with mainstream mobile and mid-range desktop parts, but the P104-100 is consistently at the top of its cluster while the CMP 70HX sits near the bottom of its group.

Architecture Differences

The foundational difference is the node and architecture generation. The P104-100 uses the GP104 chip on TSMC’s 16 nm process, while the CMP 70HX uses the GA104 chip on Samsung’s 8 nm node. This explains the dramatic transistor count difference: the GA104 packs 17,400 million transistors across a 392 mm² die, versus 7,200 million on a 314 mm² die for the GP104. The transistor density jumps from 22.9M per mm² on the Pascal card to 44.4M per mm² on the Ampere card — nearly double the packing density.

The CMP 70HX’s newer architecture brings additional hardware that the P104-100 lacks entirely. The CMP 70HX has 30 ray tracing cores and 120 tensor cores, while the P104-100 has none. The shader count also doubles: 3,840 shading units on the CMP 70HX versus 1,920 on the P104-100. Both cards keep the same 120 texture mapping units and 64 ROPs, so those counts are identical.

Clock speeds favor the older card. The P104-100 runs at a 1607 MHz base and 1733 MHz boost, while the CMP 70HX is set to 1365 MHz base and 1395 MHz boost. That is roughly 15% lower boost clock on the Ampere card. Memory also differs significantly. The P104-100 uses 4 GB of GDDR5X on a 256-bit bus, running at 10 Gbps effective for 320.3 GB/s bandwidth. The CMP 70HX doubles the capacity to 8 GB of GDDR6X on the same 256-bit bus, but at 19 Gbps effective, yielding 608.3 GB/s — nearly double the bandwidth.

Despite the CMP 70HX’s higher memory bandwidth and double the shader count, its FP32 compute is only about 61% higher at 10.71 TFLOPS versus 6.655 TFLOPS on the P104-100. The FP16 situation is starkly different: the P104-100 delivers just 104.0 GFLOPS (1:64 ratio), while the CMP 70HX delivers 10.71 TFLOPS (1:1 ratio). That is a 100x difference in FP16 throughput, which matters for certain compute workloads.

The API support also diverges. The P104-100 supports DirectX 12 (12_1), while the CMP 70HX supports DirectX 12 Ultimate (12_2). Both offer OpenGL 4.6 and Vulkan 1.4. The power connectors differ: the P104-100 uses a single 8-pin, while the CMP 70HX uses a single 12-pin. The CMP 70HX also has a listed height of 112 mm (4.4 inches), while the P104-100’s height is not specified.

The Verdict

The benchmark data is unambiguous: the P104-100 is the better performer in every measured test. It wins 2 out of 2 head-to-head benchmarks, and its average score is 8% higher. If you are choosing strictly on compute throughput in OpenCL and Vulkan, the Pascal card is the clear pick despite being from an older generation with less memory and lower bandwidth.

However, the CMP 70HX is not without reasons to exist. It has double the VRAM (8 GB vs 4 GB), which matters for workloads that exceed the 4 GB limit. It also has ray tracing and tensor cores, plus vastly superior FP16 performance (10.71 TFLOPS vs 104.0 GFLOPS) and double the FP32 throughput. The data shows the CMP 70HX is slower in the specific benchmarks recorded, but those benchmarks do not exercise the Ampere card’s unique features.

Looking at percentile rankings, the P104-100 sits at the 77th percentile of all GPUs, while the CMP 70HX sits at the 75th. That is a narrow gap in the broader GPU landscape, but the head-to-head results show it is a real one. The P104-100’s nearest rivals (T600 Mobile, T550 Mobile, RTX 3050 Mobile) are all within 0.7% of its average score, while the CMP 70HX’s rivals (Tesla M60, RX 6700, RX 6800, RTX 3070 Ti) span a 1.8% range. The P104-100 is punching near the top of its weight class; the CMP 70HX is mid-pack in its own.

For a pure compute task measured by these benchmarks, choose the P104-100. For workloads needing more memory, FP16 throughput, or hardware ray tracing/tensor cores, the CMP 70HX has capabilities the data here does not capture. But on the recorded metrics, the verdict is simple: the older card wins.

FAQ

Q: Which card has the higher average benchmark score?

A: The NVIDIA P104-100 has an average benchmark score of 32,982, while the NVIDIA CMP 70HX scores 30,476 — a lead of about 8% for the P104-100.

Q: How large is the performance gap in Geekbench OpenCL?

A: The P104-100 scores 52,368 versus the CMP 70HX’s 25,135, which is a 108.3% advantage for the Pascal card. It more than doubles the Ampere card’s OpenCL score.

Q: Does the CMP 70HX win any benchmark?

A: No. In the recorded head-to-head benchmarks, the P104-100 wins both Geekbench OpenCL and Geekbench Vulkan. The CMP 70HX has 0 wins out of 2 tests.

Q: What is the memory capacity difference?

A: The CMP 70HX has 8 GB of GDDR6X memory, while the P104-100 has 4 GB of GDDR5X. Both use a 256-bit bus, but the CMP 70HX’s bandwidth is 608.3 GB/s versus 320.3 GB/s on the P104-100.

Q: How do the shading unit counts compare?

A: The CMP 70HX has 3,840 shading units, exactly double the P104-100’s 1,920. Both cards have 120 TMUs and 64 ROPs.

Q: Do these cards have display outputs?

A: No. Both the P104-100 and CMP 70HX are listed as having no display outputs, consistent with their mining-focused design.

Where Each One Wins

The P104-100 wins in raw compute throughput as measured by the available benchmarks. Its Geekbench OpenCL score is more than double the CMP 70HX’s, and its Vulkan score is 26.1% higher. It also carries a higher percentile ranking (77th vs 75th) and a higher average benchmark score. If your workload is dominated by OpenCL compute or Vulkan compute, the data strongly favors the Pascal card. The P104-100 also runs at higher clock speeds (1733 MHz boost vs 1395 MHz boost), which may contribute to its performance advantage in these tests.

The CMP 70HX wins in memory capacity and bandwidth, offering 8 GB versus 4 GB and 608.3 GB/s versus 320.3 GB/s. It also wins decisively in FP16 compute: 10.71 TFLOPS versus 104.0 GFLOPS, which is roughly a 100x difference. For FP32, the CMP 70HX delivers 10.71 TFLOPS compared to 6.655 TFLOPS on the P104-100. It also has hardware ray tracing cores (30) and tensor cores (120) that the P104-100 lacks entirely. The CMP 70HX supports DirectX 12 Ultimate, while the P104-100 only supports DirectX 12 (12_1).

In terms of transistor density, the CMP 70HX is far ahead at 44.4M per mm² versus 22.9M per mm², reflecting its newer 8 nm Samsung process versus the 16 nm TSMC node. The CMP 70HX also has a larger die (392 mm² vs 314 mm²) and more than double the transistors (17,400 million vs 7,200 million).

For practical use, the P104-100 is the pick for anyone running the specific benchmark workloads recorded, where it consistently outperforms the CMP 70HX. The CMP 70HX is the pick for tasks that need more VRAM, higher memory bandwidth, FP16 acceleration, or ray tracing/tensor core features — none of which are measured in the recorded benchmarks. Both cards are end-of-life, dual-slot, 267 mm long, and require a 200 W PSU, so physical installation and power requirements are identical.

Specification Differences

| Specification | NVIDIA P104-100 | NVIDIA CMP 70HX |

|---|---|---|

| Architecture | Pascal | Ampere |

| Process Node | 16 nm (TSMC) | 8 nm (Samsung) |

| Transistors | 7,200 million | 17,400 million |

| Die Size | 314 mm² | 392 mm² |

| Transistor Density | 22.9M / mm² | 44.4M / mm² |

| Base Clock | 1607 MHz | 1365 MHz |

| Boost Clock | 1733 MHz | 1395 MHz |

| Memory | 4 GB GDDR5X | 8 GB GDDR6X |

| Memory Clock | 10 Gbps effective | 19 Gbps effective |

| Memory Bandwidth | 320.3 GB/s | 608.3 GB/s |

| Shading Units | 1920 | 3840 |

| RT Cores | None | 30 |

| Tensor Cores | None | 120 |

| FP32 Compute | 6.655 TFLOPS | 10.71 TFLOPS |

| FP16 Compute | 104.0 GFLOPS (1:64) | 10.71 TFLOPS (1:1) |

| Pixel Rate | 110.9 GPixel/s | 89.28 GPixel/s |

| Texture Rate | 208.0 GTexel/s | 167.4 GTexel/s |

| Power Connector | 1x 8-pin | 1x 12-pin |

| DirectX Support | 12 (12_1) | 12 Ultimate (12_2) |

| Height | Not specified | 112 mm (4.4 inches) |

DETAILED SPECIFICATIONS

SPECIFICATION
CMP 70HX
P104-100
Core Specs
Shading Units
3,840
1,920 -50.0%
Shaders
3,840
1,920 -50.0%
TMUs
120
120 0.0%
ROPs
64
64 0.0%
SM Count
30
15 -50.0%
Clocks
Base Clock
1365 MHz
1607 MHz
Boost Clock
1395 MHz
1733 MHz
Memory Clock
1188 MHz 19 Gbps effective
1251 MHz 10 Gbps effective
Memory
Memory Size
8 GB
4 GB
VRAM (MB)
8,192
4,096 -50.0%
Memory Type
GDDR6X
GDDR5X
Memory Bus
256 bit
256 bit
Bandwidth
608.3 GB/s
320.3 GB/s
Cache
L1 Cache
128 KB (per SM)
48 KB (per SM)
L2 Cache
4 MB
2 MB
Performance
Pixel Rate
89.28 GPixel/s
110.9 GPixel/s
Texture Rate
167.4 GTexel/s
208.0 GTexel/s
FP32 (TFLOPS)
10.71 TFLOPS
6.655 TFLOPS
FP64 (TFLOPS)
167.4 GFLOPS (1:64)
208.0 GFLOPS (1:32)
FP16 (TFLOPS)
10.71 TFLOPS (1:1)
104.0 GFLOPS (1:64)
AI/RT
RT Cores
30
Tensor Cores
120
Power
Suggested PSU
200 W
200 W
Power Connectors
1x 12-pin
1x 8-pin
Architecture
Architecture
Ampere
Pascal
GPU Name
GA104
GP104
Generation
Mining GPUs
Mining GPUs
Process Size
8 nm
16 nm
Transistors
17,400 million
7,200 million
Die Size
392 mm²
314 mm²
Foundry
Samsung
TSMC
Density
44.4M / mm²
22.9M / mm²
API Support
DirectX
12 Ultimate (12_2)
12 (12_1)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
3.0
3.0
CUDA
8.6
6.1
Shader Model
6.8
6.8
Physical
Slot Width
Dual-slot
Dual-slot
Length
267 mm 10.5 inches
267 mm 10.5 inches
Height
112 mm 4.4 inches
Outputs
No outputs
No outputs
Bus Interface
PCIe 1.0 x4
PCIe 1.0 x4
Other
Production
End-of-life
End-of-life
View CMP 70HX Details View P104-100 Details