NVIDIA CMP 70HX vs NVIDIA Tesla P4 Comparison

NVIDIA
GEFORCE

NVIDIA CMP 70HX

CORE STATE GA104
VRAM 8 GB
CLOCK SPEED 1395 MHz
TDP
BUS WIDTH 256 bit
ARCHITECTURE Ampere
nm
PROCESS 8 nm
LAUNCH DATE
VS
NVIDIA
GEFORCE

Tesla P4

CORE STATE GP104
VRAM 8 GB
CLOCK SPEED 1114 MHz
TDP 75 W
BUS WIDTH 256 bit
ARCHITECTURE Pascal
nm
PROCESS 16 nm
LAUNCH DATE 2016

PERFORMANCE BENCHMARKS

geekbench_opencl
25,135
34,947
geekbench_vulkan
35,817
40,309

Analysis: NVIDIA CMP 70HX vs NVIDIA Tesla P4

# NVIDIA Tesla P4 vs NVIDIA CMP 70HX: Database Analysis

The NVIDIA Tesla P4 and NVIDIA CMP 70HX occupy different segments of the GPU landscape, and the benchmark data reflects this clearly. The Tesla P4, a Pascal-era compute card, records an average benchmark score of 37,628 across the database's test suite. The CMP 70HX, built on the Ampere architecture for mining workloads, posts a considerably lower average score of 30,476. This places the Tesla P4 in the 81st percentile of all GPUs tested, while the CMP 70HX sits in the 75th percentile. The performance gap, as measured by the aggregate score, is approximately 19% in favor of the Tesla P4. In the two head-to-head comparisons available in the database, the Tesla P4 wins both: it leads by 39% in the Geekbench OpenCL test and by 12.5% in the Geekbench Vulkan test. There are no recorded benchmark wins for the CMP 70HX in these direct comparisons, making the Tesla P4 the clear performance leader in this pairing.

Head-to-Head Benchmarks

The Geekbench OpenCL benchmark is a strong indicator of raw compute throughput, and here the Tesla P4 demonstrates a decisive advantage. The Tesla P4 scores 34,947 points, while the CMP 70HX scores 25,135 points. This represents a 39% lead for the Tesla P4, a significant margin that indicates a substantial difference in compute capabilities in this workload. In the Vulkan benchmark, the Tesla P4 again comes out ahead, posting 40,309 points against 35,817 for the CMP 70HX, a 12.5% advantage. This pattern suggests that the Tesla P4's architecture is better suited to these general-purpose compute tasks. Across both workloads, the Tesla P4's advantage is consistent and substantial. The CMP 70HX, while still capable, does not match the Tesla P4 in these compute-oriented tests. For buyers prioritizing raw compute performance in general-purpose benchmarks, the Tesla P4 is the better performer by a wide margin.

Architecture Differences

The architectural differences between these two GPUs are significant and explain the performance gap. The Tesla P4 is built on the GP104 chip, using the Pascal architecture, and was manufactured on TSMC's 16 nm process node. It is a relatively compact design with a die size of 314 mm². The CMP 70HX, on the other hand, uses the GA104 chip with the newer Ampere architecture, fabricated by Samsung on an 8nm process. The CMP's die is larger at 392 mm², and its transistor count is far higher at 17,400 million versus 7,200 million for the Tesla P4. This newer process and larger chip help the CMP 70HX reach higher clock speeds, with a base clock of 1365 MHz and a boost clock of 1395 MHz, compared to the Tesla P4's base of 886 MHz and boost of 1114 MHz. This clock speed advantage, however, does not translate into benchmark performance due to other architectural differences.

The memory subsystems also differ. The Tesla P4 is equipped with 8 GB of GDDR5 memory on a 256-bit bus, delivering 192.3 GB/s of bandwidth. The CMP 70HX also has 8 GB of memory, but uses faster GDDR6X type memory, and its 256-bit bus delivers a much higher 608.3 GB/s of memory bandwidth. The memory capacity is the same, but the CMP's newer memory type provides a significant bandwidth advantage in theory, even if the benchmarks do not show it translating into a win. The Tesla P4 has 2560 shading units, 160 texture mapping units, and 64 ROPs. The CMP 70HX substantially increases these: 3840 shaders, 120 TMUs, and 64 ROPs. The larger shader and TMU counts are typical of a more recent generation and should improve geometry throughput. Additionally, the CMP 70HX has a much higher pixel rate of 89.28 GPixel/s and texture rate of 167.4 GTexel/s, compared to the Tesla P4's 71.30 GPixel/s and 178.2 GTexel/s. In FP32 compute, the Tesla P4 is rated at 5.704 TFLOPS, while the CMP 70HX is rated substantially higher at 10.71 TFLOPS. The newer CMP architecture provides a much higher raw FP32 throughput figure. As a result, the Tesla P4's overall benchmark wins despite the CMP's theoretical advantages in many compute metrics.

The Verdict

Based strictly on the recorded measurements, the NVIDIA Tesla P4 is the decisive winner. It holds a higher percentile among all tested GPUs (81st vs 75th), a higher average benchmark score (34,947 vs 30,476), and it wins both head-to-head benchmark comparisons outright. The CMP 70HX does have theoretical advantages on paper: a newer, denser architecture, higher clocks, more shaders, and a higher-rated FP32 throughput. Yet, these theoretical specifications do not show up in the Geekbench workloads. The Tesla P4 is the better card for general-purpose compute as recorded by these tests. The CMP 70HX is clearly designed for a specific workload, and the data shows it does not compete with the Tesla P4 in the database's general-purpose benchmark suite.

FAQ

Q: Which GPU is faster in the Geekbench OpenCL benchmark?

A: The NVIDIA Tesla P4 is faster. It scores 34,947 points against the CMP 70HX's 25,135 points, a 39% higher score.

Q: What is the average benchmark score difference?

A: The Tesla P4's average score is 34,628, while the CMP 70HX's is 30,476. The Tesla P4 is 1.13 times faster on average.

Q: Does the CMP 70HX win in any benchmark?

A: No. In the two recorded head-to-head tests, the NVIDIA Tesla P4 wins both. There is no benchmark where the CMP 70HX comes out ahead.

Q: Which GPU offers more memory bandwidth?

A: The CMP 70HX has more memory bandwidth on paper. It uses GDDR6X memory on a 256-bit bus rated at 608.3 GB/s, while the Tesla P4 uses GDDR5 at 192.3 GB/s on the same bus width. The benchmark results, however, still favor the Tesla P4.

Q: What explains the performance difference?

A: The Tesla P4 is based on the older Pascal architecture, but the benchmark results show it performs better in the database's general compute suite. The CMP 70HX has a newer Ampere architecture with higher clock counts and more shaders, but this does not translate to higher performance in the measured benchmarks.

Q: Are these cards for mining?

A: The CMP 70HX is from NVIDIA's CMP series, which is designed for mining workloads. The Tesla P4 is more of a data-center compute or cloud GPU. The CMP has no display outputs, and the Tesla also does not have display outputs.

Where Each One Wins

The recorded benchmarks show a consistent winner in the NVIDIA Tesla P4. The data shows no benchmark in which the CMP 70HX has the edge. For any workload that resembles the Geekbench OpenCL and Vulkan workloads, the Tesla P4 is the one to pick. The CMP 70HX is not shown to be competitive in these general-purpose tests, and it would be more suited to specialized tasks that are not reflected in this benchmark data. The Tesla P4's higher performance percentile and higher average score make it the clear choice from the database's compute perspective.

DETAILED SPECIFICATIONS

SPECIFICATION
CMP 70HX
Tesla P4
Core Specs
Shading Units
3,840
2,560 -33.3%
Shaders
3,840
2,560 -33.3%
TMUs
120
160 +33.3%
ROPs
64
64 0.0%
SM Count
30
20 -33.3%
Clocks
Base Clock
1365 MHz
886 MHz
Boost Clock
1395 MHz
1114 MHz
Memory Clock
1188 MHz 19 Gbps effective
1502 MHz 6 Gbps effective
Memory
Memory Size
8 GB
8 GB
VRAM (MB)
8,192
8,192 0.0%
Memory Type
GDDR6X
GDDR5
Memory Bus
256 bit
256 bit
Bandwidth
608.3 GB/s
192.3 GB/s
Cache
L1 Cache
128 KB (per SM)
48 KB (per SM)
L2 Cache
4 MB
2 MB
Performance
Pixel Rate
89.28 GPixel/s
71.30 GPixel/s
Texture Rate
167.4 GTexel/s
178.2 GTexel/s
FP32 (TFLOPS)
10.71 TFLOPS
5.704 TFLOPS
FP64 (TFLOPS)
167.4 GFLOPS (1:64)
178.2 GFLOPS (1:32)
FP16 (TFLOPS)
10.71 TFLOPS (1:1)
89.12 GFLOPS (1:64)
AI/RT
RT Cores
30
Tensor Cores
120
Power
TDP
75 W
TDP (W)
75
Suggested PSU
200 W
250 W
Power Connectors
1x 12-pin
None
Architecture
Architecture
Ampere
Pascal
GPU Name
GA104
GP104
Generation
Mining GPUs
Tesla Pascal (Pxx)
Process Size
8 nm
16 nm
Transistors
17,400 million
7,200 million
Die Size
392 mm²
314 mm²
Foundry
Samsung
TSMC
Density
44.4M / mm²
22.9M / mm²
API Support
DirectX
12 Ultimate (12_2)
12 (12_1)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
3.0
3.0
CUDA
8.6
6.1
Shader Model
6.8
6.8
Physical
Slot Width
Dual-slot
Single-slot
Length
267 mm 10.5 inches
168 mm 6.6 inches
Height
112 mm 4.4 inches
Outputs
No outputs
No outputs
Bus Interface
PCIe 1.0 x4
PCIe 3.0 x16
Other
Production
End-of-life
End-of-life
Predecessor
Tesla Maxwell
Successor
Tesla Volta
View CMP 70HX Details View Tesla P4 Details