NVIDIA CMP 40HX vs NVIDIA Tesla V100 PCIe 16 GB Comparison

NVIDIA
GEFORCE

NVIDIA CMP 40HX

CORE STATE TU106
VRAM 8 GB
CLOCK SPEED 1650 MHz
TDP 185 W
BUS WIDTH 256 bit
ARCHITECTURE Turing
nm
PROCESS 12 nm
LAUNCH DATE 2021
VS
NVIDIA
GEFORCE

Tesla V100 PCIe 16 GB

CORE STATE GV100
VRAM 16 GB
CLOCK SPEED 1380 MHz
TDP 300 W
BUS WIDTH 4096 bit
ARCHITECTURE Volta
nm
PROCESS 12 nm
LAUNCH DATE 2017

PERFORMANCE BENCHMARKS

geekbench_opencl
93,395
163,063
geekbench_vulkan
77,879
113,062

Analysis: NVIDIA CMP 40HX vs NVIDIA Tesla V100 PCIe 16 GB

Head-to-Head Benchmarks

The recorded data shows a decisive overall win for the NVIDIA Tesla V100 PCIe 16 GB, which takes both head-to-head tests with a 2:0 margin. The largest gap appears in the Geekbench OpenCL test, where the Tesla V100 scores 163063 against the CMP 40HX's 93395, a 74.6% advantage. This is a substantial performance gap that suggests the V100's compute architecture delivers far more raw throughput in this workload. In the Geekbench Vulkan test, the V100 again leads with 113062 versus 77879, a 45.2% delta. While the gap narrows in Vulkan, the V100 still holds a commanding lead, indicating that its advantage is not limited to a single API or workload type.

Looking at the broader database context, the Tesla V100's average benchmark score is 138063, placing it in the 96th percentile among all GPUs. Its nearest rivals are extremely close: the NVIDIA Tesla V100 SXM2 32 GB sits just 0.2% higher, while the AMD Instinct MI100 trails by 0.7%. This means the PCIe 16 GB variant is essentially performance-equivalent to its SXM2 sibling in the database's aggregate scoring, despite the latter having double the memory capacity. The CMP 40HX, by contrast, posts an average score of 85637, placing it in the 93rd percentile. Its nearest rival, the AMD Radeon PRO W7600, is 1.7% ahead, while the NVIDIA Quadro GP100 is 2.1% ahead. The CMP 40HX does outpace the AMD Radeon PRO W6600 by 4.4% and the AMD Radeon Pro Vega 64X by 5.8%, showing it is competitive within its own tier, but that tier sits roughly 38% below the V100's average score.

The per-test deltas reinforce this hierarchy. In OpenCL, the V100's 74.6% advantage is roughly double its Vulkan lead, suggesting the V100's compute capabilities are especially pronounced in OpenCL-style workloads. The CMP 40HX's Vulkan score is 16.6% lower than its OpenCL score, while the V100's Vulkan score is 30.7% lower than its OpenCL score. This pattern indicates that both GPUs lose some performance in Vulkan relative to OpenCL, but the V100 loses more in percentage terms, possibly due to driver or architecture differences in handling the Vulkan API.

FAQ

Q: Which GPU has the higher average benchmark score?

A: The NVIDIA Tesla V100 PCIe 16 GB has an average benchmark score of 138063, while the NVIDIA CMP 40HX has an average score of 85637. The V100 outperforms the CMP 40HX by roughly 61% in aggregate terms.

Q: How does the V100 compare to its closest rival, the Tesla V100 SXM2 32 GB?

A: The PCIe 16 GB version scores 138063, which is just 0.2% lower than the SXM2 32 GB's average of 137731. The two are effectively tied in the database's measurements, despite the SXM2 having double the memory.

Q: What is the CMP 40HX's standing among its nearest competitors?

A: The CMP 40HX's average score of 85637 puts it 1.7% behind the AMD Radeon PRO W7600 and 2.1% behind the NVIDIA Quadro GP100. However, it leads the AMD Radeon PRO W6600 by 4.4% and the AMD Radeon Pro Vega 64X by 5.8%.

Q: Which GPU wins the Geekbench Vulkan test, and by how much?

A: The Tesla V100 wins with a score of 113062 versus the CMP 40HX's 77879, a 45.2% advantage. This is a significant lead, though smaller than the OpenCL gap.

Q: Do both GPUs perform better in OpenCL or Vulkan?

A: Both score higher in OpenCL. The V100 scores 163063 in OpenCL versus 113062 in Vulkan, and the CMP 40HX scores 93395 in OpenCL versus 77879 in Vulkan. The V100's relative drop in Vulkan is larger (30.7% lower) than the CMP 40HX's (16.6% lower).

Q: Are these GPUs still in production?

A: According to the database, both are marked as end-of-life products. The V100 was released in June 2017, while the CMP 40HX was released in February 2021.

Where Each One Wins

The Tesla V100 PCIe 16 GB wins in every measured benchmark category, so its strengths are in raw compute performance across both OpenCL and Vulkan. The data shows a 74.6% lead in OpenCL and a 45.2% lead in Vulkan, making it the clear choice for workloads that rely on these APIs. Its 96th percentile ranking among all GPUs, combined with an average score of 138063, places it among the top tier of accelerators. The V100's nearest rivals are all within 1.7% of its score, meaning it competes at the very highest level of compute performance, and it essentially matches the SXM2 32 GB variant despite having half the memory. For users prioritizing maximum compute throughput in OpenCL-heavy tasks, the V100 is the obvious pick based on the recorded measurements.

The CMP 40HX, while losing both head-to-head tests, still holds its own in the 93rd percentile of all GPUs. Its average score of 85637 is competitive within its immediate peer group, sitting just 1.7% behind the AMD Radeon PRO W7600 and 2.1% behind the NVIDIA Quadro GP100. It also leads the AMD Radeon PRO W6600 and AMD Radeon Pro Vega 64X by 4.4% and 5.8%, respectively. This suggests the CMP 40HX is a solid mid-tier performer, but the data does not show any test where it beats the V100. Its wins are relative to other mid-range cards, not against this particular rival. For any workload measured here, the V100 is the superior choice.

Specification Differences

The two GPUs differ across nearly every major specification in the database. The V100 features a GV100 chip under the Volta architecture, while the CMP 40HX uses a TU106 chip under the Turing architecture. Process nodes are identical at 12 nm from TSMC, but the transistor counts differ dramatically: the V100 has 21,100 million transistors on an 815 mm² die, while the CMP 40HX has 10,800 million on a 445 mm² die. The V100's transistor density is 25.9M per mm² versus 24.3M for the CMP 40HX.

Clock speeds favor the CMP 40HX, which runs at a base of 1470 MHz and boost of 1650 MHz, while the V100 runs at 1245 MHz base and 1380 MHz boost. Memory configurations are also very different: the V100 has 16 GB of HBM2 on a 4096-bit bus with 897.0 GB/s bandwidth, while the CMP 40HX has 8 GB of GDDR6 on a 256-bit bus with 448.0 GB/s bandwidth. The V100 has more than double the memory bandwidth, which likely contributes to its benchmark dominance.

Compute resources strongly favor the V100: it has 5120 shading units, 320 TMUs, 128 ROPs, and 640 tensor cores, while the CMP 40HX has 2304 shading units, 144 TMUs, 64 ROPs, and 288 tensor cores. The CMP 40HX does have 36 RT cores, which the V100 lacks entirely. The V100's pixel rate is 176.6 GPixel/s and texture rate is 441.6 GTexel/s, versus 105.6 GPixel/s and 237.6 GTexel/s for the CMP 40HX. FP32 throughput is 14.13 TFLOPS for the V100 and 7.603 TFLOPS for the CMP 40HX, while FP16 is 28.26 TFLOPS and 15.21 TFLOPS, respectively.

Power and physical specs also differ. The V100 has a 300 W TDP with 2x 8-pin connectors and a suggested 700 W PSU, while the CMP 40HX has a 185 W TDP with 1x 8-pin and a 450 W suggested PSU. The CMP 40HX measures 229 mm in length, 111 mm in height, and 35 mm in width, while the V100's dimensions are not recorded. The V100 uses PCIe 3.0 x16, but the CMP 40HX uses PCIe 1.0 x4, a major interface difference. Both are dual-slot cards with no display outputs. The CMP 40HX supports DirectX 12 Ultimate (12_2), while the V100 supports DirectX 12 (12_1), though both support OpenGL 4.6 and Vulkan 1.4.

Architecture Differences

The architectural split is fundamental: Volta versus Turing. The V100 is built on the GV100 chip, which the database identifies as part of the Tesla Volta generation. The CMP 40HX is built on TU106, part of the Turing generation, and belongs to the "Mining GPUs" product line. The V100's transistor count of 21,100 million on an 815 mm² die gives it a density of 25.9M per mm², while the CMP 40HX's 10,800 million transistors on a 445 mm² die yield a density of 24.3M per mm². The V100 is a much larger and more complex chip, which aligns with its higher compute throughput.

Memory architecture is a major differentiator. The V100 uses HBM2 with a 4096-bit bus and 897.0 GB/s bandwidth, while the CMP 40HX uses GDDR6 with a 256-bit bus and 448.0 GB/s bandwidth. The V100's memory bandwidth is exactly double that of the CMP 40HX, which heavily influences compute-heavy workloads. The V100 also has 640 tensor cores versus 288 for the CMP 40HX, giving it more than twice the tensor core count. The CMP 40HX uniquely has 36 RT cores, which are absent from the V100, though this may be irrelevant for the benchmarks recorded.

The feature sets differ in API support: the CMP 40HX supports DirectX 12 Ultimate (12_2), while the V100 supports DirectX 12 (12_1). Both support OpenGL 4.6 and Vulkan 1.4. The CMP 40HX has a higher boost clock (1650 MHz versus 1380 MHz) and a higher base clock (1470 MHz versus 1245 MHz), but the V100 compensates with far more shading units, TMUs, ROPs, and tensor cores. The V100's pixel rate (176.6 GPixel/s) and texture rate (441.6 GTexel/s) are both higher than the CMP 40HX's (105.6 GPixel/s and 237.6 GTexel/s), reflecting its larger ROP and TMU counts. The V100 also has a higher TDP at 300 W versus 185 W, and requires a 700 W PSU versus 450 W.

The Verdict

The data is unambiguous: the NVIDIA Tesla V100 PCIe 16 GB is the superior performer in every recorded benchmark. It wins OpenCL by 74.6% and Vulkan by 45.2%, and its average score of 138063 places it in the 96th percentile, far above the CMP 40HX's 85637 in the 93rd percentile. The V100's nearest rivals are all within 1.7% of its score, meaning it belongs to a performance tier that the CMP 40HX does not approach. For users who need maximum compute throughput in OpenCL or Vulkan workloads, the V100 is the clear choice.

The CMP 40HX is not without merit. Its average score of 85637 beats the AMD Radeon PRO W6600 by 4.4% and the AMD Radeon Pro Vega 64X by 5.8%, and it sits within 2.1% of the NVIDIA Quadro GP100. It also has a lower TDP of 185 W versus 300 W and requires only a 450 W PSU, making it more power-efficient on paper. Its PCIe 1.0 x4 interface is a notable limitation, but for workloads that do not stress the interface, it may still be viable. However, the recorded benchmarks show no scenario where the CMP 40HX outperforms the V100.

The verdict depends on the use case. The V100 is the pick for compute-intensive tasks where raw score matters most: its 74.6% OpenCL lead and 45.2% Vulkan lead are decisive. The CMP 40HX is the pick for users who prioritize lower power draw and a smaller physical footprint, as it measures 229 mm by 111 mm by 35 mm, though the V100's dimensions are not recorded. But based strictly on benchmark performance, the V100 is the stronger card by a wide margin. The launch MSRP of the CMP 40HX was 699 USD, which may inform purchasing decisions, but the data shows that price buys a card in a lower performance tier.

DETAILED SPECIFICATIONS

SPECIFICATION
CMP 40HX
Tesla V100 PCIe 16 GB
Core Specs
Shading Units
2,304
5,120 +122.2%
Shaders
2,304
5,120 +122.2%
TMUs
144
320 +122.2%
ROPs
64
128 +100.0%
SM Count
36
80 +122.2%
Clocks
Base Clock
1470 MHz
1245 MHz
Boost Clock
1650 MHz
1380 MHz
Memory Clock
1750 MHz 14 Gbps effective
876 MHz 1752 Mbps effective
Memory
Memory Size
8 GB
16 GB
VRAM (MB)
8,192
16,384 +100.0%
Memory Type
GDDR6
HBM2
Memory Bus
256 bit
4096 bit
Bandwidth
448.0 GB/s
897.0 GB/s
Cache
L1 Cache
64 KB (per SM)
128 KB (per SM)
L2 Cache
4 MB
6 MB
Performance
Pixel Rate
105.6 GPixel/s
176.6 GPixel/s
Texture Rate
237.6 GTexel/s
441.6 GTexel/s
FP32 (TFLOPS)
7.603 TFLOPS
14.13 TFLOPS
FP64 (TFLOPS)
237.6 GFLOPS (1:32)
7.066 TFLOPS (1:2)
FP16 (TFLOPS)
15.21 TFLOPS (2:1)
28.26 TFLOPS (2:1)
AI/RT
RT Cores
36
—
Tensor Cores
288
640 +122.2%
Power
TDP
185 W
300 W
TDP (W)
185
300 +62.2%
Suggested PSU
450 W
700 W
Power Connectors
1x 8-pin
2x 8-pin
Architecture
Architecture
Turing
Volta
GPU Name
TU106
GV100
Generation
Mining GPUs
Tesla Volta (Vxx)
Process Size
12 nm
12 nm
Transistors
10,800 million
21,100 million
Die Size
445 mm²
815 mm²
Foundry
TSMC
TSMC
Density
24.3M / mm²
25.9M / mm²
API Support
DirectX
12 Ultimate (12_2)
12 (12_1)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
3.0
3.0
CUDA
7.5
7.0
Shader Model
6.8
6.8
Physical
Slot Width
Dual-slot
Dual-slot
Length
229 mm 9 inches
—
Height
111 mm 4.4 inches
—
Outputs
No outputs
No outputs
Bus Interface
PCIe 1.0 x4
PCIe 3.0 x16
Other
Launch Price
699 USD
—
Production
End-of-life
End-of-life
Predecessor
—
Tesla Pascal
Successor
—
Tesla Turing
View CMP 40HX Details View Tesla V100 PCIe 16 GB Details