NVIDIA CMP 90HX vs NVIDIA Tesla V100 PCIe 32 GB Comparison

NVIDIA
GEFORCE

NVIDIA CMP 90HX

CORE STATE GA102
VRAM 10 GB
CLOCK SPEED 1710 MHz
TDP 320 W
BUS WIDTH 320 bit
ARCHITECTURE Ampere
nm
PROCESS 8 nm
LAUNCH DATE 2021
VS
NVIDIA
GEFORCE

Tesla V100 PCIe 32 GB

CORE STATE GV100
VRAM 32 GB
CLOCK SPEED 1380 MHz
TDP 250 W
BUS WIDTH 4096 bit
ARCHITECTURE Volta
nm
PROCESS 12 nm
LAUNCH DATE 2018

PERFORMANCE BENCHMARKS

geekbench_opencl
69,000
168,763
geekbench_vulkan
N/A
131,847

Analysis: NVIDIA CMP 90HX vs NVIDIA Tesla V100 PCIe 32 GB

Head-to-Head Benchmarks

The database contains one direct benchmark comparison between the NVIDIA Tesla V100 PCIe 32 GB and the NVIDIA CMP 90HX, and the result is decisive. In the Geekbench OpenCL test, the Tesla V100 PCIe 32 GB scores 168,763 points, while the CMP 90HX scores 69,000 points. That is a delta of 144.6%, meaning the Tesla V100 PCIe 32 GB is nearly two and a half times faster in this particular compute workload. The benchmark results indicate a complete sweep: the Tesla V100 PCIe 32 GB wins the only head-to-head test, giving it 1 win against 0 for the CMP 90HX.

Looking at the broader benchmark landscape, the Tesla V100 PCIe 32 GB sits in the 96th percentile of all GPUs in the database, with an average benchmark score of 150,305 across all recorded tests. Its closest rival, the NVIDIA A10G, averages 151,963 points, a margin of only 1.1% in favor of the A10G. The AMD Radeon Pro W6800X is 6.5% ahead with 160,671 points, while the NVIDIA A100 PCIe 40 GB is 7.5% ahead with 162,504 points. The AMD Instinct MI100 trails the Tesla V100 PCIe 32 GB by 8.1%, scoring 139,035 points. These numbers show that the Tesla V100 PCIe 32 GB is competitively positioned among high-end accelerators, even if it is not the absolute leader in its class.

The CMP 90HX, by contrast, sits in the 90th percentile of all GPUs, with an average benchmark score of 69,000. Its nearest rivals are clustered tightly around that figure. The Intel Arc A770 scores 68,809, a delta of 0.3% in favor of the CMP 90HX. The AMD Radeon Instinct MI25 scores 68,562, which is 0.6% behind the CMP 90HX. The AMD Radeon Pro WX 8200 scores 69,870, putting it 1.2% ahead, and the NVIDIA Quadro P6000 scores 69,986, 1.4% ahead. The CMP 90HX is essentially at parity with these mid-range and professional cards, which is a far cry from the performance tier occupied by the Tesla V100 PCIe 32 GB.

The data shows a stark performance gap between these two products. The Tesla V100 PCIe 32 GB delivers more than double the OpenCL throughput of the CMP 90HX, and its average benchmark score across all tests is more than twice as high. The percentile rankings reinforce this: the Tesla V100 PCIe 32 GB is in the top 4% of all GPUs, while the CMP 90HX is in the top 10%. For any compute-heavy workload, the Tesla V100 PCIe 32 GB is the clear choice based on raw performance.

Where Each One Wins

The Tesla V100 PCIe 32 GB wins in every measurable compute benchmark recorded in the database. Its Geekbench OpenCL score of 168,763 is 144.6% higher than the CMP 90HX's 69,000. The Tesla V100 PCIe 32 GB also has a Geekbench Vulkan score of 131,847, a test for which the CMP 90HX has no recorded result. The average benchmark score of 150,305 for the Tesla V100 PCIe 32 GB versus 69,000 for the CMP 90HX means the Tesla V100 PCIe 32 GB is the superior performer in general compute tasks, including OpenCL and Vulkan workloads.

The CMP 90HX does not win any benchmark category in this comparison. However, its strengths lie elsewhere. The CMP 90HX has a higher FP32 throughput at 21.89 TFLOPS, compared to 14.13 TFLOPS for the Tesla V100 PCIe 32 GB. This suggests that in raw single-precision floating-point operations, the CMP 90HX has a theoretical advantage, even though the recorded OpenCL benchmark does not reflect this. The CMP 90HX also has a higher base clock of 1500 MHz and a higher boost clock of 1710 MHz, versus 1230 MHz and 1380 MHz for the Tesla V100 PCIe 32 GB. These clock speeds contribute to the CMP 90HX's higher FP32 rating, but the benchmark data shows that the Tesla V100 PCIe 32 GB still wins in real-world compute tests.

For use cases, the Tesla V100 PCIe 32 GB is the better choice for general-purpose GPU compute, machine learning inference, and any workload that leverages OpenCL or Vulkan. Its large 32 GB HBM2 memory with 897.0 GB/s of bandwidth provides ample capacity and throughput for large datasets. The CMP 90HX, with its 10 GB GDDR6X memory and 760.3 GB/s bandwidth, is more suited to tasks that do not require as much memory capacity. The CMP 90HX's higher FP32 rate may benefit certain single-precision compute workloads, but the benchmark evidence shows that the Tesla V100 PCIe 32 GB delivers better overall performance in practice.

FAQ

Q: Which GPU has the higher average benchmark score?

A: The NVIDIA Tesla V100 PCIe 32 GB has an average benchmark score of 150,305, while the NVIDIA CMP 90HX has an average benchmark score of 69,000.

Q: How much faster is the Tesla V100 PCIe 32 GB in the Geekbench OpenCL test?

A: The Tesla V100 PCIe 32 GB scores 168,763 points, which is 144.6% higher than the CMP 90HX's 69,000 points.

Q: Does the CMP 90HX win any benchmark in the head-to-head comparison?

A: No, the CMP 90HX wins zero benchmarks. The Tesla V100 PCIe 32 GB wins the only recorded head-to-head test, giving it 1 win to 0.

Q: What is the FP32 performance difference between the two GPUs?

A: The CMP 90HX has a higher FP32 rating at 21.89 TFLOPS, while the Tesla V100 PCIe 32 GB has 14.13 TFLOPS. The CMP 90HX leads by 7.76 TFLOPS in theoretical single-precision compute.

Q: How do the two GPUs compare in memory bandwidth?

A: The Tesla V100 PCIe 32 GB has a memory bandwidth of 897.0 GB/s, while the CMP 90HX has a memory bandwidth of 760.3 GB/s. The Tesla V100 PCIe 32 GB is ahead by 136.7 GB/s.

Q: What are the percentile rankings for each GPU?

A: The Tesla V100 PCIe 32 GB is in the 96th percentile of all GPUs, while the CMP 90HX is in the 90th percentile.

Specification Differences

The two GPUs differ significantly in almost every specification category. The Tesla V100 PCIe 32 GB has a die size of 815 mm², while the CMP 90HX has a die size of 628 mm². The Tesla V100 PCIe 32 GB packs 21,100 million transistors, while the CMP 90HX has 28,300 million. Transistor density also differs: the Tesla V100 PCIe 32 GB has 25.9 million transistors per mm², while the CMP 90HX has 45.1 million per mm².

Memory configurations are starkly different. The Tesla V100 PCIe 32 GB has 32 GB of HBM2 memory on a 4096-bit bus, yielding 897.0 GB/s of bandwidth. The CMP 90HX has 10 GB of GDDR6X memory on a 320-bit bus, yielding 760.3 GB/s of bandwidth. The memory clock also differs: the Tesla V100 PCIe 32 GB runs at 876 MHz (1752 Mbps effective), while the CMP 90HX runs at 1188 MHz (19 Gbps effective).

The compute units are different as well. The Tesla V100 PCIe 32 GB has 5120 shading units, 320 texture mapping units, and 128 raster operation units. The CMP 90HX has 6400 shading units, 200 texture mapping units, and 80 raster operation units. The Tesla V100 PCIe 32 GB has 640 tensor cores, while the CMP 90HX has 200 tensor cores and 50 ray tracing cores. The Tesla V100 PCIe 32 GB has no ray tracing cores.

Pixel and texture rates differ accordingly. The Tesla V100 PCIe 32 GB has a pixel rate of 176.6 GPixel/s and a texture rate of 441.6 GTexel/s. The CMP 90HX has a pixel rate of 136.8 GPixel/s and a texture rate of 342.0 GTexel/s. The FP16 performance also differs: the Tesla V100 PCIe 32 GB delivers 28.26 TFLOPS with a 2:1 ratio, while the CMP 90HX delivers 21.89 TFLOPS with a 1:1 ratio.

The bus interface is another point of divergence. The Tesla V100 PCIe 32 GB uses PCIe 3.0 x16, while the CMP 90HX uses PCIe 1.0 x4. The CMP 90HX has specified dimensions of 285 mm in length and 112 mm in height, while the Tesla V100 PCIe 32 GB has no recorded dimensions. The CMP 90HX has a TDP of 320 W and a suggested PSU of 700 W, while the Tesla V100 PCIe 32 GB has a TDP of 250 W and a suggested PSU of 600 W. Both use dual-slot cooling and 2x 8-pin power connectors. The Tesla V100 PCIe 32 GB supports DirectX 12 (12_1), while the CMP 90HX supports DirectX 12 Ultimate (12_2). Both support OpenGL 4.6 and Vulkan 1.4.

Architecture Differences

The architectural divide between these two GPUs is fundamental. The Tesla V100 PCIe 32 GB is built on the Volta architecture, using the GV100 chip, while the CMP 90HX is built on the Ampere architecture, using the GA102 chip. The Tesla V100 PCIe 32 GB belongs to the Tesla Volta generation, while the CMP 90HX belongs to the Mining GPUs generation.

The manufacturing process differs as well. The Tesla V100 PCIe 32 GB is fabricated on a 12 nm process at TSMC, while the CMP 90HX is fabricated on an 8 nm process at Samsung. This process difference explains the transistor density gap: the CMP 90HX packs 45.1 million transistors per mm², compared to 25.9 million per mm² for the Tesla V100 PCIe 32 GB, despite having a smaller die.

The memory architecture reflects different design goals. The Tesla V100 PCIe 32 GB uses HBM2 memory, which offers a very wide 4096-bit bus and high bandwidth per watt. The CMP 90HX uses GDDR6X memory, which is a more conventional design with a narrower 320-bit bus. The Tesla V100 PCIe 32 GB has 32 GB of memory, three times the capacity of the CMP 90HX's 10 GB.

The compute feature sets differ notably. The Tesla V100 PCIe 32 GB has 640 tensor cores, which are Volta-generation tensor cores designed for deep learning workloads. The CMP 90HX has 200 tensor cores and 50 ray tracing cores, reflecting the Ampere architecture's addition of ray tracing hardware. The Tesla V100 PCIe 32 GB has no ray tracing cores, as ray tracing was not part of the Volta design.

The FP16 processing differs in ratio: the Tesla V100 PCIe 32 GB delivers FP16 at a 2:1 ratio relative to FP32, meaning it can process two FP16 operations per FP32 operation. The CMP 90HX delivers FP16 at a 1:1 ratio, meaning it processes FP16 and FP32 at the same rate. This gives the Tesla V100 PCIe 32 GB a higher FP16 throughput of 28.26 TFLOPS versus 21.89 TFLOPS for the CMP 90HX.

The API support also reflects the generational gap. The Tesla V100 PCIe 32 GB supports DirectX 12 (12_1), while the CMP 90HX supports DirectX 12 Ultimate (12_2), which includes features like ray tracing and mesh shaders. Both support OpenGL 4.6 and Vulkan 1.4, but the CMP 90HX's DirectX 12 Ultimate support is a newer specification.

The bus interface is a notable divergence: the Tesla V100 PCIe 32 GB uses PCIe 3.0 x16, while the CMP 90HX uses PCIe 1.0 x4. This means the CMP 90HX has a much narrower and slower host interface, which could bottleneck data transfer in some workloads. The Tesla V100 PCIe 32 GB also has a lower TDP of 250 W versus 320 W for the CMP 90HX, despite delivering higher benchmark scores. The CMP 90HX has a physical length of 285 mm and height of 112 mm, while the Tesla V100 PCIe 32 GB has no recorded dimensions. Both are dual-slot cards with no display outputs, and both use 2x 8-pin power connectors. The Tesla V100 PCIe 32 GB was released on March 26, 2018, while the CMP 90HX was released on July 27, 2021. The Tesla V100 PCIe 32 GB has a predecessor in Tesla Pascal and a successor in Tesla Turing, while the CMP 90HX has no recorded predecessor or successor.

DETAILED SPECIFICATIONS

SPECIFICATION
CMP 90HX
Tesla V100 PCIe 32 GB
Core Specs
Shading Units
6,400
5,120 -20.0%
Shaders
6,400
5,120 -20.0%
TMUs
200
320 +60.0%
ROPs
80
128 +60.0%
SM Count
50
80 +60.0%
Clocks
Base Clock
1500 MHz
1230 MHz
Boost Clock
1710 MHz
1380 MHz
Memory Clock
1188 MHz 19 Gbps effective
876 MHz 1752 Mbps effective
Memory
Memory Size
10 GB
32 GB
VRAM (MB)
10,240
32,768 +220.0%
Memory Type
GDDR6X
HBM2
Memory Bus
320 bit
4096 bit
Bandwidth
760.3 GB/s
897.0 GB/s
Cache
L1 Cache
128 KB (per SM)
128 KB (per SM)
L2 Cache
5 MB
6 MB
Performance
Pixel Rate
136.8 GPixel/s
176.6 GPixel/s
Texture Rate
342.0 GTexel/s
441.6 GTexel/s
FP32 (TFLOPS)
21.89 TFLOPS
14.13 TFLOPS
FP64 (TFLOPS)
342.0 GFLOPS (1:64)
7.066 TFLOPS (1:2)
FP16 (TFLOPS)
21.89 TFLOPS (1:1)
28.26 TFLOPS (2:1)
AI/RT
RT Cores
50
—
Tensor Cores
200
640 +220.0%
Power
TDP
320 W
250 W
TDP (W)
320
250 -21.9%
Suggested PSU
700 W
600 W
Power Connectors
2x 8-pin
2x 8-pin
Architecture
Architecture
Ampere
Volta
GPU Name
GA102
GV100
Generation
Mining GPUs
Tesla Volta (Vxx)
Process Size
8 nm
12 nm
Transistors
28,300 million
21,100 million
Die Size
628 mm²
815 mm²
Foundry
Samsung
TSMC
Density
45.1M / mm²
25.9M / mm²
API Support
DirectX
12 Ultimate (12_2)
12 (12_1)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
3.0
3.0
CUDA
8.6
7.0
Shader Model
6.8
6.8
Physical
Slot Width
Dual-slot
Dual-slot
Length
285 mm 11.2 inches
—
Height
112 mm 4.4 inches
—
Outputs
No outputs
No outputs
Bus Interface
PCIe 1.0 x4
PCIe 3.0 x16
Other
Production
End-of-life
End-of-life
Predecessor
—
Tesla Pascal
Successor
—
Tesla Turing
View CMP 90HX Details View Tesla V100 PCIe 32 GB Details