NVIDIA CMP 90HX vs NVIDIA Tesla P100 PCIe 16 GB Comparison

NVIDIA
GEFORCE

NVIDIA CMP 90HX

CORE STATE GA102
VRAM 10 GB
CLOCK SPEED 1710 MHz
TDP 320 W
BUS WIDTH 320 bit
ARCHITECTURE Ampere
nm
PROCESS 8 nm
LAUNCH DATE 2021
VS
NVIDIA
GEFORCE

Tesla P100 PCIe 16 GB

CORE STATE GP100
VRAM 16 GB
CLOCK SPEED 1329 MHz
TDP 250 W
BUS WIDTH 4096 bit
ARCHITECTURE Pascal
nm
PROCESS 16 nm
LAUNCH DATE 2016

PERFORMANCE BENCHMARKS

geekbench_opencl
69,000
79,605

Analysis: NVIDIA CMP 90HX vs NVIDIA Tesla P100 PCIe 16 GB

NVIDIA’s Tesla P100 PCIe 16 GB and the CMP 90HX are both end-of-life, no-display accelerators, but they target entirely different eras of compute. The Pascal-based Tesla card is a 2016 compute workhorse, while the Ampere-based CMP 90HX is a 2021 mining-specific part. The benchmark data shows a clear overall winner, but the architectural split tells a more nuanced story about where each part excels.

Head-to-Head Benchmarks

The only direct benchmark comparison available is Geekbench OpenCL, and it decisively favors the older Tesla P100. The P100 scores 79,605 points, while the CMP 90HX trails at 69,000 points. That is a 15.4% advantage for the Tesla card. This is a significant margin, placing the P100 well ahead of its rival in general-purpose compute workloads.

The P100’s score is not just better than the CMP 90HX; it also holds its own against its own nearest rivals. The Tesla card sits 0.3% ahead of the Tesla P100 PCIe 12 GB, which scores 79,396. It is also 0.8% above the AMD Radeon RX 6850M XT at 78,940. Interestingly, the P100 trails the NVIDIA GeForce RTX 5090 by just 0.3% (79,842) and the AMD Radeon Pro Vega 64X by 1.7% (80,959). The data suggests that despite its age, the P100 remains competitive with much newer hardware in raw OpenCL throughput.

On the other side, the CMP 90HX’s score of 69,000 places it in a lower tier. It sits just 0.3% ahead of the Intel Arc A770 at 68,809 and 0.6% ahead of the AMD Radeon Instinct MI25 at 68,562. It falls behind the AMD Radeon Pro WX 8200 by 1.2% (69,870) and the NVIDIA Quadro P6000 by 1.4% (69,986). The CMP 90HX is not a slow card by any means, but it is clearly outclassed by the P100 in this specific test.

The delta of 15.4% is the headline number here. It shows that in a pure compute benchmark, the P100’s architecture and memory subsystem deliver substantially more performance per clock or per core than the CMP 90HX’s design. The win tally is 1-0 in favor of the Tesla P100, and there is no test in the data where the CMP 90HX takes the lead.

Where Each One Wins

The Tesla P100 wins the only head-to-head benchmark, so it is the clear choice for generic compute workloads that resemble Geekbench OpenCL. This test typically exercises a mix of memory bandwidth, shader throughput, and arithmetic operations. The P100’s 16 GB of HBM2 memory on a 4096-bit bus delivers 732.2 GB/s of bandwidth, which is a massive advantage for memory-bound tasks. Its FP32 throughput of 9.526 TFLOPS is lower than the CMP 90HX’s 21.89 TFLOPS, yet it still wins the benchmark. This suggests the P100’s advantage lies in memory latency and bandwidth efficiency rather than raw FLOPs.

The CMP 90HX, despite losing the benchmark, has specific strengths that could make it preferable in certain scenarios. Its FP32 throughput of 21.89 TFLOPS is more than double the P100’s 9.526 TFLOPS. For workloads that are purely compute-bound and fit within its 10 GB GDDR6X memory pool, the CMP 90HX could theoretically process more data per second. It also has 6,400 shading units compared to the P100’s 3,584, giving it a substantial raw execution resource advantage. However, the benchmark data does not include any test where this translates into a win.

The CMP 90HX also has dedicated RT cores (50) and tensor cores (200), which the P100 lacks entirely. This makes the CMP 90HX architecturally capable of ray tracing and AI inference tasks that the P100 cannot accelerate at all. In those specific workloads, the CMP 90HX would win by default, as the P100 has no equivalent hardware. The P100’s FP16 performance is 19.05 TFLOPS at a 2:1 ratio, while the CMP 90HX achieves 21.89 TFLOPS at a 1:1 ratio, meaning the CMP 90HX does not sacrifice half-rate throughput for FP16.

For practical use, the P100 is the better choice for scientific computing, data center workloads, and any OpenCL-based application. The CMP 90HX, with its mining-specific design and lack of display outputs, is more suited for tasks that leverage its tensor cores or require massive FP32 throughput. Yet, the lack of a benchmark win for the CMP 90HX means its advantages remain theoretical in this comparison.

Architecture Differences

The two cards are built on fundamentally different architectures and process nodes. The Tesla P100 uses the GP100 chip based on NVIDIA’s Pascal architecture, manufactured on TSMC’s 16 nm process. The CMP 90HX uses the GA102 chip based on the Ampere architecture, built on Samsung’s 8 nm process. This is a generational leap of two architecture families, and the process node shrinks from 16 nm to 8 nm.

The transistor counts reflect the process differences. The P100 packs 15,300 million transistors on a 610 mm² die, yielding a transistor density of 25.1 million per square millimeter. The CMP 90HX has 28,300 million transistors on a slightly larger 628 mm² die, achieving a density of 45.1 million per square millimeter. The CMP 90HX packs nearly twice the transistors into roughly the same physical space, a direct result of the more advanced 8 nm node.

Memory architecture is another stark divider. The P100 uses 16 GB of HBM2 with a 4096-bit bus, delivering 732.2 GB/s of bandwidth. The CMP 90HX uses 10 GB of GDDR6X on a 320-bit bus, achieving 760.3 GB/s. The CMP 90HX has slightly higher peak bandwidth, but the P100’s HBM2 memory offers different latency characteristics. The memory clock rates are also different: the P100 runs at 715 MHz (1430 Mbps effective), while the CMP 90HX runs at 1188 MHz (19 Gbps effective).

The compute units are configured very differently. The P100 has 3,584 shading units, 224 texture mapping units, and 96 ROPs. The CMP 90HX has 6,400 shading units, 200 TMUs, and 80 ROPs. The CMP 90HX has far more shaders but fewer TMUs and ROPs, indicating a design optimized for compute rather than rasterization. The CMP 90HX also includes 50 RT cores and 200 tensor cores, features absent from the Pascal-based P100.

Clock speeds also favor the CMP 90HX. Its base clock is 1500 MHz with a boost of 1710 MHz, compared to the P100’s 1190 MHz base and 1329 MHz boost. This higher clock speed contributes to the CMP 90HX’s higher FP32 throughput of 21.89 TFLOPS versus 9.526 TFLOPS for the P100. Pixel and texture rates follow the same pattern: the CMP 90HX achieves 136.8 GPixel/s and 342.0 GTexel/s, while the P100 achieves 127.6 GPixel/s and 297.7 GTexel/s.

Power and interface differences are significant. The P100 has a TDP of 250 W with a single 8-pin connector and a suggested PSU of 600 W, using a PCIe 3.0 x16 interface. The CMP 90HX has a TDP of 320 W, requires two 8-pin connectors, and suggests a 700 W PSU. Critically, the CMP 90HX uses a PCIe 1.0 x4 interface, which is a severe bottleneck for data transfer compared to the P100’s PCIe 3.0 x16.

FAQ

Q: Which card has higher raw FP32 compute performance?

A: The NVIDIA CMP 90HX has a significantly higher FP32 throughput of 21.89 TFLOPS, compared to the Tesla P100’s 9.526 TFLOPS.

Q: Why does the Tesla P100 win the Geekbench OpenCL test despite lower FP32 performance?

A: The P100 scores 79,605 versus the CMP 90HX’s 69,000, a 15.4% margin. This suggests that memory bandwidth, latency, and architectural efficiency in the P100’s HBM2 setup outweigh the CMP 90HX’s raw compute advantage in this workload.

Q: Does the CMP 90HX have any hardware features the P100 lacks?

A: Yes, the CMP 90HX includes 50 RT cores and 200 tensor cores, which are not present on the P100. These enable ray tracing and AI acceleration workloads on the CMP 90HX.

Q: What are the memory configurations of each card?

A: The Tesla P100 has 16 GB of HBM2 on a 4096-bit bus with 732.2 GB/s bandwidth. The CMP 90HX has 10 GB of GDDR6X on a 320-bit bus with 760.3 GB/s bandwidth.

Q: Which card has a more recent release date?

A: The CMP 90HX was released on 2021-07-27, while the Tesla P100 was released on 2016-06-19.

Q: How do the cards compare in terms of power requirements?

A: The P100 has a 250 W TDP with one 8-pin connector and a 600 W suggested PSU. The CMP 90HX has a 320 W TDP, two 8-pin connectors, and a 700 W suggested PSU.

Specification Differences

| Specification | NVIDIA Tesla P100 PCIe 16 GB | NVIDIA CMP 90HX |

|---|---|---|

| Architecture | Pascal | Ampere |

| Process Node | 16 nm | 8 nm |

| Foundry | TSMC | Samsung |

| Transistors | 15,300 million | 28,300 million |

| Die Size | 610 mm² | 628 mm² |

| Transistor Density | 25.1M / mm² | 45.1M / mm² |

| Base Clock | 1190 MHz | 1500 MHz |

| Boost Clock | 1329 MHz | 1710 MHz |

| Memory Size | 16 GB | 10 GB |

| Memory Type | HBM2 | GDDR6X |

| Memory Bus | 4096 bit | 320 bit |

| Memory Bandwidth | 732.2 GB/s | 760.3 GB/s |

| Shading Units | 3584 | 6400 |

| TMUs | 224 | 200 |

| ROPs | 96 | 80 |

| RT Cores | 0 | 50 |

| Tensor Cores | 0 | 200 |

| Pixel Rate | 127.6 GPixel/s | 136.8 GPixel/s |

| Texture Rate | 297.7 GTexel/s | 342.0 GTexel/s |

| FP32 Performance | 9.526 TFLOPS | 21.89 TFLOPS |

| FP16 Performance | 19.05 TFLOPS (2:1) | 21.89 TFLOPS (1:1) |

| TDP | 250 W | 320 W |

| Power Connectors | 1x 8-pin | 2x 8-pin |

| Suggested PSU | 600 W | 700 W |

| Bus Interface | PCIe 3.0 x16 | PCIe 1.0 x4 |

| Length | 267 mm | 285 mm |

| Height | Not specified | 112 mm |

| Launch MSRP | 5,699 USD | Not specified |

| Release Date | 2016-06-19 | 2021-07-27 |

| DirectX Support | 12 (12_1) | 12 Ultimate (12_2) |

| Vulkan Support | 1.3 | 1.4 |

DETAILED SPECIFICATIONS

SPECIFICATION
CMP 90HX
Tesla P100 PCIe 16 GB
Core Specs
Shading Units
6,400
3,584 -44.0%
Shaders
6,400
3,584 -44.0%
TMUs
200
224 +12.0%
ROPs
80
96 +20.0%
SM Count
50
56 +12.0%
Clocks
Base Clock
1500 MHz
1190 MHz
Boost Clock
1710 MHz
1329 MHz
Memory Clock
1188 MHz 19 Gbps effective
715 MHz 1430 Mbps effective
Memory
Memory Size
10 GB
16 GB
VRAM (MB)
10,240
16,384 +60.0%
Memory Type
GDDR6X
HBM2
Memory Bus
320 bit
4096 bit
Bandwidth
760.3 GB/s
732.2 GB/s
Cache
L1 Cache
128 KB (per SM)
24 KB (per SM)
L2 Cache
5 MB
4 MB
Performance
Pixel Rate
136.8 GPixel/s
127.6 GPixel/s
Texture Rate
342.0 GTexel/s
297.7 GTexel/s
FP32 (TFLOPS)
21.89 TFLOPS
9.526 TFLOPS
FP64 (TFLOPS)
342.0 GFLOPS (1:64)
4.763 TFLOPS (1:2)
FP16 (TFLOPS)
21.89 TFLOPS (1:1)
19.05 TFLOPS (2:1)
AI/RT
RT Cores
50
—
Tensor Cores
200
—
Power
TDP
320 W
250 W
TDP (W)
320
250 -21.9%
Suggested PSU
700 W
600 W
Power Connectors
2x 8-pin
1x 8-pin
Architecture
Architecture
Ampere
Pascal
GPU Name
GA102
GP100
Generation
Mining GPUs
Tesla Pascal (Pxx)
Process Size
8 nm
16 nm
Transistors
28,300 million
15,300 million
Die Size
628 mm²
610 mm²
Foundry
Samsung
TSMC
Density
45.1M / mm²
25.1M / mm²
API Support
DirectX
12 Ultimate (12_2)
12 (12_1)
OpenGL
4.6
4.6
Vulkan
1.4
1.3
OpenCL
3.0
3.0
CUDA
8.6
6.0
Shader Model
6.8
6.0
Physical
Slot Width
Dual-slot
Dual-slot
Length
285 mm 11.2 inches
267 mm 10.5 inches
Height
112 mm 4.4 inches
—
Outputs
No outputs
No outputs
Bus Interface
PCIe 1.0 x4
PCIe 3.0 x16
Other
Launch Price
—
5,699 USD
Production
End-of-life
End-of-life
Predecessor
—
Tesla Maxwell
Successor
—
Tesla Volta
View CMP 90HX Details View Tesla P100 PCIe 16 GB Details