NVIDIA CMP 90HX vs NVIDIA P102-100 Comparison
NVIDIA CMP 90HX
P102-100
PERFORMANCE BENCHMARKS
Analysis: NVIDIA CMP 90HX vs NVIDIA P102-100
The NVIDIA CMP 90HX and NVIDIA P102-100 are both mining-focused graphics cards with no display outputs, but they represent two very different generations of NVIDIA architecture. The data shows a clear performance hierarchy between them, though the older card retains some relevance in specific workloads. This analysis breaks down where each card wins, the architectural chasm between them, and what the benchmark results actually mean for a builder deciding between these two end-of-life mining parts.
Where Each One Wins
The benchmark data is unambiguous: the CMP 90HX wins the only head-to-head test available. In Geekbench OpenCL, the CMP 90HX scores 69000 against the P102-100’s 49602, a 39.1% advantage. That is a decisive margin, not a narrow one. The CMP 90HX’s average benchmark score of 69000 places it at the 90th percentile of all GPUs, while the P102-100’s average of 58528 puts it at the 88th percentile. So in raw compute throughput, the newer card is not just faster—it is in a different performance tier.
However, the P102-100 has one unique advantage: it is the only card in this comparison with a Vulkan benchmark result. Its Geekbench Vulkan score of 67454 is remarkably close to the CMP 90HX’s OpenCL score of 69000. This suggests that in Vulkan-based workloads, the older Pascal card can approach the newer Ampere card’s performance, even if they never go head-to-head in the same API test. The data does not include a Vulkan score for the CMP 90HX, so this remains an area where the P102-100 shows surprising strength relative to its overall compute deficit.
For the CMP 90HX, the wins extend to every architectural metric that matters for compute density. It has double the shading units (6400 vs 3200), double the FP32 throughput (21.89 TFLOPS vs 10.77 TFLOPS), and double the memory capacity (10 GB vs 5 GB). The 90HX also brings ray tracing cores and tensor cores to the table, features entirely absent from the P102-100. If your workload can use those, the 90HX is the only choice.
Architecture Differences
The CMP 90HX is built on the GA102 chip using the Ampere architecture, fabricated on Samsung’s 8 nm process. It packs 28,300 million transistors into a 628 mm² die, giving a transistor density of 45.1M per mm². The P102-100 uses the GP102 chip with the Pascal architecture, built on TSMC’s 16 nm process. That older node holds 11,800 million transistors across a 471 mm² die, with a density of just 25.1M per mm². The density difference is stark: the Ampere chip crams nearly twice the transistors into only a third more die area.
Memory configurations diverge sharply. The CMP 90HX uses 10 GB of GDDR6X on a 320-bit bus, delivering 760.3 GB/s of bandwidth at 19 Gbps effective. The P102-100 has 5 GB of GDDR5X, also on a 320-bit bus, but bandwidth drops to 440.3 GB/s at 11 Gbps effective. That is a 72.7% bandwidth advantage for the newer card, which matters enormously for memory-bound mining workloads.
Clock speeds tell a different story. The P102-100 has a higher base clock at 1582 MHz versus 1500 MHz, though the boost clocks are nearly identical (1683 MHz vs 1710 MHz). The older card’s memory clock of 1376 MHz is also higher than the CMP 90HX’s 1188 MHz, but the GDDR6X architecture on the newer card more than compensates with the higher effective data rate.
Feature sets are generations apart. The CMP 90HX supports DirectX 12 Ultimate (12_2) and includes 50 RT cores and 200 tensor cores. The P102-100 is limited to DirectX 12 (12_1) and has neither RT nor tensor cores. Both support OpenGL 4.6 and Vulkan 1.4, and both use a PCIe 1.0 x4 bus interface—a strange limitation for mining cards that suggests neither was designed for general compute workloads that rely on host transfer rates.
Head-to-Head Benchmarks
The single head-to-head benchmark is Geekbench OpenCL, and the result is decisive. The CMP 90HX scores 69000, the P102-100 scores 49602, giving the newer card a 39.1% win. To put that in perspective, the CMP 90HX’s nearest rival is the Intel Arc A770 at 68809 (a 0.3% gap), while the P102-100 sits near the AMD Radeon PRO V710 at 58657 (a 0.2% gap in the P102-100’s favor). The two cards occupy different neighborhoods in the performance landscape entirely.
The FP32 compute figures corroborate the benchmark gap. The CMP 90HX delivers 21.89 TFLOPS, exactly double the P102-100’s 10.77 TFLOPS. The pixel rate is nearly identical (136.8 GPixel/s vs 134.6 GPixel/s), and the texture rates are close (342.0 GTexel/s vs 336.6 GTexel/s). This tells you the fixed-function units scaled modestly, but the shader horsepower doubled. That is why OpenCL—a compute API—shows such a large delta.
FP16 performance is where the architectures truly diverge. The CMP 90HX offers 21.89 TFLOPS FP16 with a 1:1 ratio to FP32. The P102-100 offers only 168.3 GFLOPS FP16, a 1:64 ratio. That means the Pascal card is effectively crippled for FP16 workloads, while the Ampere card treats them as first-class citizens. Any mining or compute task using FP16 will see an even larger gap than the OpenCL score suggests.
FAQ
Q: Which card has better overall compute performance?
A: The CMP 90HX wins decisively. Its Geekbench OpenCL score of 69000 beats the P102-100’s 49602 by 39.1%. Its FP32 throughput of 21.89 TFLOPS is exactly double the P102-100’s 10.77 TFLOPS.
Q: Does the P102-100 have any performance advantage?
A: The data shows one: it has a Geekbench Vulkan score of 67454, which is close to the CMP 90HX’s OpenCL score of 69000. The CMP 90HX has no Vulkan benchmark listed, so in Vulkan-specific workloads the Pascal card may be competitive, though this is not confirmed by a direct head-to-head test.
Q: Are these cards usable for gaming?
A: No. Both cards have no display outputs, making them unusable for any task requiring a monitor connection. They are mining-only parts, as indicated by their generation label "Mining GPUs."
Q: Which card has more memory and bandwidth?
A: The CMP 90HX has 10 GB of GDDR6X with 760.3 GB/s bandwidth. The P102-100 has 5 GB of GDDR5X with 440.3 GB/s bandwidth. The newer card has double the capacity and 72.7% more bandwidth.
Q: What are the power requirements for each card?
A: The CMP 90HX has a TDP of 320 W and requires a 700 W power supply. The P102-100 has a TDP of 250 W and requires a 600 W power supply. Both use two 8-pin power connectors and are dual-slot cards.
Q: Which card supports ray tracing and tensor cores?
A: Only the CMP 90HX. It has 50 RT cores and 200 tensor cores, enabling DirectX 12 Ultimate features. The P102-100 has neither, limited to DirectX 12 (12_1) support.
The Verdict
The choice here is straightforward for most use cases. The CMP 90HX is the superior mining and compute card by every measurable metric in the data. It wins the only head-to-head benchmark by 39.1%, has double the FP32 throughput, double the memory capacity, 72.7% more bandwidth, and adds RT and tensor cores that the P102-100 lacks entirely. Its 90th percentile ranking versus the P102-100’s 88th confirms the performance tier gap.
The P102-100 is only worth considering if your workload is Vulkan-specific, because its Vulkan score of 67454 is close to the 90HX’s OpenCL score. It also runs cooler and cheaper—250 W TDP versus 320 W, and a 600 W PSU recommendation versus 700 W. But those are secondary concerns when the compute performance gap is so large.
For a builder setting up a mining rig or a compute node, the CMP 90HX is the pick. It is newer, faster, and far more capable. The P102-100 is a legacy part that only makes sense if you already own it and are targeting Vulkan workloads, or if power draw is the absolute priority. Otherwise, the data says the 90HX is worth the extra wattage.
Specification Differences
| Specification | NVIDIA CMP 90HX | NVIDIA P102-100 |
|---|---|---|
| Chip | GA102 | GP102 |
| Architecture | Ampere | Pascal |
| Process Node | 8 nm (Samsung) | 16 nm (TSMC) |
| Transistors | 28,300 million | 11,800 million |
| Die Size | 628 mm² | 471 mm² |
| Transistor Density | 45.1M / mm² | 25.1M / mm² |
| Base Clock | 1500 MHz | 1582 MHz |
| Boost Clock | 1710 MHz | 1683 MHz |
| Memory Clock | 1188 MHz (19 Gbps effective) | 1376 MHz (11 Gbps effective) |
| Memory Size | 10 GB | 5 GB |
| Memory Type | GDDR6X | GDDR5X |
| Memory Bus | 320 bit | 320 bit |
| Bandwidth | 760.3 GB/s | 440.3 GB/s |
| Shading Units | 6400 | 3200 |
| TMUs | 200 | 200 |
| ROPs | 80 | 80 |
| RT Cores | 50 | None |
| Tensor Cores | 200 | None |
| Pixel Rate | 136.8 GPixel/s | 134.6 GPixel/s |
| Texture Rate | 342.0 GTexel/s | 336.6 GTexel/s |
| FP32 | 21.89 TFLOPS | 10.77 TFLOPS |
| FP16 | 21.89 TFLOPS (1:1) | 168.3 GFLOPS (1:64) |
| TDP | 320 W | 250 W |
| Suggested PSU | 700 W | 600 W |
| DirectX | 12 Ultimate (12_2) | 12 (12_1) |
| Length | 285 mm (11.2 inches) | 267 mm (10.5 inches) |
| Release Date | 2021-07-27 | 2018-02-11 |