NVIDIA CMP 90HX vs NVIDIA Quadro GP100 Comparison
NVIDIA CMP 90HX
Quadro GP100
PERFORMANCE BENCHMARKS
Analysis: NVIDIA CMP 90HX vs NVIDIA Quadro GP100
The Geekbench OpenCL data places the NVIDIA Quadro GP100 clearly ahead of the NVIDIA CMP 90HX, with a score of 87445 against 69000. This is a 26.7% advantage for the Quadro GP100, a decisive margin in a single benchmark. While the CMP 90HX is not without merit, the data shows that in raw compute throughput as measured by this test, the older Pascal-based professional card outpaces the newer Ampere-based mining card significantly.
Head-to-Head Benchmarks
The only direct benchmark comparison available is Geekbench OpenCL, and the results are unambiguous. The Quadro GP100 scores 87445, while the CMP 90HX scores 69000. This yields a delta of 26.7% in favor of the Quadro GP100, meaning the GP100 completes the workload roughly a quarter faster than the CMP 90HX. This is not a marginal win; it is a substantial performance gap that defines their relative standing.
Looking at the broader competitive landscape reinforces this. The Quadro GP100 sits at the 93rd percentile of all GPUs, while the CMP 90HX sits at the 90th. The GP100’s nearest rivals include the AMD Radeon PRO W7600, which scores 87108 (a 0.4% deficit), and the NVIDIA CMP 40HX, which scores 85637 (a 2.1% deficit). It also outperforms the NVIDIA RTX A4500 Mobile (91134, -4%) and the NVIDIA RTX A4500 (91671, -4.6%) by smaller margins, though those are technically ahead. The CMP 90HX, by contrast, trades blows with the Intel Arc A770 (68809, 0.3% ahead) and the AMD Radeon Instinct MI25 (68562, 0.6% ahead), while trailing the AMD Radeon Pro WX 8200 (69870, -1.2%) and NVIDIA Quadro P6000 (69986, -1.4%).
In direct terms, the Quadro GP100 is more than 18% faster than the CMP 90HX’s closest competitor, the Quadro P6000. The CMP 90HX’s score places it in a cluster where small percentage swings decide positioning, whereas the GP100 enjoys a clear buffer above its nearest rivals. For any workload reliant on OpenCL compute, the data heavily favors the Quadro GP100.
Architecture Differences
The two cards come from different architectural generations and are built for different purposes. The Quadro GP100 uses the GP100 chip on the Pascal architecture, manufactured by TSMC on a 16 nm process. It packs 15,300 million transistors on a 610 mm² die, yielding a transistor density of 25.1 million per mm². In contrast, the CMP 90HX uses the GA102 chip on the Ampere architecture, built by Samsung on an 8 nm process. It contains 28,300 million transistors on a 628 mm² die, giving a density of 45.1 million per mm².
The memory subsystems diverge sharply. The Quadro GP100 has 16 GB of HBM2 on a 4096-bit bus, delivering 732.2 GB/s of bandwidth. The CMP 90HX has 10 GB of GDDR6X on a 320-bit bus, which achieves a slightly higher 760.3 GB/s. Despite the smaller bus, the faster GDDR6X memory (19 Gbps effective vs 1430 Mbps effective) gives the CMP 90HX a modest bandwidth edge.
Compute resources also differ. The Quadro GP100 has 3584 shading units, 224 TMUs, and 96 ROPs. The CMP 90HX has 6400 shading units, 200 TMUs, and 80 ROPs. The CMP 90HX also includes 50 RT cores and 200 tensor cores, features entirely absent from the Pascal-based GP100. Clock speeds favor the CMP 90HX: its base is 1500 MHz and boost is 1710 MHz, versus 1304 MHz and 1443 MHz for the Quadro GP100.
The architectural differences show up in raw throughput. The CMP 90HX achieves 21.89 TFLOPS in FP32 and 21.89 TFLOPS in FP16 (1:1), while the Quadro GP100 delivers 10.34 TFLOPS in FP32 and 20.69 TFLOPS in FP16 (2:1). Texture and pixel rates are closer: the CMP 90HX has 342.0 GTexel/s and 136.8 GPixel/s, against 323.2 GTexel/s and 138.5 GPixel/s for the Quadro GP100. The CMP 90HX also has higher API support, including DirectX 12 Ultimate (12_2) and Vulkan 1.4, whereas the Quadro GP100 tops out at DirectX 12 (12_1) and Vulkan 1.3.
The Verdict
The data points to a clear winner for general compute: the Quadro GP100. Its 26.7% lead in Geekbench OpenCL is decisive, and its 93rd percentile standing versus the CMP 90HX’s 90th confirms it is the stronger performer in this test. The GP100 also offers more memory (16 GB vs 10 GB) and a wider bus (4096-bit vs 320-bit), which can matter for large datasets. However, the CMP 90HX is not without its own advantages: it has higher FP32 throughput (21.89 TFLOPS vs 10.34 TFLOPS), more shading units (6400 vs 3584), and includes RT and tensor cores, plus a newer API feature set.
Who should pick which? If the workload is OpenCL-based and memory capacity or bandwidth is critical, the Quadro GP100 is the safer choice based on benchmark results. Its higher score and larger memory pool make it suitable for professional compute tasks. The CMP 90HX, being a mining card with no display outputs and a PCIe 1.0 x4 interface, is clearly specialized for a niche. Its higher FP32 rate and tensor cores could benefit certain compute workloads, but its lack of display outputs and limited bus interface make it impractical for general use. The data suggests the Quadro GP100 is the more versatile and faster card in the tested metric, making it the recommendation for anyone prioritizing OpenCL performance.
Specification Differences
- Process Node: Quadro GP100 is 16 nm (TSMC); CMP 90HX is 8 nm (Samsung).
- Transistors: 15,300 million vs 28,300 million.
- Die Size: 610 mm² vs 628 mm².
- Transistor Density: 25.1M / mm² vs 45.1M / mm².
- Base Clock: 1304 MHz vs 1500 MHz.
- Boost Clock: 1443 MHz vs 1710 MHz.
- Memory Clock: 715 MHz / 1430 Mbps effective vs 1188 MHz / 19 Gbps effective.
- Memory Size: 16 GB vs 10 GB.
- Memory Type: HBM2 vs GDDR6X.
- Memory Bus: 4096 bit vs 320 bit.
- Memory Bandwidth: 732.2 GB/s vs 760.3 GB/s.
- Shading Units: 3584 vs 6400.
- TMUs: 224 vs 200.
- ROPs: 96 vs 80.
- RT Cores: None vs 50.
- Tensor Cores: None vs 200.
- Pixel Rate: 138.5 GPixel/s vs 136.8 GPixel/s.
- Texture Rate: 323.2 GTexel/s vs 342.0 GTexel/s.
- FP32: 10.34 TFLOPS vs 21.89 TFLOPS.
- FP16: 20.69 TFLOPS (2:1) vs 21.89 TFLOPS (1:1).
- TDP: 235 W vs 320 W.
- Power Connectors: 1x 8-pin vs 2x 8-pin.
- Suggested PSU: 550 W vs 700 W.
- Bus Interface: PCIe 3.0 x16 vs PCIe 1.0 x4.
- Display Outputs: 1x DVI, 4x DisplayPort 1.4a vs No outputs.
- DirectX: 12 (12_1) vs 12 Ultimate (12_2).
- Vulkan: 1.3 vs 1.4.
- Length: 267 mm (10.5 inches) vs 285 mm (11.2 inches).
- Release Date: 2016-09-30 vs 2021-07-27.
FAQ
Q: Which card has a higher Geekbench OpenCL score?
A: The Quadro GP100 scores 87445, which is 26.7% higher than the CMP 90HX’s 69000.
Q: Does the CMP 90HX have any compute advantages over the Quadro GP100?
A: Yes. The CMP 90HX has higher FP32 throughput (21.89 TFLOPS vs 10.34 TFLOPS), more shading units (6400 vs 3584), and includes 50 RT cores and 200 tensor cores, which the Quadro GP100 lacks.
Q: What are the memory differences between the two cards?
A: The Quadro GP100 has 16 GB of HBM2 on a 4096-bit bus with 732.2 GB/s bandwidth. The CMP 90HX has 10 GB of GDDR6X on a 320-bit bus with 760.3 GB/s bandwidth.
Q: Can the CMP 90HX be used for display output?
A: No. The CMP 90HX has no display outputs, while the Quadro GP100 offers 1x DVI and 4x DisplayPort 1.4a.
Q: How does each card compare to its nearest rivals?
A: The Quadro GP100 is 0.4% ahead of the AMD Radeon PRO W7600 and 2.1% ahead of the NVIDIA CMP 40HX. The CMP 90HX is 0.3% ahead of the Intel Arc A770 and 0.6% ahead of the AMD Radeon Instinct MI25.
Q: Which card has a higher transistor density?
A: The CMP 90HX has a density of 45.1 million transistors per mm², versus 25.1 million per mm² for the Quadro GP100.
Where Each One Wins
The Quadro GP100 wins in the only benchmark tested, Geekbench OpenCL, with a 26.7% lead. It also wins on memory capacity (16 GB vs 10 GB) and memory bus width (4096-bit vs 320-bit), which can be decisive for large working sets. Its lower TDP (235 W vs 320 W) and single 8-pin connector mean it is easier to power, and its display outputs make it usable in a standard workstation setup. Its PCIe 3.0 x16 interface is far more flexible than the CMP 90HX’s PCIe 1.0 x4, which is a major practical advantage.
The CMP 90HX wins on raw FP32 compute, delivering more than double the TFLOPS (21.89 vs 10.34) and offering FP16 at a 1:1 ratio versus the GP100’s 2:1. It also has more shading units (6400 vs 3584) and is the only one with RT and tensor cores, making it more capable for workloads that leverage those features. Its memory bandwidth is slightly higher (760.3 GB/s vs 732.2 GB/s), and it supports newer APIs, including DirectX 12 Ultimate and Vulkan 1.4. However, these advantages do not translate into a win in the OpenCL benchmark, where the GP100 remains faster. For compute tasks that use tensor cores or require the latest API features, the CMP 90HX may be preferable, but the data shows the Quadro GP100 is the stronger all-around performer.