NVIDIA CMP 90HX vs NVIDIA Quadro P6000 Comparison

NVIDIA
GEFORCE

NVIDIA CMP 90HX

CORE STATE GA102
VRAM 10 GB
CLOCK SPEED 1710 MHz
TDP 320 W
BUS WIDTH 320 bit
ARCHITECTURE Ampere
nm
PROCESS 8 nm
LAUNCH DATE 2021
VS
NVIDIA
GEFORCE

Quadro P6000

CORE STATE GP102
VRAM 24 GB
CLOCK SPEED 1645 MHz
TDP 250 W
BUS WIDTH 384 bit
ARCHITECTURE Pascal
nm
PROCESS 16 nm
LAUNCH DATE 2016

PERFORMANCE BENCHMARKS

geekbench_opencl
69,000
66,382
geekbench_vulkan
N/A
73,590

Analysis: NVIDIA CMP 90HX vs NVIDIA Quadro P6000

NVIDIA’s Quadro P6000 and CMP 90HX are both end-of-life workstation and mining-oriented GPUs, respectively, yet they deliver nearly identical average benchmark scores. The CMP 90HX edges out the Quadro P6000 in the available Geekbench OpenCL test, but the architectural gap between the Pascal and Ampere generations tells a more nuanced story. The data shows a 1.4% advantage for the Quadro P6000 in average score (69,986 vs. 69,000), placing both at the 90th percentile among all GPUs. This is a close contest decided by workload-specific strengths rather than outright dominance.

Head-to-Head Benchmarks

The only direct benchmark comparison available is Geekbench OpenCL, and the CMP 90HX wins it with a score of 69,000 against the Quadro P6000’s 66,382. That is a 3.8% margin — a meaningful but not overwhelming lead. The CMP 90HX’s advantage here aligns with its raw compute specifications, though the Quadro P6000 remains competitive due to its higher memory capacity and mature driver optimizations.

Looking at broader rivalries, the Quadro P6000 sits within 0.2% of the AMD Radeon Pro WX 8200 (69,870) and 0.2% behind the NVIDIA RTX A3000 Mobile (70,140). The CMP 90HX, meanwhile, is 0.3% ahead of the Intel Arc A770 (68,809) and 0.6% ahead of the AMD Radeon Instinct MI25 (68,562). These deltas are small enough to suggest that both cards are trading blows with contemporaries from different vendors, but the CMP 90HX’s single benchmark victory does not translate into a clear overall superiority.

The Quadro P6000 also has a Geekbench Vulkan score of 73,590, which is not available for the CMP 90HX due to its lack of display outputs. That absence matters: the Vulkan result demonstrates the Quadro’s ability to handle graphics APIs, a capability the CMP 90HX simply does not possess. In compute-only terms, the CMP 90HX’s 3.8% OpenCL win is its sole claim to fame, and it is a narrow one.

Architecture Differences

The two GPUs come from different process nodes and foundries. The Quadro P6000 uses a 16 nm TSMC process with 11,800 million transistors on a 471 mm² die, yielding a transistor density of 25.1M per mm². The CMP 90HX is built on Samsung’s 8 nm node, packing 28,300 million transistors into a larger 628 mm² die, with a density of 45.1M per mm². The newer node gives the CMP 90HX a 58% higher transistor count and nearly double the density, enabling more complex compute units.

Core configurations diverge sharply. The Quadro P6000 has 3,840 shading units, 240 TMUs, and 96 ROPs, while the CMP 90HX offers 6,400 shading units, 200 TMUs, and 80 ROPs. The CMP 90HX has 67% more shading units but 17% fewer TMUs and 17% fewer ROPs. It also includes 50 RT cores and 200 tensor cores — features entirely absent from the Pascal-based Quadro P6000, which predates ray tracing and tensor core acceleration.

Memory is another stark differentiator. The Quadro P6000 carries 24 GB of GDDR5X on a 384-bit bus, delivering 432.8 GB/s of bandwidth. The CMP 90HX has only 10 GB of GDDR6X but on a narrower 320-bit bus, achieving 760.3 GB/s. That is a 75.7% bandwidth advantage for the CMP 90HX, despite the smaller capacity. Clock speeds are similar: the Quadro P6000 runs at 1506 MHz base and 1645 MHz boost, while the CMP 90HX is at 1500 MHz base and 1710 MHz boost. The memory clocks differ more, with the CMP 90HX at 19 Gbps effective versus the Quadro’s 9 Gbps.

Raw throughput figures favor the CMP 90HX. It produces 21.89 TFLOPS FP32 and 21.89 TFLOPS FP16 (1:1 ratio), whereas the Quadro P6000 delivers 12.63 TFLOPS FP32 and a paltry 197.4 GFLOPS FP16 (1:64 ratio). The CMP 90HX is 73% faster in FP32 and over 100x faster in FP16, reflecting its Ampere architecture’s dedicated tensor cores. Pixel rate goes the other way: the Quadro P6000 hits 157.9 GPixel/s versus 136.8 GPixel/s for the CMP 90HX, a 15% advantage for the older card.

Where Each One Wins

The CMP 90HX is the compute monster. Its FP32 throughput of 21.89 TFLOPS and FP16 performance of 21.89 TFLOPS make it suited for tasks that rely on massive parallel math, such as machine learning inference or scientific simulations. The 200 tensor cores amplify this further, though benchmark data only covers OpenCL. The 760.3 GB/s memory bandwidth also gives it an edge in bandwidth-bound workloads, and the 3.8% OpenCL victory reflects that.

The Quadro P6000 wins on capacity and graphics features. Its 24 GB VRAM is 140% larger than the CMP 90HX’s 10 GB, making it the clear choice for datasets or rendering scenes that exceed 10 GB. The Quadro also has display outputs (1x DVI, 4x DisplayPort 1.4a), enabling professional visualization workflows that the CMP 90HX, with no outputs, cannot handle. The Vulkan score of 73,590 demonstrates its graphics capability, and its higher pixel rate (157.9 GPixel/s) helps with rasterization-heavy tasks.

For traditional workstation use — CAD, 3D modeling, video editing with GPU acceleration — the Quadro P6000’s driver support and display connectivity are decisive. For headless compute farms or mining, the CMP 90HX’s raw numbers and bandwidth win out. The CMP 90HX also requires a 700 W PSU versus 600 W for the Quadro, and uses two 8-pin connectors instead of one, reflecting its higher 320 W TDP against the Quadro’s 250 W.

FAQ

Q: Which GPU has higher average benchmark performance?

A: The Quadro P6000 leads with an average benchmark score of 69,986, which is 1.4% higher than the CMP 90HX’s 69,000.

Q: Does the CMP 90HX support ray tracing?

A: Yes, the CMP 90HX includes 50 RT cores, a feature absent from the Pascal-based Quadro P6000.

Q: Can the CMP 90HX be used for display output?

A: No, the CMP 90HX has no display outputs, whereas the Quadro P6000 offers 1x DVI and 4x DisplayPort 1.4a.

Q: How much faster is the CMP 90HX in FP32 compute?

A: The CMP 90HX delivers 21.89 TFLOPS FP32, which is 73% higher than the Quadro P6000’s 12.63 TFLOPS.

Q: Which card has more memory?

A: The Quadro P6000 has 24 GB of GDDR5X, while the CMP 90HX has 10 GB of GDDR6X.

Q: What is the memory bandwidth difference?

A: The CMP 90HX achieves 760.3 GB/s, which is 75.7% higher than the Quadro P6000’s 432.8 GB/s.

The Verdict

The CMP 90HX wins on raw compute and bandwidth. Its 21.89 TFLOPS FP32, 21.89 TFLOPS FP16, and 760.3 GB/s memory bandwidth are all far superior to the Quadro P6000’s 12.63 TFLOPS, 197.4 GFLOPS, and 432.8 GB/s. The Geekbench OpenCL score of 69,000 versus 66,382 confirms this in practice. Anyone building a headless compute node or mining rig should choose the CMP 90HX, provided they can handle its 320 W TDP and dual 8-pin power requirement.

The Quadro P6000 is the better workstation card. Its 24 GB VRAM dwarfs the CMP 90HX’s 10 GB, and its display outputs allow it to drive professional monitors. The Vulkan score of 73,590 shows it can handle graphics APIs effectively, and its higher pixel rate (157.9 GPixel/s) benefits rendering tasks. The Quadro also has a slightly higher average benchmark score (69,986 vs. 69,000), indicating more consistent performance across a wider range of tests.

The choice hinges on use case. If the workload is compute-heavy, memory-bandwidth-sensitive, and requires no display, the CMP 90HX’s 3.8% OpenCL win and 73% FP32 lead make it the obvious pick. If the workload involves large datasets, graphics output, or legacy software optimized for Pascal, the Quadro P6000’s capacity and features win out. The 90th percentile ranking for both cards shows they are both strong performers, but they serve different masters.

Specification Differences

| Specification | NVIDIA Quadro P6000 | NVIDIA CMP 90HX |

|----------------|---------------------|------------------|

| Architecture | Pascal | Ampere |

| Process Node | 16 nm | 8 nm |

| Foundry | TSMC | Samsung |

| Transistors | 11,800 million | 28,300 million |

| Die Size | 471 mm² | 628 mm² |

| Transistor Density | 25.1M / mm² | 45.1M / mm² |

| Base Clock | 1506 MHz | 1500 MHz |

| Boost Clock | 1645 MHz | 1710 MHz |

| Memory Clock | 1127 MHz (9 Gbps effective) | 1188 MHz (19 Gbps effective) |

| Memory Size | 24 GB | 10 GB |

| Memory Type | GDDR5X | GDDR6X |

| Memory Bus Width | 384 bit | 320 bit |

| Memory Bandwidth | 432.8 GB/s | 760.3 GB/s |

| Shading Units | 3840 | 6400 |

| TMUs | 240 | 200 |

| ROPs | 96 | 80 |

| RT Cores | N/A | 50 |

| Tensor Cores | N/A | 200 |

| Pixel Rate | 157.9 GPixel/s | 136.8 GPixel/s |

| Texture Rate | 394.8 GTexel/s | 342.0 GTexel/s |

| FP32 | 12.63 TFLOPS | 21.89 TFLOPS |

| FP16 | 197.4 GFLOPS (1:64) | 21.89 TFLOPS (1:1) |

| TDP | 250 W | 320 W |

| Power Connectors | 1x 8-pin | 2x 8-pin |

| Suggested PSU | 600 W | 700 W |

| Bus Interface | PCIe 3.0 x16 | PCIe 1.0 x4 |

| Display Outputs | 1x DVI, 4x DisplayPort 1.4a | No outputs |

| DirectX | 12 (12_1) | 12 Ultimate (12_2) |

| Vulkan | 1.4 | 1.4 |

| Length | 267 mm (10.5 inches) | 285 mm (11.2 inches) |

| Height | 111 mm (4.4 inches) | 112 mm (4.4 inches) |

| Release Date | 2016-09-30 | 2021-07-27 |

| Launch MSRP | 5,999 USD | N/A |

DETAILED SPECIFICATIONS

SPECIFICATION
CMP 90HX
Quadro P6000
Core Specs
Shading Units
6,400
3,840 -40.0%
Shaders
6,400
3,840 -40.0%
TMUs
200
240 +20.0%
ROPs
80
96 +20.0%
SM Count
50
30 -40.0%
Clocks
Base Clock
1500 MHz
1506 MHz
Boost Clock
1710 MHz
1645 MHz
Memory Clock
1188 MHz 19 Gbps effective
1127 MHz 9 Gbps effective
Memory
Memory Size
10 GB
24 GB
VRAM (MB)
10,240
24,576 +140.0%
Memory Type
GDDR6X
GDDR5X
Memory Bus
320 bit
384 bit
Bandwidth
760.3 GB/s
432.8 GB/s
Cache
L1 Cache
128 KB (per SM)
48 KB (per SM)
L2 Cache
5 MB
3 MB
Performance
Pixel Rate
136.8 GPixel/s
157.9 GPixel/s
Texture Rate
342.0 GTexel/s
394.8 GTexel/s
FP32 (TFLOPS)
21.89 TFLOPS
12.63 TFLOPS
FP64 (TFLOPS)
342.0 GFLOPS (1:64)
394.8 GFLOPS (1:32)
FP16 (TFLOPS)
21.89 TFLOPS (1:1)
197.4 GFLOPS (1:64)
AI/RT
RT Cores
50
Tensor Cores
200
Power
TDP
320 W
250 W
TDP (W)
320
250 -21.9%
Suggested PSU
700 W
600 W
Power Connectors
2x 8-pin
1x 8-pin
Architecture
Architecture
Ampere
Pascal
GPU Name
GA102
GP102
Generation
Mining GPUs
Quadro Pascal (Px000)
Process Size
8 nm
16 nm
Transistors
28,300 million
11,800 million
Die Size
628 mm²
471 mm²
Foundry
Samsung
TSMC
Density
45.1M / mm²
25.1M / mm²
API Support
DirectX
12 Ultimate (12_2)
12 (12_1)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
3.0
3.0
CUDA
8.6
6.1
Shader Model
6.8
6.8
Physical
Slot Width
Dual-slot
Dual-slot
Length
285 mm 11.2 inches
267 mm 10.5 inches
Height
112 mm 4.4 inches
111 mm 4.4 inches
Outputs
No outputs
1x DVI4x DisplayPort 1.4a
Bus Interface
PCIe 1.0 x4
PCIe 3.0 x16
Other
Launch Price
5,999 USD
Production
End-of-life
End-of-life
Predecessor
Quadro Maxwell
Successor
Quadro Volta
View CMP 90HX Details View Quadro P6000 Details