GPU Comparison

NVIDIA
GEFORCE

NVIDIA CMP 40HX

CORE STATE TU106
VRAM 8 GB
CLOCK SPEED 1650 MHz
TDP 185 W
BUS WIDTH 256 bit
ARCHITECTURE Turing
nm
PROCESS 12 nm
LAUNCH DATE 2021
VS
NVIDIA
GEFORCE

Quadro RTX 6000

CORE STATE TU102
VRAM 24 GB
CLOCK SPEED 1770 MHz
TDP 260 W
BUS WIDTH 384 bit
ARCHITECTURE Turing
nm
PROCESS 12 nm
LAUNCH DATE 2018

PERFORMANCE BENCHMARKS

geekbench_opencl
93,395
74,179
geekbench_vulkan
77,879
129,564

Analysis: NVIDIA CMP 40HX vs NVIDIA Quadro RTX 6000

The NVIDIA Quadro RTX 6000 and the NVIDIA CMP 40HX are both Turing-architecture cards, but they target completely different workloads. The data shows a clear split: the CMP 40HX wins the OpenCL benchmark, while the Quadro RTX 6000 dominates Vulkan. This head-to-head is less about overall superiority and more about which API and use case matters more to the buyer.

Head-to-Head Benchmarks

The two cards split their benchmark results exactly one win each. In Geekbench OpenCL, the NVIDIA CMP 40HX scores 93,395, which is 20.6% higher than the Quadro RTX 6000’s 74,179. That is a decisive margin, placing the mining-focused card clearly ahead in compute workloads that leverage OpenCL. The CMP 40HX’s average benchmark score of 85,637 also exceeds the Quadro’s 101,872 average, but that average includes both tests, so the OpenCL result is the CMP’s primary strength.

In Geekbench Vulkan, the tables turn dramatically. The Quadro RTX 6000 posts 129,564, which is 66.4% ahead of the CMP 40HX’s 77,879. That is a massive gap, more than three times the size of the CMP’s OpenCL advantage. For any graphics or GPU-compute workload that uses Vulkan, the Quadro RTX 6000 is the clear choice. The Quadro’s average benchmark score of 101,872 reflects this Vulkan dominance, pulling its combined average above the CMP’s.

Looking at the broader context, the Quadro RTX 6000 sits in the 94th percentile of all GPUs, while the CMP 40HX is just one point behind at the 93rd percentile. Both are high-performing parts, but their strengths are polar opposites. The CMP 40HX’s nearest rivals include the AMD Radeon PRO W7600 (deltaPct -1.7%), NVIDIA Quadro GP100 (-2.1%), AMD Radeon PRO W6600 (+4.4%), and AMD Radeon Pro Vega 64X (+5.8%). The Quadro RTX 6000, meanwhile, is 4.5% ahead of the AMD Radeon RX 7900M, 4.9% ahead of the AMD Radeon Pro VII, and 4.6% behind the AMD Radeon Pro Vega II Duo, with the AMD Radeon Pro W6600X 5.1% ahead. These rival deltas show that both cards are competitive in their respective niches, but neither is a runaway leader against all comers.

Architecture Differences

Both cards use the Turing architecture and are built on TSMC’s 12 nm process, but the silicon underneath is very different. The Quadro RTX 6000 uses the TU102 chip, which packs 18,600 million transistors on a 754 mm² die. The CMP 40HX uses the smaller TU106 chip, with 10,800 million transistors on a 445 mm² die. Transistor density is nearly identical, 24.7M per mm² for the Quadro versus 24.3M per mm² for the CMP, so the performance gap comes from sheer scale, not efficiency.

The Quadro RTX 6000 has 4,608 shading units, 288 texture mapping units, and 96 ROPs. The CMP 40HX is cut down to 2,304 shading units, 144 TMUs, and 64 ROPs. That is exactly half the shaders and TMUs, and two-thirds of the ROPs, which explains the large FP32 compute gap: 16.31 TFLOPS for the Quadro versus 7.603 TFLOPS for the CMP. Ray tracing and tensor cores follow the same pattern, the Quadro has 72 RT cores and 576 tensor cores, while the CMP has 36 RT cores and 288 tensor cores.

Clock speeds are surprisingly close. The CMP 40HX has a slightly higher base clock at 1470 MHz versus 1440 MHz, but the Quadro boosts higher at 1770 MHz versus 1650 MHz. Memory is another major divider. The Quadro RTX 6000 features 24 GB of GDDR6 on a 384-bit bus, yielding 672.0 GB/s of bandwidth. The CMP 40HX has only 8 GB of GDDR6 on a 256-bit bus, delivering 448.0 GB/s. Both run memory at 1750 MHz (14 Gbps effective), so the bandwidth difference is purely a function of bus width.

The most striking architectural difference is the bus interface. The Quadro RTX 6000 uses PCIe 3.0 x16, while the CMP 40HX is limited to PCIe 1.0 x4. That is a severe bottleneck for any data transfer between the card and the host system. The CMP 40HX also has no display outputs at all, whereas the Quadro offers 4x DisplayPort 1.4a and 1x USB Type-C. In terms of API support, both cards list DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4, so software compatibility is identical.

Where Each One Wins

The CMP 40HX wins in OpenCL compute. Its 93,395 OpenCL score is 20.6% higher than the Quadro’s, making it the better choice for workloads that rely on that API. This is likely why the card was designed for mining, where OpenCL-based kernels are common. The CMP also has a lower TDP at 185 W versus 260 W, and a smaller footprint at 229 mm length versus 267 mm. Its 8-pin power connector and 450 W suggested PSU are simpler than the Quadro’s 1x 6-pin + 1x 8-pin setup with a 600 W PSU suggestion.

The Quadro RTX 6000 wins in Vulkan, and it wins big. Its 129,564 Vulkan score is 66.4% ahead of the CMP’s 77,879. For any gaming, rendering, or compute workload that leverages Vulkan, the Quadro is overwhelmingly superior. It also has 3x the memory (24 GB versus 8 GB), which matters for large datasets or high-resolution textures. The Quadro’s 72 RT cores and 576 tensor cores enable hardware ray tracing and AI acceleration, features that are present in the CMP but at half the count, 36 RT cores and 288 tensor cores. The Quadro’s display outputs and USB Type-C port make it usable in a workstation environment, while the CMP has none.

In terms of compute throughput, the Quadro’s 16.31 TFLOPS FP32 is more than double the CMP’s 7.603 TFLOPS. Pixel rate is 169.9 GPixel/s versus 105.6 GPixel/s, and texture rate is 509.8 GTexel/s versus 237.6 GTexel/s. These are raw hardware advantages that translate to higher frame rates and faster processing in any graphics-heavy task.

Specification Differences

The key specification differences between the two cards are stark:

  • Chip: TU102 (Quadro) vs TU106 (CMP)
  • Transistors: 18,600 million vs 10,800 million
  • Die Size: 754 mm² vs 445 mm²
  • Base Clock: 1440 MHz vs 1470 MHz
  • Boost Clock: 1770 MHz vs 1650 MHz
  • Memory Size: 24 GB vs 8 GB
  • Memory Bus: 384 bit vs 256 bit
  • Memory Bandwidth: 672.0 GB/s vs 448.0 GB/s
  • Shading Units: 4608 vs 2304
  • TMUs: 288 vs 144
  • ROPs: 96 vs 64
  • RT Cores: 72 vs 36
  • Tensor Cores: 576 vs 288
  • Pixel Rate: 169.9 GPixel/s vs 105.6 GPixel/s
  • Texture Rate: 509.8 GTexel/s vs 237.6 GTexel/s
  • FP32: 16.31 TFLOPS vs 7.603 TFLOPS
  • FP16: 32.62 TFLOPS vs 15.21 TFLOPS
  • TDP: 260 W vs 185 W
  • Power Connectors: 1x 6-pin + 1x 8-pin vs 1x 8-pin
  • Suggested PSU: 600 W vs 450 W
  • Bus Interface: PCIe 3.0 x16 vs PCIe 1.0 x4
  • Display Outputs: 4x DisplayPort 1.4a, 1x USB Type-C vs No outputs
  • Length: 267 mm vs 229 mm
  • Release Date: 2018-08-12 vs 2021-02-24

Both cards share the same process node, foundry (TSMC), memory type (GDDR6), memory clock (1750 MHz, 14 Gbps effective), slot width (Dual-slot), height (111 mm), and API support (DirectX 12 Ultimate, OpenGL 4.6, Vulkan 1.4).

FAQ

Q: Which card has a higher average benchmark score?

A: The NVIDIA Quadro RTX 6000 has an average benchmark score of 101,872, while the NVIDIA CMP 40HX scores 85,637.

Q: How do the two cards compare in OpenCL performance?

A: The CMP 40HX wins in Geekbench OpenCL with a score of 93,395, which is 20.6% higher than the Quadro RTX 6000’s 74,179.

Q: What is the Vulkan performance difference?

A: The Quadro RTX 6000 scores 129,564 in Geekbench Vulkan, which is 66.4% higher than the CMP 40HX’s 77,879.

Q: Do both cards support the same APIs?

A: Yes, both list DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4.

Q: What are the memory specifications for each card?

A: The Quadro RTX 6000 has 24 GB of GDDR6 on a 384-bit bus with 672.0 GB/s bandwidth. The CMP 40HX has 8 GB of GDDR6 on a 256-bit bus with 448.0 GB/s bandwidth.

Q: Can the CMP 40HX be used for display output?

A: No, the CMP 40HX has no display outputs. The Quadro RTX 6000 offers 4x DisplayPort 1.4a and 1x USB Type-C.

The Verdict

The data is unambiguous: these cards serve opposite purposes. The NVIDIA CMP 40HX is a specialized compute card that wins in OpenCL by 20.6%, but it lacks display outputs, uses a crippled PCIe 1.0 x4 interface, and has half the shaders and memory of its rival. Its 93rd percentile ranking and 85,637 average score are respectable, but its architecture is fundamentally cut down for mining workloads.

The NVIDIA Quadro RTX 6000 is the workstation part. It wins Vulkan by 66.4%, offers 24 GB of memory, 72 RT cores, 576 tensor cores, and full display connectivity. Its 94th percentile ranking and 101,872 average score reflect its broader capability. The 4.5% edge over the AMD Radeon RX 7900M and 4.9% over the AMD Radeon Pro VII in nearest-rival comparisons show it holds its own in the high-end segment.

For buyers, the choice depends entirely on the workload. If the task is OpenCL compute with no need for display output or host bandwidth, the CMP 40HX’s 20.6% OpenCL advantage makes it a viable option, especially given its lower 185 W TDP and 450 W PSU requirement. But for anything involving Vulkan, graphics, ray tracing, tensor operations, or large memory footprints, the Quadro RTX 6000 is the only rational pick, 66.4% ahead in Vulkan and double the FP32 throughput. The Quadro’s 6,299 USD launch MSRP reflects its professional positioning, while the CMP 40HX’s 699 USD launch MSRP aligns with its mining-focused design. The verdict is straightforward: the Quadro RTX 6000 is the superior all-around GPU, while the CMP 40HX is a niche compute accelerator with one clear strength.

DETAILED SPECIFICATIONS

SPECIFICATION
CMP 40HX
Quadro RTX 6000
Core Specs
Shading Units
2,304
4,608 +100.0%
Shaders
2,304
4,608 +100.0%
TMUs
144
288 +100.0%
ROPs
64
96 +50.0%
SM Count
36
72 +100.0%
Clocks
Base Clock
1470 MHz
1440 MHz
Boost Clock
1650 MHz
1770 MHz
Memory Clock
1750 MHz 14 Gbps effective
1750 MHz 14 Gbps effective
Memory
Memory Size
8 GB
24 GB
VRAM (MB)
8,192
24,576 +200.0%
Memory Type
GDDR6
GDDR6
Memory Bus
256 bit
384 bit
Bandwidth
448.0 GB/s
672.0 GB/s
Cache
L1 Cache
64 KB (per SM)
64 KB (per SM)
L2 Cache
4 MB
6 MB
Performance
Pixel Rate
105.6 GPixel/s
169.9 GPixel/s
Texture Rate
237.6 GTexel/s
509.8 GTexel/s
FP32 (TFLOPS)
7.603 TFLOPS
16.31 TFLOPS
FP64 (TFLOPS)
237.6 GFLOPS (1:32)
509.8 GFLOPS (1:32)
FP16 (TFLOPS)
15.21 TFLOPS (2:1)
32.62 TFLOPS (2:1)
AI/RT
RT Cores
36
72 +100.0%
Tensor Cores
288
576 +100.0%
Power
TDP
185 W
260 W
TDP (W)
185
260 +40.5%
Suggested PSU
450 W
600 W
Power Connectors
1x 8-pin
1x 6-pin + 1x 8-pin
Architecture
Architecture
Turing
Turing
GPU Name
TU106
TU102
Generation
Mining GPUs
Quadro Turing (Tx000)
Process Size
12 nm
12 nm
Transistors
10,800 million
18,600 million
Die Size
445 mm²
754 mm²
Foundry
TSMC
TSMC
Density
24.3M / mm²
24.7M / mm²
API Support
DirectX
12 Ultimate (12_2)
12 Ultimate (12_2)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
3.0
3.0
CUDA
7.5
7.5
Shader Model
6.8
6.8
Physical
Slot Width
Dual-slot
Dual-slot
Length
229 mm 9 inches
267 mm 10.5 inches
Height
111 mm 4.4 inches
111 mm 4.4 inches
Outputs
No outputs
4x DisplayPort 1.4a1x USB Type-C
Bus Interface
PCIe 1.0 x4
PCIe 3.0 x16
Other
Launch Price
699 USD
6,299 USD
Production
End-of-life
End-of-life
Predecessor
Quadro Volta
Successor
Workstation Ampere
View CMP 40HX Details View Quadro RTX 6000 Details