NVIDIA CMP 40HX vs NVIDIA GB10 Comparison

NVIDIA
GEFORCE

NVIDIA CMP 40HX

CORE STATE TU106
VRAM 8 GB
CLOCK SPEED 1650 MHz
TDP 185 W
BUS WIDTH 256 bit
ARCHITECTURE Turing
nm
PROCESS 12 nm
LAUNCH DATE 2021
VS
NVIDIA
GEFORCE

GB10

CORE STATE GB20B
VRAM 128 GB
CLOCK SPEED 2418 MHz
TDP 140 W
BUS WIDTH 256 bit
ARCHITECTURE Blackwell 2.0
nm
PROCESS 5 nm
LAUNCH DATE 2025

PERFORMANCE BENCHMARKS

geekbench_opencl
93,395
120,137
geekbench_vulkan
77,879
114,648

Analysis: NVIDIA CMP 40HX vs NVIDIA GB10

The NVIDIA GB10 is a 2025 Blackwell 2.0 server part with a 29.71 TFLOPS FP32 rating, while the NVIDIA CMP 40HX is a 2021 Turing-based mining card with 7.603 TFLOPS FP32. Across the two shared Geekbench tests, the GB10 wins decisively in both, with a 28.6% OpenCL advantage and a 47.2% Vulkan advantage. The data shows a generational gap that extends well beyond raw compute, touching memory capacity, bandwidth, and platform integration.

Head-to-Head Benchmarks

The two benchmark results are one-sided, but the margins tell a nuanced story. In Geekbench OpenCL, the GB10 scores 120,137 against the CMP 40HX’s 93,395, a delta of 28.6%. That gap is substantial, but it is smaller than the Vulkan spread. In Geekbench Vulkan, the GB10 posts 114,648 versus 77,879, yielding a 47.2% lead. The GB10’s Vulkan advantage is nearly double its OpenCL advantage, suggesting the newer architecture handles the lower-level API more efficiently relative to the older Turing design.

The GB10’s average benchmark score of 117,393 places it at the 95th percentile of all GPUs. Its nearest rival, the NVIDIA RTX 4000 SFF Ada Generation, averages 117,088, meaning the GB10 is just 0.3% ahead. That is effectively a tie. The AMD Radeon PRO W7700 averages 118,976, putting the GB10 1.3% behind. The GB10 also beats the NVIDIA Tesla V100 SXM2 16 GB by 2.6% and the NVIDIA RTX A5500 Mobile by 3%.

The CMP 40HX’s 85,637 average score sits at the 93rd percentile. Its closest competitor is the AMD Radeon PRO W7600 at 87,108, which is 1.7% ahead. The NVIDIA Quadro GP100 averages 87,445, 2.1% ahead. The CMP 40HX does beat the AMD Radeon PRO W6600 by 4.4% and the AMD Radeon Pro Vega 64X by 5.8%. So while the CMP 40HX is competitive within its own peer group, that group is far below the GB10’s tier. The 28.6% OpenCL and 47.2% Vulkan deltas dwarf the single-digit percentage differences the CMP 40HX sees against its own rivals.

Looking at the raw specs that drive these scores, the GB10’s 6,144 shading units and 384 tensor cores dwarf the CMP 40HX’s 2,304 shading units and 288 tensor cores. The GB10 also has more RT cores (48 vs 36) and TMUs (384 vs 144). However, the CMP 40HX has more ROPs (64 vs 48) and a higher memory bandwidth (448.0 GB/s vs 273.2 GB/s). The bandwidth advantage does not translate into benchmark wins, likely because the GB10 compensates with 128 GB of LPDDR5X memory versus 8 GB of GDDR6, a 16x capacity difference that matters for large working sets.

The Verdict

The data supports only one choice for anyone running OpenCL or Vulkan workloads: the NVIDIA GB10. It wins both head-to-head tests, has a higher average score (117,393 vs 85,637), and sits at a higher percentile (95th vs 93rd). The GB10’s 28.6% OpenCL lead is significant, but its 47.2% Vulkan lead is overwhelming. The CMP 40HX is not just slower; it is in a different performance class, closer to the AMD Radeon PRO W6600 (4.4% ahead) than to the GB10.

The CMP 40HX’s only technical advantages are memory bandwidth (448.0 GB/s vs 273.2 GB/s), ROP count (64 vs 48), and API support — it supports DirectX 12 Ultimate, OpenGL 4.6, and Vulkan 1.4, while the GB10 lists N/A for all three APIs. But those API features are irrelevant if the card cannot deliver competitive frame rates or compute throughput. The GB10’s higher pixel rate (116.1 GPixel/s vs 105.6 GPixel/s) and texture rate (928.5 GTexel/s vs 237.6 GTexel/s) reinforce its compute lead.

For a buyer choosing between these two, the GB10 is the only rational pick for performance. The CMP 40HX is end-of-life, while the GB10 is active. The GB10 also has a display output (1x HDMI), while the CMP 40HX has none. The GB10 uses no power connectors and is an IGP form factor, whereas the CMP 40HX requires a 1x 8-pin connector and is dual-slot. The GB10’s suggested PSU is 300 W versus 450 W for the CMP 40HX, despite the GB10 having a lower TDP (140 W vs 185 W).

Where Each One Wins

The GB10 wins every benchmark category where both have data. Its OpenCL score of 120,137 is 28.6% higher than the CMP 40HX’s 93,395. Its Vulkan score of 114,648 is 47.2% higher than the CMP 40HX’s 77,879. The GB10 also wins on compute throughput — 29.71 TFLOPS FP32 versus 7.603 TFLOPS FP32 — and on FP16, where the GB10 delivers 29.71 TFLOPS (1:1 ratio) versus the CMP 40HX’s 15.21 TFLOPS (2:1 ratio). The GB10’s texture rate of 928.5 GTexel/s is nearly 4x the CMP 40HX’s 237.6 GTexel/s.

The CMP 40HX wins only on memory bandwidth (448.0 GB/s vs 273.2 GB/s) and ROP count (64 vs 48). That bandwidth advantage could help in memory-bound scenarios, but the GB10’s 128 GB capacity and 16x larger memory pool make it the better choice for datasets that exceed 8 GB. The CMP 40HX also wins on API support — it has DirectX 12 Ultimate, OpenGL 4.6, and Vulkan 1.4, while the GB10 lists N/A for all three. For legacy applications requiring those APIs, the CMP 40HX is the only option, but it is a narrow use case.

The CMP 40HX’s PCIe 1.0 x4 interface is a significant bottleneck, whereas the GB10 uses PCIe 5.0 x16. Even with its higher bandwidth, the CMP 40HX’s older bus interface limits data transfer rates. The GB10 also has a lower TDP (140 W vs 185 W) and a lower suggested PSU (300 W vs 450 W), making it more power-efficient per unit of performance, despite the CMP 40HX being the older, smaller card.

FAQ

Q: Which card has the higher average benchmark score?

A: The NVIDIA GB10 averages 117,393, while the NVIDIA CMP 40HX averages 85,637. The GB10 sits at the 95th percentile of all GPUs, versus the 93rd percentile for the CMP 40HX.

Q: How big is the gap in Vulkan performance?

A: The GB10 scores 114,648 in Geekbench Vulkan, which is 47.2% higher than the CMP 40HX’s 77,879. This is the largest performance delta between the two cards in any benchmark.

Q: Does the CMP 40HX have any advantages over the GB10?

A: Yes. The CMP 40HX has higher memory bandwidth (448.0 GB/s vs 273.2 GB/s), more ROPs (64 vs 48), and supports DirectX 12 Ultimate, OpenGL 4.6, and Vulkan 1.4, whereas the GB10 lists N/A for those APIs.

Q: What is the memory capacity difference?

A: The GB10 has 128 GB of LPDDR5X memory, while the CMP 40HX has 8 GB of GDDR6. The GB10 offers 16 times the capacity, which is critical for large datasets.

Q: Which card is more power-hungry?

A: The CMP 40HX has a 185 W TDP and requires a 450 W suggested PSU, while the GB10 has a 140 W TDP and a 300 W suggested PSU. The GB10 also uses no power connectors, whereas the CMP 40HX needs a 1x 8-pin connector.

Q: How does each card compare to its nearest rival?

A: The GB10 is 0.3% ahead of the NVIDIA RTX 4000 SFF Ada Generation and 1.3% behind the AMD Radeon PRO W7700. The CMP 40HX is 1.7% behind the AMD Radeon PRO W7600 and 4.4% ahead of the AMD Radeon PRO W6600.

Architecture Differences

The GB10 uses the GB20B chip on a 5 nm process from TSMC, with a 382 mm² die size. The CMP 40HX uses the TU106 chip on a 12 nm process, also from TSMC, with a 445 mm² die and 10,800 million transistors. The GB10’s transistor count is listed as unknown, but its die is smaller despite being on a more advanced node. The GB10 belongs to the Server Blackwell (Bxx) generation with a Blackwell 2.0 architecture, while the CMP 40HX is part of the Mining GPUs generation with a Turing architecture. The GB10’s predecessor is Server Hopper and its successor is Server Rubin, whereas the CMP 40HX has no predecessor or successor listed.

The GB10 has 6,144 shading units, 384 TMUs, 48 ROPs, 48 RT cores, and 384 tensor cores. The CMP 40HX has 2,304 shading units, 144 TMUs, 64 ROPs, 36 RT cores, and 288 tensor cores. The GB10’s FP32 throughput is 29.71 TFLOPS, nearly four times the CMP 40HX’s 7.603 TFLOPS. The GB10’s FP16 is also 29.71 TFLOPS at a 1:1 ratio, while the CMP 40HX reaches 15.21 TFLOPS at a 2:1 ratio. The GB10’s texture rate of 928.5 GTexel/s is 3.9x the CMP 40HX’s 237.6 GTexel/s, and its pixel rate of 116.1 GPixel/s is 10% higher.

Memory and platform features differ sharply. The GB10 has 128 GB of LPDDR5X on a 256-bit bus with 273.2 GB/s bandwidth, while the CMP 40HX has 8 GB of GDDR6 on a 256-bit bus with 448.0 GB/s bandwidth. The GB10 uses PCIe 5.0 x16, whereas the CMP 40HX uses PCIe 1.0 x4. The GB10 is an IGP form factor with no power connectors and a 140 W TDP, while the CMP 40HX is dual-slot, requires a 1x 8-pin connector, and has a 185 W TDP. The GB10 has one HDMI output; the CMP 40HX has no display outputs. The GB10’s suggested PSU is 300 W, versus 450 W for the CMP 40HX. The GB10 is 150 mm long, 51 mm high, and 150 mm wide; the CMP 40HX is 229 mm long, 111 mm high, and 35 mm wide. The GB10 was released on October 14, 2025, with a launch MSRP of 3,999 USD, while the CMP 40HX was released on February 24, 2021, with a launch MSRP of 699 USD. The GB10 is active in production, while the CMP 40HX is end-of-life.

DETAILED SPECIFICATIONS

SPECIFICATION
CMP 40HX
GB10
Core Specs
Shading Units
2,304
6,144 +166.7%
Shaders
2,304
6,144 +166.7%
TMUs
144
384 +166.7%
ROPs
64
48 -25.0%
SM Count
36
48 +33.3%
Clocks
Base Clock
1470 MHz
1665 MHz
Boost Clock
1650 MHz
2418 MHz
Memory Clock
1750 MHz 14 Gbps effective
1067 MHz 8.5 Gbps effective
Memory
Memory Size
8 GB
128 GB
VRAM (MB)
8,192
131,072 +1500.0%
Memory Type
GDDR6
LPDDR5X
Memory Bus
256 bit
256 bit
Bandwidth
448.0 GB/s
273.2 GB/s
Cache
L1 Cache
64 KB (per SM)
128 KB (per SM)
L2 Cache
4 MB
50 MB
Performance
Pixel Rate
105.6 GPixel/s
116.1 GPixel/s
Texture Rate
237.6 GTexel/s
928.5 GTexel/s
FP32 (TFLOPS)
7.603 TFLOPS
29.71 TFLOPS
FP64 (TFLOPS)
237.6 GFLOPS (1:32)
464.3 GFLOPS (1:64)
FP16 (TFLOPS)
15.21 TFLOPS (2:1)
29.71 TFLOPS (1:1)
AI/RT
RT Cores
36
48 +33.3%
Tensor Cores
288
384 +33.3%
Power
TDP
185 W
140 W
TDP (W)
185
140 -24.3%
Suggested PSU
450 W
300 W
Power Connectors
1x 8-pin
None
Architecture
Architecture
Turing
Blackwell 2.0
GPU Name
TU106
GB20B
Generation
Mining GPUs
Server Blackwell (Bxx)
Process Size
12 nm
5 nm
Transistors
10,800 million
unknown
Die Size
445 mm²
382 mm²
Foundry
TSMC
TSMC
Density
24.3M / mm²
API Support
DirectX
12 Ultimate (12_2)
OpenGL
4.6
Vulkan
1.4
OpenCL
3.0
3.0
CUDA
7.5
12.1
Shader Model
6.8
Physical
Slot Width
Dual-slot
IGP
Length
229 mm 9 inches
150 mm 5.9 inches
Height
111 mm 4.4 inches
51 mm 2 inches
Outputs
No outputs
1x HDMI
Bus Interface
PCIe 1.0 x4
PCIe 5.0 x16
Other
Launch Price
699 USD
3,999 USD
Production
End-of-life
Active
Predecessor
Server Hopper
Successor
Server Rubin
View CMP 40HX Details View GB10 Details