NVIDIA CMP 40HX vs NVIDIA CMP 90HX Comparison

NVIDIA
GEFORCE

NVIDIA CMP 40HX

CORE STATE TU106
VRAM 8 GB
CLOCK SPEED 1650 MHz
TDP 185 W
BUS WIDTH 256 bit
ARCHITECTURE Turing
nm
PROCESS 12 nm
LAUNCH DATE 2021
VS
NVIDIA
GEFORCE

CMP 90HX

CORE STATE GA102
VRAM 10 GB
CLOCK SPEED 1710 MHz
TDP 320 W
BUS WIDTH 320 bit
ARCHITECTURE Ampere
nm
PROCESS 8 nm
LAUNCH DATE 2021

PERFORMANCE BENCHMARKS

geekbench_opencl
93,395
69,000
geekbench_vulkan
77,879
N/A

Analysis: NVIDIA CMP 40HX vs NVIDIA CMP 90HX

The NVIDIA CMP 40HX and NVIDIA CMP 90HX are both end-of-life mining-oriented GPUs with no display outputs, but they serve very different performance tiers. Based on the available benchmark data, the CMP 40HX is the clear winner in the only head-to-head test, delivering a 35.4% higher Geekbench OpenCL score than the CMP 90HX. This result is counterintuitive given the CMP 90HX’s much larger chip and newer architecture, so the data demands a closer look at where each card actually excels.

Where Each One Wins

The CMP 40HX wins the only directly comparable benchmark, Geekbench OpenCL, with a score of 93395 versus 69000 for the CMP 90HX. That is a decisive 35.4% margin, placing the 40HX in the 93rd percentile of all GPUs, while the 90HX sits at the 90th percentile. The 40HX also has a second benchmark result — a Geekbench Vulkan score of 77879 — which the 90HX lacks entirely, giving the 40HX a broader software compatibility profile in the data.

However, the CMP 90HX is not without its own strengths in the specification sheet. Its raw compute resources are substantially higher: 6400 shading units versus 2304, 200 texture mapping units versus 144, and 80 render output units versus 64. It also carries 50 RT cores and 200 tensor cores, compared to 36 RT cores and 288 tensor cores on the 40HX. These figures suggest the 90HX should dominate in compute-heavy workloads that scale with shading unit count, even though the benchmark data does not capture that advantage in the OpenCL test.

The 90HX also wins on memory capacity and bandwidth, with 10 GB of GDDR6X on a 320-bit bus delivering 760.3 GB/s, versus 8 GB of GDDR6 on a 256-bit bus at 448.0 GB/s. For workloads that are memory-bound rather than compute-bound, the 90HX’s 69.7% bandwidth advantage could translate into real-world wins that the single OpenCL score does not reflect. The 40HX, by contrast, wins on efficiency per watt in the data: it achieves its higher benchmark score at 185 W TDP, while the 90HX requires 320 W.

Architecture Differences

The two cards come from different NVIDIA architectures and foundries. The CMP 40HX uses the TU106 chip built on TSMC’s 12 nm process, while the CMP 90HX uses the GA102 chip on Samsung’s 8 nm node. This is a generational leap: the 90HX packs 28,300 million transistors into a 628 mm² die, versus 10,800 million transistors on a 445 mm² die for the 40HX. Transistor density tells the story of the node advantage — the 90HX achieves 45.1 million transistors per mm², nearly double the 40HX’s 24.3 million.

The architecture shift from Turing to Ampere brings a fundamental change in compute precision handling. The 40HX’s FP16 throughput is listed as 15.21 TFLOPS with a 2:1 ratio relative to FP32, meaning it halves FP32 throughput when doing FP16 work. The 90HX, in contrast, delivers 21.89 TFLOPS for both FP16 and FP32 with a 1:1 ratio, indicating it can do full-rate FP16 without sacrificing FP32 performance. This makes the 90HX more flexible for mixed-precision workloads.

Memory technology also diverges: the 40HX uses GDDR6 at 14 Gbps effective, while the 90HX uses GDDR6X at 19 Gbps effective. The 90HX’s memory clock is listed at 1188 MHz base, translating to the higher effective data rate. Both cards share the same PCIe 1.0 x4 bus interface, which is an unusual and restrictive choice for mining cards, and both have no display outputs. The 40HX is physically smaller at 229 mm in length versus 285 mm for the 90HX, and the 40HX requires a single 8-pin power connector while the 90HX needs two.

Head-to-Head Benchmarks

The single head-to-head benchmark is Geekbench OpenCL, and the result is emphatic: the CMP 40HX scores 93395, beating the CMP 90HX’s 69000 by 35.4%. This is a substantial margin that flips the expected hierarchy based on specs. The 40HX’s average benchmark score across its two tests is 85637, while the 90HX’s average is 69000 — a 24.1% gap in favor of the 40HX.

Looking at rival comparisons, the 40HX sits just 1.7% below the AMD Radeon PRO W7600 and 2.1% below the NVIDIA Quadro GP100, while leading the AMD Radeon PRO W6600 by 4.4% and the AMD Radeon Pro Vega 64X by 5.8%. The 90HX, on the other hand, is nearly tied with the Intel Arc A770 (0.3% ahead) and the AMD Radeon Instinct MI25 (0.6% ahead), while trailing the AMD Radeon Pro WX 8200 by 1.2% and the NVIDIA Quadro P6000 by 1.4%. These rival deltas show that the 40HX competes in a higher performance tier than the 90HX in OpenCL, despite the 90HX’s newer architecture.

The 40HX also has a Geekbench Vulkan score of 77879, which is 16.6% below its own OpenCL score. This suggests the 40HX performs better in OpenCL than Vulkan, though both results are strong. The 90HX has no Vulkan benchmark in the data, so no comparison is possible for that API. The overall wins tally is 1 for the 40HX and 0 for the 90HX, but this is based on a single shared test, which limits the conclusiveness of the comparison.

FAQ

Q: Which card has the higher benchmark score?

A: The NVIDIA CMP 40HX scores 93395 in Geekbench OpenCL, while the CMP 90HX scores 69000, giving the 40HX a 35.4% advantage.

Q: Does the CMP 90HX have any benchmark where it wins?

A: No. In the only head-to-head benchmark (Geekbench OpenCL), the CMP 40HX wins. The 90HX has no other benchmark results to compare.

Q: What are the memory differences between the two cards?

A: The CMP 90HX has 10 GB of GDDR6X on a 320-bit bus with 760.3 GB/s bandwidth. The CMP 40HX has 8 GB of GDDR6 on a 256-bit bus with 448.0 GB/s bandwidth.

Q: Which card has more shading units?

A: The CMP 90HX has 6400 shading units, compared to 2304 on the CMP 40HX. The 90HX also has more TMUs (200 vs 144) and ROPs (80 vs 64).

Q: Are these cards different in power requirements?

A: Yes. The CMP 40HX has a 185 W TDP and uses a single 8-pin connector with a 450 W suggested PSU. The CMP 90HX has a 320 W TDP, uses two 8-pin connectors, and requires a 700 W suggested PSU.

Q: What is the transistor count difference?

A: The CMP 90HX has 28,300 million transistors on a 628 mm² die, while the CMP 40HX has 10,800 million transistors on a 445 mm² die. The 90HX’s 8 nm process yields 45.1M transistors per mm² versus 24.3M for the 40HX’s 12 nm process.

The Verdict

The data is unambiguous in one respect: if you are choosing based purely on the available benchmark results, the NVIDIA CMP 40HX is the superior card. It wins the only head-to-head test by 35.4%, holds a higher percentile rank (93rd vs 90th), and has additional benchmark coverage with its Vulkan score. The 40HX also achieves this at a lower TDP of 185 W versus 320 W, making it the more efficient choice in the data.

However, the CMP 90HX should not be dismissed outright. Its specification sheet reveals a card with 2.8 times the shading units, 69.7% more memory bandwidth, and double the transistor count. The 90HX’s 1:1 FP16/FP32 ratio and full-rate FP16 performance at 21.89 TFLOPS indicate it is built for compute workloads that the benchmark data does not capture. If your workload scales with shading unit count or memory bandwidth rather than the specific OpenCL test used here, the 90HX could be the stronger performer.

The practical choice depends on which metric matters more. For a straightforward compute benchmark, the CMP 40HX is the verdict — it simply outperforms the 90HX in the test that matters. For raw theoretical compute and memory resources, the CMP 90HX is the more capable silicon, but the data does not confirm that advantage in practice. Buyers should note that both cards are end-of-life, have no display outputs, and use a restrictive PCIe 1.0 x4 interface, which limits their utility beyond mining or compute tasks.

Specification Differences

| Specification | NVIDIA CMP 40HX | NVIDIA CMP 90HX |

|---|---|---|

| Chip | TU106 | GA102 |

| Architecture | Turing | Ampere |

| Process Node | 12 nm | 8 nm |

| Foundry | TSMC | Samsung |

| Transistors | 10,800 million | 28,300 million |

| Die Size | 445 mm² | 628 mm² |

| Transistor Density | 24.3M / mm² | 45.1M / mm² |

| Base Clock | 1470 MHz | 1500 MHz |

| Boost Clock | 1650 MHz | 1710 MHz |

| Memory Clock | 1750 MHz (14 Gbps effective) | 1188 MHz (19 Gbps effective) |

| Memory Size | 8 GB | 10 GB |

| Memory Type | GDDR6 | GDDR6X |

| Memory Bus Width | 256 bit | 320 bit |

| Memory Bandwidth | 448.0 GB/s | 760.3 GB/s |

| Shading Units | 2304 | 6400 |

| TMUs | 144 | 200 |

| ROPs | 64 | 80 |

| RT Cores | 36 | 50 |

| Tensor Cores | 288 | 200 |

| Pixel Rate | 105.6 GPixel/s | 136.8 GPixel/s |

| Texture Rate | 237.6 GTexel/s | 342.0 GTexel/s |

| FP32 Performance | 7.603 TFLOPS | 21.89 TFLOPS |

| FP16 Performance | 15.21 TFLOPS (2:1) | 21.89 TFLOPS (1:1) |

| TDP | 185 W | 320 W |

| Power Connectors | 1x 8-pin | 2x 8-pin |

| Suggested PSU | 450 W | 700 W |

| Length | 229 mm (9 inches) | 285 mm (11.2 inches) |

| Height | 111 mm (4.4 inches) | 112 mm (4.4 inches) |

| Width | 35 mm (1.4 inches) | Not specified |

| Launch MSRP | 699 USD | Not specified |

| Release Date | 2021-02-24 | 2021-07-27 |

DETAILED SPECIFICATIONS

SPECIFICATION
CMP 40HX
CMP 90HX
Core Specs
Shading Units
2,304
6,400 +177.8%
Shaders
2,304
6,400 +177.8%
TMUs
144
200 +38.9%
ROPs
64
80 +25.0%
SM Count
36
50 +38.9%
Clocks
Base Clock
1470 MHz
1500 MHz
Boost Clock
1650 MHz
1710 MHz
Memory Clock
1750 MHz 14 Gbps effective
1188 MHz 19 Gbps effective
Memory
Memory Size
8 GB
10 GB
VRAM (MB)
8,192
10,240 +25.0%
Memory Type
GDDR6
GDDR6X
Memory Bus
256 bit
320 bit
Bandwidth
448.0 GB/s
760.3 GB/s
Cache
L1 Cache
64 KB (per SM)
128 KB (per SM)
L2 Cache
4 MB
5 MB
Performance
Pixel Rate
105.6 GPixel/s
136.8 GPixel/s
Texture Rate
237.6 GTexel/s
342.0 GTexel/s
FP32 (TFLOPS)
7.603 TFLOPS
21.89 TFLOPS
FP64 (TFLOPS)
237.6 GFLOPS (1:32)
342.0 GFLOPS (1:64)
FP16 (TFLOPS)
15.21 TFLOPS (2:1)
21.89 TFLOPS (1:1)
AI/RT
RT Cores
36
50 +38.9%
Tensor Cores
288
200 -30.6%
Power
TDP
185 W
320 W
TDP (W)
185
320 +73.0%
Suggested PSU
450 W
700 W
Power Connectors
1x 8-pin
2x 8-pin
Architecture
Architecture
Turing
Ampere
GPU Name
TU106
GA102
Generation
Mining GPUs
Mining GPUs
Process Size
12 nm
8 nm
Transistors
10,800 million
28,300 million
Die Size
445 mm²
628 mm²
Foundry
TSMC
Samsung
Density
24.3M / mm²
45.1M / mm²
API Support
DirectX
12 Ultimate (12_2)
12 Ultimate (12_2)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
3.0
3.0
CUDA
7.5
8.6
Shader Model
6.8
6.8
Physical
Slot Width
Dual-slot
Dual-slot
Length
229 mm 9 inches
285 mm 11.2 inches
Height
111 mm 4.4 inches
112 mm 4.4 inches
Outputs
No outputs
No outputs
Bus Interface
PCIe 1.0 x4
PCIe 1.0 x4
Other
Launch Price
699 USD
Production
End-of-life
End-of-life
View CMP 40HX Details View CMP 90HX Details