NVIDIA CMP 40HX vs NVIDIA CMP 90HX Comparison
NVIDIA CMP 40HX
CMP 90HX
PERFORMANCE BENCHMARKS
Analysis: NVIDIA CMP 40HX vs NVIDIA CMP 90HX
The NVIDIA CMP 40HX and NVIDIA CMP 90HX are both end-of-life mining-oriented GPUs with no display outputs, but they serve very different performance tiers. Based on the available benchmark data, the CMP 40HX is the clear winner in the only head-to-head test, delivering a 35.4% higher Geekbench OpenCL score than the CMP 90HX. This result is counterintuitive given the CMP 90HX’s much larger chip and newer architecture, so the data demands a closer look at where each card actually excels.
Where Each One Wins
The CMP 40HX wins the only directly comparable benchmark, Geekbench OpenCL, with a score of 93395 versus 69000 for the CMP 90HX. That is a decisive 35.4% margin, placing the 40HX in the 93rd percentile of all GPUs, while the 90HX sits at the 90th percentile. The 40HX also has a second benchmark result — a Geekbench Vulkan score of 77879 — which the 90HX lacks entirely, giving the 40HX a broader software compatibility profile in the data.
However, the CMP 90HX is not without its own strengths in the specification sheet. Its raw compute resources are substantially higher: 6400 shading units versus 2304, 200 texture mapping units versus 144, and 80 render output units versus 64. It also carries 50 RT cores and 200 tensor cores, compared to 36 RT cores and 288 tensor cores on the 40HX. These figures suggest the 90HX should dominate in compute-heavy workloads that scale with shading unit count, even though the benchmark data does not capture that advantage in the OpenCL test.
The 90HX also wins on memory capacity and bandwidth, with 10 GB of GDDR6X on a 320-bit bus delivering 760.3 GB/s, versus 8 GB of GDDR6 on a 256-bit bus at 448.0 GB/s. For workloads that are memory-bound rather than compute-bound, the 90HX’s 69.7% bandwidth advantage could translate into real-world wins that the single OpenCL score does not reflect. The 40HX, by contrast, wins on efficiency per watt in the data: it achieves its higher benchmark score at 185 W TDP, while the 90HX requires 320 W.
Architecture Differences
The two cards come from different NVIDIA architectures and foundries. The CMP 40HX uses the TU106 chip built on TSMC’s 12 nm process, while the CMP 90HX uses the GA102 chip on Samsung’s 8 nm node. This is a generational leap: the 90HX packs 28,300 million transistors into a 628 mm² die, versus 10,800 million transistors on a 445 mm² die for the 40HX. Transistor density tells the story of the node advantage — the 90HX achieves 45.1 million transistors per mm², nearly double the 40HX’s 24.3 million.
The architecture shift from Turing to Ampere brings a fundamental change in compute precision handling. The 40HX’s FP16 throughput is listed as 15.21 TFLOPS with a 2:1 ratio relative to FP32, meaning it halves FP32 throughput when doing FP16 work. The 90HX, in contrast, delivers 21.89 TFLOPS for both FP16 and FP32 with a 1:1 ratio, indicating it can do full-rate FP16 without sacrificing FP32 performance. This makes the 90HX more flexible for mixed-precision workloads.
Memory technology also diverges: the 40HX uses GDDR6 at 14 Gbps effective, while the 90HX uses GDDR6X at 19 Gbps effective. The 90HX’s memory clock is listed at 1188 MHz base, translating to the higher effective data rate. Both cards share the same PCIe 1.0 x4 bus interface, which is an unusual and restrictive choice for mining cards, and both have no display outputs. The 40HX is physically smaller at 229 mm in length versus 285 mm for the 90HX, and the 40HX requires a single 8-pin power connector while the 90HX needs two.
Head-to-Head Benchmarks
The single head-to-head benchmark is Geekbench OpenCL, and the result is emphatic: the CMP 40HX scores 93395, beating the CMP 90HX’s 69000 by 35.4%. This is a substantial margin that flips the expected hierarchy based on specs. The 40HX’s average benchmark score across its two tests is 85637, while the 90HX’s average is 69000 — a 24.1% gap in favor of the 40HX.
Looking at rival comparisons, the 40HX sits just 1.7% below the AMD Radeon PRO W7600 and 2.1% below the NVIDIA Quadro GP100, while leading the AMD Radeon PRO W6600 by 4.4% and the AMD Radeon Pro Vega 64X by 5.8%. The 90HX, on the other hand, is nearly tied with the Intel Arc A770 (0.3% ahead) and the AMD Radeon Instinct MI25 (0.6% ahead), while trailing the AMD Radeon Pro WX 8200 by 1.2% and the NVIDIA Quadro P6000 by 1.4%. These rival deltas show that the 40HX competes in a higher performance tier than the 90HX in OpenCL, despite the 90HX’s newer architecture.
The 40HX also has a Geekbench Vulkan score of 77879, which is 16.6% below its own OpenCL score. This suggests the 40HX performs better in OpenCL than Vulkan, though both results are strong. The 90HX has no Vulkan benchmark in the data, so no comparison is possible for that API. The overall wins tally is 1 for the 40HX and 0 for the 90HX, but this is based on a single shared test, which limits the conclusiveness of the comparison.
FAQ
Q: Which card has the higher benchmark score?
A: The NVIDIA CMP 40HX scores 93395 in Geekbench OpenCL, while the CMP 90HX scores 69000, giving the 40HX a 35.4% advantage.
Q: Does the CMP 90HX have any benchmark where it wins?
A: No. In the only head-to-head benchmark (Geekbench OpenCL), the CMP 40HX wins. The 90HX has no other benchmark results to compare.
Q: What are the memory differences between the two cards?
A: The CMP 90HX has 10 GB of GDDR6X on a 320-bit bus with 760.3 GB/s bandwidth. The CMP 40HX has 8 GB of GDDR6 on a 256-bit bus with 448.0 GB/s bandwidth.
Q: Which card has more shading units?
A: The CMP 90HX has 6400 shading units, compared to 2304 on the CMP 40HX. The 90HX also has more TMUs (200 vs 144) and ROPs (80 vs 64).
Q: Are these cards different in power requirements?
A: Yes. The CMP 40HX has a 185 W TDP and uses a single 8-pin connector with a 450 W suggested PSU. The CMP 90HX has a 320 W TDP, uses two 8-pin connectors, and requires a 700 W suggested PSU.
Q: What is the transistor count difference?
A: The CMP 90HX has 28,300 million transistors on a 628 mm² die, while the CMP 40HX has 10,800 million transistors on a 445 mm² die. The 90HX’s 8 nm process yields 45.1M transistors per mm² versus 24.3M for the 40HX’s 12 nm process.
The Verdict
The data is unambiguous in one respect: if you are choosing based purely on the available benchmark results, the NVIDIA CMP 40HX is the superior card. It wins the only head-to-head test by 35.4%, holds a higher percentile rank (93rd vs 90th), and has additional benchmark coverage with its Vulkan score. The 40HX also achieves this at a lower TDP of 185 W versus 320 W, making it the more efficient choice in the data.
However, the CMP 90HX should not be dismissed outright. Its specification sheet reveals a card with 2.8 times the shading units, 69.7% more memory bandwidth, and double the transistor count. The 90HX’s 1:1 FP16/FP32 ratio and full-rate FP16 performance at 21.89 TFLOPS indicate it is built for compute workloads that the benchmark data does not capture. If your workload scales with shading unit count or memory bandwidth rather than the specific OpenCL test used here, the 90HX could be the stronger performer.
The practical choice depends on which metric matters more. For a straightforward compute benchmark, the CMP 40HX is the verdict — it simply outperforms the 90HX in the test that matters. For raw theoretical compute and memory resources, the CMP 90HX is the more capable silicon, but the data does not confirm that advantage in practice. Buyers should note that both cards are end-of-life, have no display outputs, and use a restrictive PCIe 1.0 x4 interface, which limits their utility beyond mining or compute tasks.
Specification Differences
| Specification | NVIDIA CMP 40HX | NVIDIA CMP 90HX |
|---|---|---|
| Chip | TU106 | GA102 |
| Architecture | Turing | Ampere |
| Process Node | 12 nm | 8 nm |
| Foundry | TSMC | Samsung |
| Transistors | 10,800 million | 28,300 million |
| Die Size | 445 mm² | 628 mm² |
| Transistor Density | 24.3M / mm² | 45.1M / mm² |
| Base Clock | 1470 MHz | 1500 MHz |
| Boost Clock | 1650 MHz | 1710 MHz |
| Memory Clock | 1750 MHz (14 Gbps effective) | 1188 MHz (19 Gbps effective) |
| Memory Size | 8 GB | 10 GB |
| Memory Type | GDDR6 | GDDR6X |
| Memory Bus Width | 256 bit | 320 bit |
| Memory Bandwidth | 448.0 GB/s | 760.3 GB/s |
| Shading Units | 2304 | 6400 |
| TMUs | 144 | 200 |
| ROPs | 64 | 80 |
| RT Cores | 36 | 50 |
| Tensor Cores | 288 | 200 |
| Pixel Rate | 105.6 GPixel/s | 136.8 GPixel/s |
| Texture Rate | 237.6 GTexel/s | 342.0 GTexel/s |
| FP32 Performance | 7.603 TFLOPS | 21.89 TFLOPS |
| FP16 Performance | 15.21 TFLOPS (2:1) | 21.89 TFLOPS (1:1) |
| TDP | 185 W | 320 W |
| Power Connectors | 1x 8-pin | 2x 8-pin |
| Suggested PSU | 450 W | 700 W |
| Length | 229 mm (9 inches) | 285 mm (11.2 inches) |
| Height | 111 mm (4.4 inches) | 112 mm (4.4 inches) |
| Width | 35 mm (1.4 inches) | Not specified |
| Launch MSRP | 699 USD | Not specified |
| Release Date | 2021-02-24 | 2021-07-27 |