NVIDIA CMP 90HX vs NVIDIA GeForce RTX 4090 Comparison

NVIDIA
GEFORCE

NVIDIA CMP 90HX

CORE STATE GA102
VRAM 10 GB
CLOCK SPEED 1710 MHz
TDP 320 W
BUS WIDTH 320 bit
ARCHITECTURE Ampere
nm
PROCESS 8 nm
LAUNCH DATE 2021
VS
NVIDIA
GEFORCE

GeForce RTX 4090

CORE STATE AD102
VRAM 24 GB
CLOCK SPEED 2520 MHz
TDP 450 W
BUS WIDTH 384 bit
ARCHITECTURE Ada Lovelace
nm
PROCESS 5 nm
LAUNCH DATE 2022

PERFORMANCE BENCHMARKS

geekbench_opencl
69,000
255,416
3dmark_3dmark_steel_nomad_dx12
N/A
9,223
geekbench_vulkan
N/A
271,631
passmark_directx_10
N/A
224
passmark_directx_11
N/A
326
passmark_directx_12
N/A
150
passmark_directx_9
N/A
397
passmark_g2d
N/A
1,299
passmark_g3d
N/A
38,194
passmark_gpu_compute
N/A
26,613

Analysis: NVIDIA CMP 90HX vs NVIDIA GeForce RTX 4090

The data presents a stark generational contrast between two NVIDIA GPUs designed for entirely different purposes. The NVIDIA CMP 90HX, a dedicated mining part from the Ampere era, and the NVIDIA GeForce RTX 4090, a flagship consumer graphics card from the Ada Lovelace generation, share a manufacturer but little else in terms of performance DNA. The single head-to-head benchmark available, Geekbench OpenCL, shows a decisive victory for the RTX 4090, but the story is more nuanced when examining the architectural and specification chasms between them.

Head-to-Head Benchmarks

The only direct benchmark comparison in the data is the Geekbench OpenCL test. Here, the NVIDIA GeForce RTX 4090 scores a massive 255,416 points, while the NVIDIA CMP 90HX manages 69,000 points. This results in a delta of -73% for the CMP 90HX, meaning the RTX 4090 is approximately 270% faster in this compute-oriented workload. The margin is overwhelming and reflects the fundamental gap in compute resources and architecture between the two.

Looking at the broader context of the average benchmark scores, the RTX 4090's average is 60,347, which is actually lower than its Geekbench OpenCL score due to other benchmarks in its suite, such as Passmark DirectX 9 (397) and Passmark DirectX 12 (150). The CMP 90HX's average benchmark score is exactly 69,000, as that is its sole entry. The RTX 4090's percentile ranking of 88th versus the CMP 90HX's 90th percentile is an anomaly explained by the different benchmark suites each card is subjected to. The CMP 90HX sits near the top of the database because its single score is high for a mining-focused part, while the RTX 4090's percentile is pulled down by its diverse range of DirectX and compute tests.

The nearest rivals for each card further highlight their distinct performance classes. The CMP 90HX is flanked by the Intel Arc A770 (68,809, delta 0.3%) and the AMD Radeon Pro WX 8200 (69,870, delta -1.2%), showing it is a mid-to-high-tier performer in the overall database. In contrast, the RTX 4090's nearest rivals include the Intel Arc Pro A60 (60,326, delta 0%) and the AMD Radeon Pro W6600M (61,896, delta -2.5%), which are all far less powerful than the RTX 4090's raw Geekbench score suggests, indicating that its average is heavily weighted by lower DirectX scores from its other tests.

Architecture Differences

The architectural divide is the primary reason for the performance gap. The CMP 90HX is built on the GA102 chip using the Ampere architecture, fabricated on Samsung's 8 nm process node. It packs 28,300 million transistors onto a 628 mm² die, resulting in a transistor density of 45.1 million per mm². In contrast, the RTX 4090 uses the AD102 chip with the Ada Lovelace architecture, built on TSMC's 5 nm process. This newer node allows for a staggering 76,300 million transistors on a slightly smaller 609 mm² die, achieving a density of 125.3 million per mm². The move to a denser, more advanced process node is the foundational advantage for the RTX 4090.

The core configurations differ by an order of magnitude. The CMP 90HX has 6,400 shading units, 200 texture mapping units (TMUs), and 80 render output units (ROPs). Its ray tracing and tensor core counts are 50 and 200, respectively. The RTX 4090 dwarfs this with 16,384 shading units, 512 TMUs, and 176 ROPs. It also features 128 ray tracing cores and 512 tensor cores, representing a 2.56x increase in every core category. This massive parallel processing capability directly translates into higher fill rates and compute throughput.

Clock speeds and derived performance metrics reinforce the gap. The CMP 90HX has a base clock of 1500 MHz and a boost clock of 1710 MHz, producing a pixel rate of 136.8 GPixel/s and a texture rate of 342.0 GTexel/s. Its FP32 and FP16 compute are both pegged at 21.89 TFLOPS. The RTX 4090 operates at a base clock of 2235 MHz and a boost of 2520 MHz, which, combined with its larger core count, yields a pixel rate of 443.5 GPixel/s and a texture rate of 1,290.2 GTexel/s. Its FP32 and FP16 performance is 82.58 TFLOPS, which is roughly 3.8 times higher than the CMP 90HX. The clock speed advantage is notable, but the core count difference is the dominant factor.

The Verdict

From a pure performance standpoint, the data leaves no ambiguity: the NVIDIA GeForce RTX 4090 is in a completely different league. Its 82.58 TFLOPS of FP32 compute is nearly four times that of the CMP 90HX's 21.89 TFLOPS, and its 24 GB of memory with 1.01 TB/s bandwidth obliterates the CMP 90HX's 10 GB and 760.3 GB/s. For any compute, rendering, or gaming workload, the RTX 4090 is the definitive choice.

The CMP 90HX’s sole purpose, as its "Mining GPUs" generation label suggests, was not general-purpose computing. Its lack of display outputs and its unusual PCIe 1.0 x4 bus interface (a severe bottleneck for data transfer) make it unsuitable for normal use. Its 10 GB memory buffer is also small for modern workloads. The RTX 4090, with its triple-slot cooler, 1x 16-pin power connector, and 450 W TDP, is a power-hungry but complete consumer product. The CMP 90HX, with a 320 W TDP and dual-slot design, is simpler but crippled for anything other than its mining function.

FAQ

Q: Which GPU is faster in the Geekbench OpenCL benchmark?

A: The NVIDIA GeForce RTX 4090 is significantly faster, scoring 255,416 compared to the NVIDIA CMP 90HX's 69,000, a delta of -73% for the CMP 90HX.

Q: How does the memory bandwidth compare between the two?

A: The RTX 4090 offers 1.01 TB/s of bandwidth, while the CMP 90HX provides 760.3 GB/s. The RTX 4090 also has a larger 384-bit memory bus versus the CMP 90HX's 320-bit bus.

Q: What are the key architectural differences?

A: The CMP 90HX uses the Ampere architecture on an 8 nm Samsung process, while the RTX 4090 uses the Ada Lovelace architecture on a 5 nm TSMC process. The RTX 4090 also has significantly more cores: 16,384 shading units versus 6,400, and 128 RT cores versus 50.

Q: Do both cards support the same APIs?

A: Yes, both support DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4. However, the RTX 4090 has display outputs (1x HDMI 2.1 and 3x DisplayPort 1.4a), while the CMP 90HX has none.

Q: What is the power requirement difference?

A: The RTX 4090 has a TDP of 450 W and requires a 850 W PSU, while the CMP 90HX has a lower 320 W TDP and a suggested 700 W PSU. The RTX 4090 uses a single 16-pin connector, whereas the CMP 90HX uses two 8-pin connectors.

Q: How do their percentile rankings compare?

A: The CMP 90HX ranks in the 90th percentile of all GPUs, while the RTX 4090 ranks in the 88th percentile. This is due to the different benchmark suites used; the CMP 90HX's single score is high, but the RTX 4090's average is lowered by its DirectX 9 and 12 scores.

Where Each One Wins

The NVIDIA GeForce RTX 4090 wins in every conceivable performance category measured. It is the clear victor in raw compute (82.58 vs 21.89 TFLOPS FP32), memory capacity (24 GB vs 10 GB), and memory bandwidth (1.01 TB/s vs 760.3 GB/s). Its higher pixel rate (443.5 GPixel/s vs 136.8 GPixel/s) and texture rate (1,290.2 GTexel/s vs 342.0 GTexel/s) make it superior for any graphically intensive application, from gaming to professional 3D rendering. Its connectivity, with PCIe 4.0 x16 and multiple display outputs, makes it a functional, complete product.

The NVIDIA CMP 90HX has no benchmark wins in the data. Its only theoretical advantage is its lower power consumption (320 W vs 450 W) and a smaller physical footprint (285 mm length vs 304 mm). However, its lack of display outputs and its crippled PCIe 1.0 x4 interface mean it cannot be used for standard tasks. It is a specialized tool for a single purpose—mining—where its 10 GB memory and 320-bit bus are sufficient. Its sole benchmark score of 69,000 does place it in the 90th percentile of all GPUs, but this is an artifact of the limited testing and does not reflect real-world versatility.

Specification Differences

The specification sheet reveals a total divergence in design philosophy. The process node differs from Samsung's 8 nm to TSMC's 5 nm, and the transistor count more than doubles from 28,300 million to 76,300 million. The core architecture changes from GA102 to AD102, and every core count jumps: shading units from 6,400 to 16,384, TMUs from 200 to 512, ROPs from 80 to 176, RT cores from 50 to 128, and tensor cores from 200 to 512. The memory subsystem is also different, with the CMP 90HX using a 320-bit bus for 10 GB, while the RTX 4090 uses a 384-bit bus for 24 GB. Clock speeds rise from a 1500/1710 MHz base/boost to 2235/2520 MHz.

Power and physical specifications follow suit. The CMP 90HX is a dual-slot card with a 320 W TDP and two 8-pin power connectors, while the RTX 4090 is a triple-slot card with a 450 W TDP and a single 16-pin connector. The suggested PSU increases from 700 W to 850 W. The bus interface is a major differentiator: the CMP 90HX is limited to PCIe 1.0 x4, whereas the RTX 4090 uses the modern PCIe 4.0 x16. Most importantly, the CMP 90HX has no display outputs, while the RTX 4090 includes 1x HDMI 2.1 and 3x DisplayPort 1.4a. The dimensions also differ, with the CMP 90HX at 285 mm length and 112 mm height, and the RTX 4090 at 304 mm length, 137 mm height, and 61 mm width. The RTX 4090 has a launch MSRP of 1,599 USD.

DETAILED SPECIFICATIONS

SPECIFICATION
CMP 90HX
RTX 4090
Core Specs
Shading Units
6,400
16,384 +156.0%
Shaders
6,400
16,384 +156.0%
TMUs
200
512 +156.0%
ROPs
80
176 +120.0%
SM Count
50
128 +156.0%
Clocks
Base Clock
1500 MHz
2235 MHz
Boost Clock
1710 MHz
2520 MHz
Memory Clock
1188 MHz 19 Gbps effective
1313 MHz 21 Gbps effective
Memory
Memory Size
10 GB
24 GB
VRAM (MB)
10,240
24,576 +140.0%
Memory Type
GDDR6X
GDDR6X
Memory Bus
320 bit
384 bit
Bandwidth
760.3 GB/s
1.01 TB/s
Cache
L1 Cache
128 KB (per SM)
128 KB (per SM)
L2 Cache
5 MB
72 MB
Performance
Pixel Rate
136.8 GPixel/s
443.5 GPixel/s
Texture Rate
342.0 GTexel/s
1,290.2 GTexel/s
FP32 (TFLOPS)
21.89 TFLOPS
82.58 TFLOPS
FP64 (TFLOPS)
342.0 GFLOPS (1:64)
1,290.2 GFLOPS (1:64)
FP16 (TFLOPS)
21.89 TFLOPS (1:1)
82.58 TFLOPS (1:1)
AI/RT
RT Cores
50
128 +156.0%
Tensor Cores
200
512 +156.0%
Power
TDP
320 W
450 W
TDP (W)
320
450 +40.6%
Suggested PSU
700 W
850 W
Power Connectors
2x 8-pin
1x 16-pin
Architecture
Architecture
Ampere
Ada Lovelace
GPU Name
GA102
AD102
Generation
Mining GPUs
GeForce 40
Process Size
8 nm
5 nm
Transistors
28,300 million
76,300 million
Die Size
628 mm²
609 mm²
Foundry
Samsung
TSMC
Density
45.1M / mm²
125.3M / mm²
API Support
DirectX
12 Ultimate (12_2)
12 Ultimate (12_2)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
3.0
3.0
CUDA
8.6
8.9
Shader Model
6.8
6.8
Physical
Slot Width
Dual-slot
Triple-slot
Length
285 mm 11.2 inches
304 mm 12 inches
Height
112 mm 4.4 inches
137 mm 5.4 inches
Outputs
No outputs
1x HDMI 2.13x DisplayPort 1.4a
Bus Interface
PCIe 1.0 x4
PCIe 4.0 x16
Other
Launch Price
1,599 USD
Production
End-of-life
End-of-life
Predecessor
GeForce 30
Successor
GeForce 50
View CMP 90HX Details View GeForce RTX 4090 Details