NVIDIA CMP 40HX vs NVIDIA GeForce RTX 3090 Ti Comparison

NVIDIA
GEFORCE

NVIDIA CMP 40HX

CORE STATE TU106
VRAM 8 GB
CLOCK SPEED 1650 MHz
TDP 185 W
BUS WIDTH 256 bit
ARCHITECTURE Turing
nm
PROCESS 12 nm
LAUNCH DATE 2021
VS
NVIDIA
GEFORCE

GeForce RTX 3090 Ti

CORE STATE GA102
VRAM 24 GB
CLOCK SPEED 1860 MHz
TDP 450 W
BUS WIDTH 384 bit
ARCHITECTURE Ampere
nm
PROCESS 8 nm
LAUNCH DATE 2022

PERFORMANCE BENCHMARKS

geekbench_opencl
93,395
174,441
geekbench_vulkan
77,879
215,633
3dmark_3dmark_steel_nomad_dx12
N/A
5,741

Analysis: NVIDIA CMP 40HX vs NVIDIA GeForce RTX 3090 Ti

Where Each One Wins

The recorded benchmark data splits cleanly between these two NVIDIA cards, and the split is not subtle. The GeForce RTX 3090 Ti wins every head-to-head benchmark in the database, and it does so by margins that are difficult to overlook. The CMP 40HX, by contrast, does not register a single win in the shared tests. That does not mean the CMP 40HX is without purpose, but its purpose is clearly not measured by the two tests both cards took.

The RTX 3090 Ti dominates in compute-heavy workloads that stress raw throughput. Its Geekbench OpenCL score of 174,441 positions it as a high-end general compute card, while its Geekbench Vulkan score of 215,633 shows strength in graphics-API workloads. The CMP 40HX trails in both, with Geekbench OpenCL at 93,395 and Geekbench Vulkan at 77,879. The Vulkan gap is especially stark: the RTX 3090 Ti scores 176.9% higher, which means it more than doubles the CMP 40HX in that test.

The CMP 40HX has a different identity. It belongs to the "Mining GPUs" generation, has no display outputs, and uses a PCIe 1.0 x4 bus interface. Those traits point toward a card designed for a narrow, non-visual workload: compute tasks that do not require a monitor connection. The database shows its average benchmark score at 85,637, which places it in the 93rd percentile of all GPUs. That percentile is respectable, but it is two points below the RTX 3090 Ti's 95th percentile.

The use-case split is therefore clear. The RTX 3090 Ti is the general-purpose powerhouse, capable of graphics output (1x HDMI 2.1, 3x DisplayPort 1.4a) and top-tier compute. The CMP 40HX is a specialized compute-only card with no display outputs, a lower memory ceiling, and a much lower performance envelope. If the workload is rendering, gaming, or any task requiring a framebuffer, the CMP 40HX cannot participate. If the workload is pure compute without visual output, the CMP 40HX exists, but the data still shows the RTX 3090 Ti outperforming it by a wide margin.

FAQ

Q: Which card has the higher average benchmark score?

A: The NVIDIA GeForce RTX 3090 Ti scores 131,938 on average, while the NVIDIA CMP 40HX scores 85,637. That is a difference of roughly 54%, favoring the RTX 3090 Ti.

Q: How large is the Vulkan performance gap between the two?

A: In Geekbench Vulkan, the RTX 3090 Ti scores 215,633 versus the CMP 40HX's 77,879. The database records a delta of 176.9%, meaning the RTX 3090 Ti is nearly three times faster in that test.

Q: Does the CMP 40HX have any benchmark win over the RTX 3090 Ti?

A: No. Across the two shared head-to-head benchmarks (Geekbench OpenCL and Geekbench Vulkan), the CMP 40HX wins zero tests. The RTX 3090 Ti wins both.

Q: What memory configurations do the two cards use?

A: The RTX 3090 Ti has 24 GB of GDDR6X on a 384-bit bus, yielding 1.01 TB/s of bandwidth. The CMP 40HX has 8 GB of GDDR6 on a 256-bit bus, yielding 448.0 GB/s of bandwidth.

Q: Can the CMP 40HX be used for display output?

A: No. The CMP 40HX lists "No outputs" for its display connections. The RTX 3090 Ti, by contrast, provides 1x HDMI 2.1 and 3x DisplayPort 1.4a.

Q: How do the two cards compare in transistor density?

A: The RTX 3090 Ti packs 28,300 million transistors into a 628 mm² die, for a density of 45.1M per mm². The CMP 40HX has 10,800 million transistors on a 445 mm² die, for a density of 24.3M per mm². The RTX 3090 Ti is denser by almost 21M transistors per square millimeter.

Head-to-Head Benchmarks

The shared benchmark suite consists of exactly two tests, and the RTX 3090 Ti wins both. The first is Geekbench OpenCL. Here the RTX 3090 Ti scores 174,441, while the CMP 40HX scores 93,395. The delta is 86.8%, meaning the RTX 3090 Ti is 86.8% faster in OpenCL compute. That is not a marginal lead; it is a near-doubling of performance.

The second test is Geekbench Vulkan, and the margin grows. The RTX 3090 Ti hits 215,633, while the CMP 40HX manages only 77,879. The delta of 176.9% means the RTX 3090 Ti more than doubles the CMP 40HX's Vulkan score. This is the single largest win in the head-to-head data, and it highlights the architectural gulf between the two.

Looking at the broader context, the RTX 3090 Ti's average score of 131,938 sits between two close rivals: the NVIDIA L4 at 131,072 (0.7% behind) and the NVIDIA RTX 4000 Ada Generation at 135,218 (2.4% ahead). The CMP 40HX's average of 85,637 is closest to the AMD Radeon PRO W7600 at 87,108 (1.7% behind) and the NVIDIA Quadro GP100 at 87,445 (2.1% behind). So while the RTX 3090 Ti competes in the upper tier of workstation-class cards, the CMP 40HX lands in a lower tier, competing with mid-range professional GPUs.

The takeaway from the head-to-head numbers is straightforward: the RTX 3090 Ti is not just slightly better, it is categorically faster in both recorded workloads. The CMP 40HX's best showing is still 86.8% behind, and its worst is 176.9% behind. No amount of workload tuning can close that gap within the measured tests.

Specification Differences

The two cards differ on nearly every specification that matters. The RTX 3090 Ti uses the GA102 chip on an 8 nm Samsung process, while the CMP 40HX uses the TU106 chip on a 12 nm TSMC process. The RTX 3090 Ti has 28,300 million transistors on a 628 mm² die; the CMP 40HX has 10,800 million transistors on a 445 mm² die.

Clocks: the RTX 3090 Ti runs at a base of 1560 MHz and boosts to 1860 MHz, with memory at 1313 MHz (21 Gbps effective). The CMP 40HX runs at 1470 MHz base and 1650 MHz boost, with memory at 1750 MHz (14 Gbps effective). The RTX 3090 Ti has higher base and boost clocks, plus faster memory.

Memory capacity is a major split: 24 GB of GDDR6X on a 384-bit bus versus 8 GB of GDDR6 on a 256-bit bus. Bandwidth follows: 1.01 TB/s for the RTX 3090 Ti, 448.0 GB/s for the CMP 40HX. The RTX 3090 Ti offers more than double the bandwidth.

Compute resources diverge sharply. The RTX 3090 Ti has 10,752 shading units, 336 TMUs, 112 ROPs, 84 RT cores, and 336 tensor cores. The CMP 40HX has 2,304 shading units, 144 TMUs, 64 ROPs, 36 RT cores, and 288 tensor cores. Pixel rate: 208.3 GPixel/s versus 105.6 GPixel/s. Texture rate: 625.0 GTexel/s versus 237.6 GTexel/s. FP32 throughput: 40.00 TFLOPS versus 7.603 TFLOPS. The RTX 3090 Ti is more than five times faster in FP32.

Power and physical specs also differ. The RTX 3090 Ti has a 450 W TDP, requires a 850 W suggested PSU, uses a triple-slot cooler, and needs a 1x 16-pin power connector. The CMP 40HX has a 185 W TDP, a 450 W suggested PSU, a dual-slot cooler, and a 1x 8-pin connector. The RTX 3090 Ti is 336 mm long, 140 mm tall, and 61 mm wide; the CMP 40HX is 229 mm long, 111 mm tall, and 35 mm wide. The RTX 3090 Ti is physically much larger.

Bus interface is another differentiator: the RTX 3090 Ti uses PCIe 4.0 x16, while the CMP 40HX uses PCIe 1.0 x4. The latter is a severe bandwidth limitation for a card that still supports DirectX 12 Ultimate, OpenGL 4.6, and Vulkan 1.4. Both cards share the same API support, but the CMP 40HX has no display outputs.

Architecture Differences

The architectural gap is fundamental. The RTX 3090 Ti is built on NVIDIA's Ampere architecture, designed for the GeForce 30 series. The CMP 40HX is built on the older Turing architecture, from the "Mining GPUs" generation. Ampere targets general-purpose graphics and compute with a focus on ray tracing and tensor operations; Turing is an earlier design that also includes RT and tensor cores, but with far lower counts.

The RTX 3090 Ti uses an 8 nm Samsung process, while the CMP 40HX uses a 12 nm TSMC process. The process node difference explains part of the transistor density gap: 45.1M per mm² versus 24.3M per mm². More transistors on a smaller node gives the RTX 3090 Ti both higher raw counts and tighter packing.

Core composition differs beyond simple counts. The RTX 3090 Ti has 84 RT cores and 336 tensor cores, while the CMP 40HX has 36 RT cores and 288 tensor cores. The RTX 3090 Ti has more than double the RT cores, and 48 more tensor cores. The shading unit count is the largest gap: 10,752 versus 2,304, a factor of 4.7.

Memory architecture is also distinct. GDDR6X on a 384-bit bus (RTX 3090 Ti) versus GDDR6 on a 256-bit bus (CMP 40HX). GDDR6X offers higher data rates, and the wider bus multiplies the advantage. The result is 1.01 TB/s versus 448.0 GB/s, a 2.25x bandwidth lead.

The CMP 40HX's PCIe 1.0 x4 interface is a architectural oddity for a card with 12 Ultimate API support. It suggests the card was designed for compute tasks that transfer data in bulk rather than interactively. The RTX 3090 Ti's PCIe 4.0 x16 interface is standard for a high-end graphics card. The lack of display outputs on the CMP 40HX reinforces its mining-only design, while the RTX 3090 Ti's full display suite (1x HDMI 2.1, 3x DisplayPort 1.4a) makes it a complete workstation card.

The Verdict

The data points to a clear choice for most users. The RTX 3090 Ti wins both head-to-head benchmarks, has a higher average score (131,938 versus 85,637), sits in the 95th percentile versus the 93rd, and offers more than double the memory bandwidth. It also has display outputs, which the CMP 40HX lacks entirely. Anyone needing a card for rendering, gaming, or general compute should choose the RTX 3090 Ti without hesitation.

The CMP 40HX is a different product for a different purpose. It was designed for mining, not for interactive use. Its PCIe 1.0 x4 interface and lack of display outputs make it unsuitable for any task that requires a monitor. Its 8 GB of memory and 448.0 GB/s bandwidth are enough for some compute workloads, but the RTX 3090 Ti outperforms it in every recorded test. The CMP 40HX's average score of 85,637 places it near the AMD Radeon PRO W7600 (87,108) and NVIDIA Quadro GP100 (87,445), which are mid-range professional cards. That is its competitive tier.

The verdict is simple: the RTX 3090 Ti is the superior card by every measured metric. The CMP 40HX is a niche product that cannot match the RTX 3090 Ti in any benchmark the database recorded. If the workload is compute-only and power efficiency is a priority, the CMP 40HX's 185 W TDP versus 450 W TDP might matter. But the performance gap of 86.8% to 176.9% is too large to ignore. The RTX 3090 Ti is the pick for performance; the CMP 40HX is only relevant for a specific mining use case where its lower power draw and smaller size (229 mm versus 336 mm) could be an advantage.

DETAILED SPECIFICATIONS

SPECIFICATION
CMP 40HX
RTX 3090 Ti
Core Specs
Shading Units
2,304
10,752 +366.7%
Shaders
2,304
10,752 +366.7%
TMUs
144
336 +133.3%
ROPs
64
112 +75.0%
SM Count
36
84 +133.3%
Clocks
Base Clock
1470 MHz
1560 MHz
Boost Clock
1650 MHz
1860 MHz
Memory Clock
1750 MHz 14 Gbps effective
1313 MHz 21 Gbps effective
Memory
Memory Size
8 GB
24 GB
VRAM (MB)
8,192
24,576 +200.0%
Memory Type
GDDR6
GDDR6X
Memory Bus
256 bit
384 bit
Bandwidth
448.0 GB/s
1.01 TB/s
Cache
L1 Cache
64 KB (per SM)
128 KB (per SM)
L2 Cache
4 MB
6 MB
Performance
Pixel Rate
105.6 GPixel/s
208.3 GPixel/s
Texture Rate
237.6 GTexel/s
625.0 GTexel/s
FP32 (TFLOPS)
7.603 TFLOPS
40.00 TFLOPS
FP64 (TFLOPS)
237.6 GFLOPS (1:32)
625.0 GFLOPS (1:64)
FP16 (TFLOPS)
15.21 TFLOPS (2:1)
40.00 TFLOPS (1:1)
AI/RT
RT Cores
36
84 +133.3%
Tensor Cores
288
336 +16.7%
Power
TDP
185 W
450 W
TDP (W)
185
450 +143.2%
Suggested PSU
450 W
850 W
Power Connectors
1x 8-pin
1x 16-pin
Architecture
Architecture
Turing
Ampere
GPU Name
TU106
GA102
Generation
Mining GPUs
GeForce 30
Process Size
12 nm
8 nm
Transistors
10,800 million
28,300 million
Die Size
445 mm²
628 mm²
Foundry
TSMC
Samsung
Density
24.3M / mm²
45.1M / mm²
API Support
DirectX
12 Ultimate (12_2)
12 Ultimate (12_2)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
3.0
3.0
CUDA
7.5
8.6
Shader Model
6.8
6.8
Physical
Slot Width
Dual-slot
Triple-slot
Length
229 mm 9 inches
336 mm 13.2 inches
Height
111 mm 4.4 inches
140 mm 5.5 inches
Outputs
No outputs
1x HDMI 2.13x DisplayPort 1.4a
Bus Interface
PCIe 1.0 x4
PCIe 4.0 x16
Other
Launch Price
699 USD
1,999 USD
Production
End-of-life
End-of-life
Predecessor
GeForce 20
Successor
GeForce 40
View CMP 40HX Details View GeForce RTX 3090 Ti Details