NVIDIA CMP 40HX vs NVIDIA GeForce RTX 5090 Comparison
NVIDIA CMP 40HX
GeForce RTX 5090
PERFORMANCE BENCHMARKS
Analysis: NVIDIA CMP 40HX vs NVIDIA GeForce RTX 5090
NVIDIA CMP 40HX and NVIDIA GeForce RTX 5090 sit in the 94th percentile of all GPUs, yet their average benchmark scores tell a surprising story—the CMP 40HX edges out the RTX 5090 by 1.6% in aggregate, despite the latter being a flagship Blackwell part. The data shows a mining-focused Turing card from 2021 outscoring a 2025 consumer behemoth in the two shared compute tests, which raises immediate questions about what these benchmarks actually measure and where each card's architectural strengths truly lie.
The Verdict
The RTX 5090 is the clear winner for any workload involving modern graphics APIs, given its massive 307% lead in Geekbench OpenCL and 387% lead in Geekbench Vulkan over the CMP 40HX. The data indicates the RTX 5090 delivers 380,114 points in OpenCL against 93,395 for the CMP 40HX, a -75.4% delta, and 379,571 versus 77,879 in Vulkan, a -79.5% delta. If the task involves rendering, gaming, or any DirectX 12 Ultimate workload, the RTX 5090 is the only rational choice from these two.
However, the aggregate benchmark picture is more nuanced. The CMP 40HX averages 85,637 points across its two tests, while the RTX 5090 averages 84,306 across ten tests, giving the older card a 1.6% edge in the nearestRivals comparison. This suggests the CMP 40HX's compute-heavy design—with its 288 tensor cores and 36 RT cores in a mining-focused package—can outperform the RTX 5090 in certain raw compute scenarios, particularly those that don't leverage the newer architecture's full feature set. The RTX 5090's PassMark scores (39,650 in G3D, 26,756 in compute) show strength in DirectX 10, 11, and 12 paths, but these tests don't appear in the CMP 40HX's benchmark list, making direct comparison impossible.
For a buyer strictly using these two options, the RTX 5090 is the pick for any display-connected workstation or gaming rig, while the CMP 40HX—with no display outputs—only makes sense for headless compute farms where its lower power draw and surprising aggregate score matter more than raw graphics throughput.
Architecture Differences
The two cards are generations apart in every meaningful architectural dimension. The CMP 40HX uses the TU106 chip on TSMC's 12 nm process, packing 10,800 million transistors into a 445 mm² die with a transistor density of 24.3M per mm². The RTX 5090 uses GB202 on a 5 nm node, with 92,200 million transistors across 750 mm², achieving 122.9M transistors per mm²—a 5x density improvement that fundamentally changes what can fit on a chip.
Clock speeds reflect this generational leap. The CMP 40HX runs at 1470 MHz base and 1650 MHz boost, while the RTX 5090 operates at 2017 MHz base and 2407 MHz boost. The newer card's 1:1 FP16 ratio (104.8 TFLOPS) versus the CMP 40HX's 2:1 ratio (15.21 TFLOPS FP16) indicates the RTX 5090 handles half-precision workloads at full rate, a critical difference for AI inference. Shading units tell the scale: 21,760 on the RTX 5090 versus 2,304 on the CMP 40HX, with 680 TMUs and 176 ROPs compared to 144 and 64 respectively.
Memory architecture diverges completely. The CMP 40HX uses 8 GB of GDDR6 on a 256-bit bus at 448.0 GB/s, while the RTX 5090 uses 32 GB of GDDR7 on a 512-bit bus at 1.79 TB/s—4x the capacity and 4x the bandwidth. The RTX 5090 also supports PCIe 5.0 x16 versus the CMP 40HX's PCIe 1.0 x4 interface, which severely limits data transfer for the older card. Display outputs exist only on the RTX 5090 (1x HDMI 2.1b, 3x DisplayPort 2.1b); the CMP 40HX has none, confirming its mining-only design intent.
Where Each One Wins
The RTX 5090 wins decisively in every shared benchmark, but the data suggests distinct use-case advantages. For OpenCL compute—which often reflects general-purpose GPU workloads like physics simulation, image processing, or scientific computing—the RTX 5090's 380,114 score versus 93,395 shows a 307% advantage. This aligns with its 104.8 TFLOPS FP32 throughput versus 7.603 TFLOPS on the CMP 40HX, a 13.8x raw compute difference that the benchmark confirms.
Vulkan performance follows the same pattern: 379,571 for the RTX 5090 against 77,879 for the CMP 40HX, a 387% gap. This is where the RTX 5090's 170 RT cores and 680 tensor cores outperform the CMP 40HX's 36 RT cores and 288 tensor cores in modern rendering paths. Any workload that uses Vulkan—from games to CAD visualization—belongs to the RTX 5090.
The CMP 40HX's aggregate win (85,637 average versus 84,306) stems from its two-test benchmark suite being narrower and possibly more favorable to its Turing compute architecture. The card's 1.6% delta over the RTX 5090 in nearestRivals suggests that in specific, non-graphics compute tasks—particularly those that scale well with tensor core count relative to power draw—the older card can hold its own. Its 185 W TDP versus 575 W means it achieves comparable aggregate scores at less than one-third the power, though the RTX 5090's additional benchmark coverage makes direct efficiency claims unreliable.
FAQ
Q: Which card has better raw compute performance?
A: The RTX 5090 dominates in both shared benchmarks, with 380,114 versus 93,395 in OpenCL and 379,571 versus 77,879 in Vulkan, representing 307% and 387% advantages respectively.
Q: Why does the CMP 40HX have a higher average benchmark score than the RTX 5090?
A: The CMP 40HX averages 85,637 across its two benchmarks, while the RTX 5090 averages 84,306 across ten. The RTX 5090's additional tests (PassMark DirectX 10, 11, 12, 9, G2D, G3D, and GPU Compute) drag its average down, despite winning the two tests both cards share.
Q: Can the CMP 40HX be used for gaming or display output?
A: No. The CMP 40HX has no display outputs, making it unsuitable for any visual output. The RTX 5090 offers 1x HDMI 2.1b and 3x DisplayPort 2.1b for display connectivity.
Q: What are the memory capacity and bandwidth differences?
A: The RTX 5090 has 32 GB of GDDR7 on a 512-bit bus with 1.79 TB/s bandwidth, while the CMP 40HX has 8 GB of GDDR6 on a 256-bit bus with 448.0 GB/s—the RTX 5090 offers 4x capacity and 4x bandwidth.
Q: How do the power requirements compare?
A: The CMP 40HX has a 185 W TDP with a 450 W suggested PSU and 1x 8-pin connector, while the RTX 5090 has a 575 W TDP with a 950 W suggested PSU and 1x 16-pin connector.
Q: Which card has better tensor core performance?
A: The RTX 5090 has 680 tensor cores versus 288 on the CMP 40HX, and its FP16 throughput is 104.8 TFLOPS (1:1 ratio) compared to 15.21 TFLOPS (2:1 ratio) on the CMP 40HX.
Head-to-Head Benchmarks
The two shared benchmarks produce lopsided results that reveal the architectural gulf between these cards. In Geekbench OpenCL, the RTX 5090 scores 380,114 against the CMP 40HX's 93,395, a delta of -75.4% from the CMP 40HX's perspective. This 286,719-point gap reflects the RTX 5090's 21760 shading units and 104.8 TFLOPS FP32, compared to 2,304 shading units and 7.603 TFLOPS on the CMP 40HX. The OpenCL test likely stresses raw ALU throughput, where the newer card's 5 nm node and 122.9M transistors per mm² density provide an insurmountable advantage.
Geekbench Vulkan shows an even wider relative gap: 379,571 for the RTX 5090 versus 77,879 for the CMP 40HX, a -79.5% delta. This benchmark exercises graphics pipeline features like ray tracing and compute shaders, where the RTX 5090's 170 RT cores and 680 tensor cores dwarf the CMP 40HX's 36 RT cores and 288 tensor cores. The pixel rate difference (423.6 GPixel/s versus 105.6 GPixel/s) and texture rate difference (1,636.8 GTexel/s versus 237.6 GTexel/s) further explain why Vulkan performance diverges so sharply.
The RTX 5090's additional PassMark scores—39,650 in G3D, 26,756 in GPU Compute, 1,413 in G2D—provide context for its aggregate 84,306 average. These tests show strong DirectX 12 performance (185) and DirectX 11 (341), but their inclusion in the average dilutes the card's score relative to the CMP 40HX, which only has two high-scoring compute benchmarks. The CMP 40HX's 94th percentile ranking, matching the RTX 5090's percentile, suggests that in the broader GPU landscape, both cards sit at similar overall performance tiers despite their wildly different architectures.
Specification Differences
The most striking difference is transistor count: the RTX 5090 packs 92,200 million transistors versus 10,800 million on the CMP 40HX, an 8.5x increase. Die size grows from 445 mm² to 750 mm², while transistor density jumps from 24.3M to 122.9M per mm². Process node shrinks from 12 nm to 5 nm, both TSMC, enabling the RTX 5090's higher clocks and core counts.
Memory specifications differ completely: 8 GB GDDR6 versus 32 GB GDDR7, 256-bit versus 512-bit bus width, and 448.0 GB/s versus 1.79 TB/s bandwidth. The RTX 5090 also has faster effective memory speed at 28 Gbps versus 14 Gbps, though both run at 1750 MHz base memory clock.
Core counts scale dramatically: shading units go from 2,304 to 21,760 (9.4x), TMUs from 144 to 680 (4.7x), ROPs from 64 to 176 (2.75x), RT cores from 36 to 170 (4.7x), and tensor cores from 288 to 680 (2.4x). Pixel rate rises from 105.6 to 423.6 GPixel/s, texture rate from 237.6 to 1,636.8 GTexel/s, and FP32 from 7.603 to 104.8 TFLOPS. FP16 throughput changes from 15.21 TFLOPS (2:1) to 104.8 TFLOPS (1:1), indicating architectural efficiency improvements.
Physical and power specifications also diverge: the CMP 40HX measures 229 mm x 111 mm x 35 mm with 185 W TDP and 450 W suggested PSU, while the RTX 5090 is 304 mm x 137 mm x 40 mm with 575 W TDP and 950 W suggested PSU. Both are dual-slot, but the CMP 40HX uses 1x 8-pin power while the RTX 5090 requires 1x 16-pin. Bus interface differs from PCIe 1.0 x4 on the CMP 40HX to PCIe 5.0 x16 on the RTX 5090—a substantial bandwidth difference for data transfer.
Release dates and production status reflect their different lifecycles: the CMP 40HX launched 2021-02-24 and is end-of-life, while the RTX 5090 launched 2025-01-29 and remains active. The CMP 40HX's launch MSRP was 699 USD; the RTX 5090's was 1,999 USD. Both support DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4, but only the RTX 5090 offers display outputs (1x HDMI 2.1b, 3x DisplayPort 2.1b).