NVIDIA CMP 40HX vs NVIDIA L40 Comparison

NVIDIA
GEFORCE

NVIDIA CMP 40HX

CORE STATE TU106
VRAM 8 GB
CLOCK SPEED 1650 MHz
TDP 185 W
BUS WIDTH 256 bit
ARCHITECTURE Turing
nm
PROCESS 12 nm
LAUNCH DATE 2021
VS
NVIDIA
GEFORCE

L40

CORE STATE AD102
VRAM 48 GB
CLOCK SPEED 2490 MHz
TDP 300 W
BUS WIDTH 384 bit
ARCHITECTURE Ada Lovelace
nm
PROCESS 5 nm
LAUNCH DATE 2022

PERFORMANCE BENCHMARKS

geekbench_opencl
93,395
330,926
geekbench_vulkan
77,879
237,295

Analysis: NVIDIA CMP 40HX vs NVIDIA L40

# NVIDIA L40 vs NVIDIA CMP 40HX

The NVIDIA L40 and NVIDIA CMP 40HX occupy opposite ends of the GPU spectrum, despite sharing the NVIDIA brand. The L40 is a server-class Ada Lovelace accelerator designed for compute and professional visualization, while the CMP 40HX is a Turing-based mining card with no display outputs. Benchmark data from the database shows a decisive performance gap: the L40 averages 284,111 points across tests, placing it in the 99th percentile of all GPUs, whereas the CMP 40HX averages 85,637 points, sitting in the 93rd percentile. That difference of roughly 232% in average score reflects not just generational progress but entirely different design goals, power envelopes, and target workloads.

Where Each One Wins

The NVIDIA L40 wins every recorded benchmark in this comparison, and by overwhelming margins. In Geekbench OpenCL, the L40 scores 330,926 points against the CMP 40HX's 93,395, a 254.3% advantage. In Geekbench Vulkan, the L40 posts 237,295 points versus 77,879, a 204.7% lead. There is no workload category in the recorded data where the CMP 40HX comes out ahead.

The L40's dominance stems from its architecture and configuration. It carries 18,176 shading units, 568 texture mapping units, and 192 raster operation units, along with 142 ray tracing cores and 568 tensor cores. That hardware translates into 90.52 TFLOPS of FP32 compute and 1,414.3 GTexel/s of texture throughput. For AI inference, deep learning training, scientific simulation, or high-end rendering, the L40 is built to crush these tasks. The CMP 40HX, by contrast, offers 2,304 shading units, 144 TMUs, and 64 ROPs, with 36 ray tracing cores and 288 tensor cores. Its FP32 peak is 7.603 TFLOPS, and its texture rate is 237.6 GTexel/s, roughly one-sixth of the L40's capabilities.

The CMP 40HX was designed for cryptocurrency mining, a workload that relies heavily on memory bandwidth and raw arithmetic but not on display output or advanced graphics features. It has no display outputs, which immediately disqualifies it from any interactive or visualization role. The L40, with four DisplayPort 1.4a outputs, can drive professional displays and is suited for workstation use. In the database's recorded benchmarks, the CMP 40HX does not achieve a single win, so its practical use case lies outside the tested metrics, primarily in mining operations where its 185 W power draw and 8 GB GDDR6 memory might have been adequate at launch.

For any compute task involving FP32, FP16, ray tracing, or tensor operations, the L40 is the only viable choice of these two. The CMP 40HX's Turing architecture does support FP16 at a 2:1 ratio (15.21 TFLOPS), but even that is far below the L40's FP16 figure of 90.52 TFLOPS, which runs at 1:1 with FP32. In short, the L40 wins everywhere; the CMP 40HX is a niche product whose strengths, such as they are, were never measured by mainstream benchmark suites.

FAQ

Q: Which GPU has higher benchmark scores?

A: The NVIDIA L40 wins both recorded tests. In Geekbench OpenCL, it scores 330,926 versus 93,395, a 254.3% lead. In Geekbench Vulkan, it scores 237,295 versus 77,879, a 204.7% advantage.

Q: What is the average benchmark score for each card?

A: The L40 averages 284,111 points across all recorded benchmarks, placing it in the 99th percentile of all GPUs. The CMP 40HX averages 85,637 points, placing it in the 93rd percentile.

Q: How does the L40 compare to its nearest rivals?

A: The L40 sits close to the NVIDIA RTX 6000 Ada Generation, which scores 287,237 on average, a 1.1% difference. It trails the NVIDIA L40S (295,763) by 3.9% but leads the NVIDIA L20 (251,147) by 13.1% and the AMD Instinct MI300X (317,994) by 10.7% in the reverse direction.

Q: Does the CMP 40HX have any display outputs?

A: No. The CMP 40HX has no display outputs, making it unsuitable for any interactive or visual workload. The L40, in contrast, has four DisplayPort 1.4a outputs.

Q: What is the memory configuration difference?

A: The L40 has 48 GB of GDDR6 on a 384-bit bus, delivering 864.0 GB/s of bandwidth. The CMP 40HX has 8 GB of GDDR6 on a 256-bit bus, delivering 448.0 GB/s.

Q: What is the launch MSRP of the CMP 40HX?

A: The CMP 40HX had a launch MSRP of 699 USD. The L40's launch MSRP is not recorded in the database.

Head-to-Head Benchmarks

The database records two head-to-head comparisons, and both are lopsided. In Geekbench OpenCL, the L40 scores 330,926 points, while the CMP 40HX manages only 93,395. That is a 254.3% delta, meaning the L40 delivers more than three and a half times the raw compute performance in this test. The OpenCL workload stresses general-purpose compute, including memory bandwidth and shader throughput, areas where the L40's 48 GB frame buffer and 864.0 GB/s bandwidth vastly outperform the CMP 40HX's 8 GB and 448.0 GB/s.

In Geekbench Vulkan, the margin narrows slightly but remains enormous. The L40 scores 237,295 points, and the CMP 40HX scores 77,879, a 204.7% difference. Vulkan is a lower-level API that can expose hardware efficiency differences, but here the L40's Ada Lovelace architecture still dominates. The L40's pixel rate of 478.1 GPixel/s and texture rate of 1,414.3 GTexel/s dwarf the CMP 40HX's 105.6 GPixel/s and 237.6 GTexel/s, respectively.

What is notable is that the CMP 40HX's closest rivals in the database are not modern cards but older or midrange workstation GPUs. Its nearest competitor is the AMD Radeon PRO W7600, which averages 87,108 points, a 1.7% difference, followed by the NVIDIA Quadro GP100 at 87,445 (2.1% difference). The CMP 40HX leads the AMD Radeon PRO W6600 by 4.4% and the AMD Radeon Pro Vega 64X by 5.8%. These are modest margins, indicating that the CMP 40HX performs at a level roughly comparable to a midrange workstation GPU from a few years ago. The L40, by contrast, sits within 1.1% of the RTX 6000 Ada Generation and within 3.9% of the L40S, both of which are current server-class parts.

The performance gap between the two cards is not linear with their specifications; it is geometric. The L40 has 7.9 times the shading units, 3.9 times the TMUs, 3 times the ROPs, and 3.9 times the ray tracing cores of the CMP 40HX. Its FP32 throughput is 11.9 times higher. These disparities compound in real workloads, which is why the benchmark deltas exceed 200% in both tests.

Specification Differences

The two cards differ in nearly every measurable specification. The L40 uses the AD102 chip on a 5 nm process, while the CMP 40HX uses the TU106 chip on a 12 nm process. Transistor counts reflect this: the L40 packs 76,300 million transistors on a 609 mm² die (125.3 million per square millimeter), whereas the CMP 40HX has 10,800 million transistors on a 445 mm² die (24.3 million per square millimeter). The L40's transistor density is more than five times higher.

Clock speeds tell a different story. The CMP 40HX has a higher base clock (1470 MHz versus 735 MHz) and a lower boost clock (1650 MHz versus 2490 MHz). The L40's boost clock is significantly higher, which, combined with its massive shader count, explains its performance advantage. Memory clocks also differ: the L40 runs at 2250 MHz (18 Gbps effective), while the CMP 40HX runs at 1750 MHz (14 Gbps effective).

Memory capacity and bandwidth are major separators. The L40 has 48 GB of GDDR6 on a 384-bit bus, yielding 864.0 GB/s. The CMP 40HX has 8 GB of GDDR6 on a 256-bit bus, yielding 448.0 GB/s. That is a 6x capacity difference and a 1.93x bandwidth difference. The bus interface also differs: the L40 uses PCIe 4.0 x16, while the CMP 40HX uses PCIe 1.0 x4, a legacy interface that severely limits data transfer to and from the host system.

Power and physical specifications diverge as well. The L40 has a 300 W TDP and requires a 700 W suggested PSU with a 1x 16-pin power connector. The CMP 40HX has a 185 W TDP, a 450 W suggested PSU, and a 1x 8-pin connector. Both are dual-slot cards, but the L40 is longer at 267 mm (10.5 inches) versus 229 mm (9 inches). Both are 111 mm tall (4.4 inches), and the CMP 40HX has a width of 35 mm (1.4 inches), while the L40's width is not recorded. The L40 has four DisplayPort 1.4a outputs; the CMP 40HX has none.

Architecture Differences

The L40 is built on the Ada Lovelace architecture, a direct descendant of the Ampere generation, and is part of the Server Ada (Lxx) product line. The CMP 40HX is built on the Turing architecture and belongs to the Mining GPUs generation. These are fundamentally different designs separated by two architectural generations. Ada Lovelace introduced significant improvements in ray tracing performance, tensor core throughput, and power efficiency compared to Turing. The L40's 142 ray tracing cores and 568 tensor cores are not just more numerous but also more capable per core than the CMP 40HX's 36 ray tracing cores and 288 tensor cores.

The process node is a critical differentiator. The L40 uses TSMC's 5 nm process, while the CMP 40HX uses TSMC's 12 nm process. This explains the L40's higher transistor density (125.3M per mm² versus 24.3M per mm²) and its ability to fit 76,300 million transistors on a die that is only 37% larger by area. The CMP 40HX's TU106 chip is a smaller, older design that was originally used in consumer Turing GPUs.

Feature sets also differ. The L40 supports DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4, as does the CMP 40HX. However, the L40's inclusion of display outputs enables full graphics and visualization workloads, while the CMP 40HX's lack of outputs makes it a compute-only device. The L40 also has a higher FP16 to FP32 ratio (1:1 versus 2:1 for the CMP 40HX), meaning it can handle FP16 workloads without halving throughput. The CMP 40HX's FP16 rate of 15.21 TFLOPS is actually double its FP32 rate, but in absolute terms it is still far below the L40's 90.52 TFLOPS for both precisions.

The L40's release date is October 12, 2022, while the CMP 40HX launched on February 24, 2021. The L40's production status is end-of-life, as is the CMP 40HX's, but the L40 has a successor (Server Hopper) and a predecessor (Server Ampere), while the CMP 40HX has neither. The L40's architecture is designed for long-term server deployment, whereas the CMP 40HX was a short-lived product tied to the cryptocurrency mining boom. In every architectural metric, the L40 is the more advanced, more capable, and more versatile GPU.

DETAILED SPECIFICATIONS

SPECIFICATION
CMP 40HX
L40
Core Specs
Shading Units
2,304
18,176 +688.9%
Shaders
2,304
18,176 +688.9%
TMUs
144
568 +294.4%
ROPs
64
192 +200.0%
SM Count
36
142 +294.4%
Clocks
Base Clock
1470 MHz
735 MHz
Boost Clock
1650 MHz
2490 MHz
Memory Clock
1750 MHz 14 Gbps effective
2250 MHz 18 Gbps effective
Memory
Memory Size
8 GB
48 GB
VRAM (MB)
8,192
49,152 +500.0%
Memory Type
GDDR6
GDDR6
Memory Bus
256 bit
384 bit
Bandwidth
448.0 GB/s
864.0 GB/s
Cache
L1 Cache
64 KB (per SM)
128 KB (per SM)
L2 Cache
4 MB
96 MB
Performance
Pixel Rate
105.6 GPixel/s
478.1 GPixel/s
Texture Rate
237.6 GTexel/s
1,414.3 GTexel/s
FP32 (TFLOPS)
7.603 TFLOPS
90.52 TFLOPS
FP64 (TFLOPS)
237.6 GFLOPS (1:32)
1,414.3 GFLOPS (1:64)
FP16 (TFLOPS)
15.21 TFLOPS (2:1)
90.52 TFLOPS (1:1)
AI/RT
RT Cores
36
142 +294.4%
Tensor Cores
288
568 +97.2%
Power
TDP
185 W
300 W
TDP (W)
185
300 +62.2%
Suggested PSU
450 W
700 W
Power Connectors
1x 8-pin
1x 16-pin
Architecture
Architecture
Turing
Ada Lovelace
GPU Name
TU106
AD102
Generation
Mining GPUs
Server Ada (Lxx)
Process Size
12 nm
5 nm
Transistors
10,800 million
76,300 million
Die Size
445 mm²
609 mm²
Foundry
TSMC
TSMC
Density
24.3M / mm²
125.3M / mm²
API Support
DirectX
12 Ultimate (12_2)
12 Ultimate (12_2)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
3.0
3.0
CUDA
7.5
8.9
Shader Model
6.8
6.8
Physical
Slot Width
Dual-slot
Dual-slot
Length
229 mm 9 inches
267 mm 10.5 inches
Height
111 mm 4.4 inches
111 mm 4.4 inches
Outputs
No outputs
4x DisplayPort 1.4a
Bus Interface
PCIe 1.0 x4
PCIe 4.0 x16
Other
Launch Price
699 USD
Production
End-of-life
End-of-life
Predecessor
Server Ampere
Successor
Server Hopper
View CMP 40HX Details View L40 Details