NVIDIA CMP 40HX vs NVIDIA L4 Comparison

NVIDIA
GEFORCE

NVIDIA CMP 40HX

CORE STATE TU106
VRAM 8 GB
CLOCK SPEED 1650 MHz
TDP 185 W
BUS WIDTH 256 bit
ARCHITECTURE Turing
nm
PROCESS 12 nm
LAUNCH DATE 2021
VS
NVIDIA
GEFORCE

L4

CORE STATE AD104
VRAM 24 GB
CLOCK SPEED 2040 MHz
TDP 72 W
BUS WIDTH 192 bit
ARCHITECTURE Ada Lovelace
nm
PROCESS 5 nm
LAUNCH DATE 2023

PERFORMANCE BENCHMARKS

geekbench_opencl
93,395
140,838
geekbench_vulkan
77,879
121,306

Analysis: NVIDIA CMP 40HX vs NVIDIA L4

FAQ

Q: Which GPU has the higher average benchmark score?

A: The NVIDIA L4 records an average benchmark score of 131072, while the NVIDIA CMP 40HX records 85637. This places the L4 in the 95th percentile of all GPUs, compared to the 93rd percentile for the CMP 40HX.

Q: How much faster is the NVIDIA L4 in OpenCL and Vulkan?

A: In Geekbench OpenCL, the L4 scores 140838 versus 93395 for the CMP 40HX, a 50.8% advantage. In Geekbench Vulkan, the L4 scores 121306 versus 77879, a 55.8% advantage.

Q: What are the memory configurations of both cards?

A: The NVIDIA L4 has 24 GB of GDDR6 memory on a 192-bit bus with 300.1 GB/s bandwidth. The NVIDIA CMP 40HX has 8 GB of GDDR6 on a 256-bit bus with 448.0 GB/s bandwidth.

Q: Which GPU consumes less power?

A: The NVIDIA L4 has a TDP of 72 W, while the NVIDIA CMP 40HX has a TDP of 185 W. The L4 also requires no power connectors, whereas the CMP 40HX needs a single 8-pin connector.

Q: What is the release timeline for these two cards?

A: The NVIDIA CMP 40HX was released on 2021-02-24 and is now end-of-life. The NVIDIA L4 was released on 2023-03-20 and is still active in production.

Q: Which GPU has the higher transistor count?

A: The NVIDIA L4 uses 35,800 million transistors on a 294 mm² die, while the NVIDIA CMP 40HX uses 10,800 million transistors on a 445 mm² die.

Architecture Differences

The NVIDIA L4 and NVIDIA CMP 40HX represent two fundamentally different design eras from NVIDIA. The L4 is built on the Ada Lovelace architecture using TSMC's 5 nm process, while the CMP 40HX uses the older Turing architecture on TSMC's 12 nm process. This process difference alone explains much of the performance and efficiency gap between the two.

The L4's chip, designated AD104, packs 35,800 million transistors into a 294 mm² die, yielding a transistor density of 121.8 million transistors per square millimeter. The CMP 40HX uses the TU106 chip with 10,800 million transistors spread across a much larger 445 mm² die, giving a density of only 24.3 million transistors per square millimeter. This means the L4 fits over five times more transistors per area, which directly translates to higher compute throughput at lower power.

Clock speeds tell a contrasting story. The CMP 40HX runs at a 1470 MHz base clock and 1650 MHz boost clock, while the L4 runs at 795 MHz base and 2040 MHz boost. The L4's boost clock is significantly higher, and its lower base clock is offset by the massive architectural efficiency gains of Ada Lovelace.

The shading resources differ substantially. The L4 features 7424 shading units, 240 texture mapping units, 80 raster operation pipelines, 60 ray tracing cores, and 240 tensor cores. The CMP 40HX has 2304 shading units, 144 TMUs, 64 ROPs, 36 RT cores, and 288 tensor cores. While the CMP 40HX has more tensor cores on paper, the L4's tensor cores are far more capable per unit due to the newer architecture.

Memory architecture also diverges. The L4 offers 24 GB of GDDR6 on a 192-bit bus, while the CMP 40HX offers 8 GB on a 256-bit bus. Interestingly, the CMP 40HX has higher memory bandwidth at 448.0 GB/s versus 300.1 GB/s for the L4, but the L4 compensates with far greater compute density.

The interface differences are stark: the L4 uses PCIe 4.0 x16, while the CMP 40HX is limited to PCIe 1.0 x4. This makes the L4 far more suitable for modern server workloads that require high data transfer rates. The L4 is also single-slot with no power connectors, while the CMP 40HX is dual-slot with a single 8-pin connector. Neither card has display outputs, confirming their compute-only roles.

Head-to-Head Benchmarks

The recorded data shows a decisive victory for the NVIDIA L4 across both benchmark tests. In Geekbench OpenCL, the L4 scores 140838 against the CMP 40HX's 93395, representing a 50.8% performance lead. This is not a marginal difference; it is a substantial gap that places the two cards in entirely different performance tiers.

The Vulkan results amplify this trend. The L4 scores 121306, while the CMP 40HX manages 77879. The delta widens to 55.8%, meaning the L4 is more than half again as fast as the CMP 40HX in this workload. The consistency of the lead across both APIs suggests this is an architectural superiority rather than a workload-specific quirk.

Looking at the nearest rivals provides context for these scores. The L4's average benchmark score of 131072 sits just 0.7% below the NVIDIA GeForce RTX 3090 Ti (131938) and 3.1% below both the NVIDIA RTX 4000 Ada Generation (135218) and NVIDIA A10M (135230). This places the L4 in the company of high-end workstation and server GPUs, despite its modest 72 W TDP.

The CMP 40HX, by contrast, sits in a lower tier. Its average score of 85637 is 1.7% below the AMD Radeon PRO W7600 (87108) and 2.1% below the NVIDIA Quadro GP100 (87445). It does lead the AMD Radeon PRO W6600 by 4.4% and the AMD Radeon Pro Vega 64X by 5.8%, but these are mid-range competitors.

The FP32 compute figures reinforce the benchmark results. The L4 delivers 30.29 TFLOPS of FP32 performance, while the CMP 40HX delivers 7.603 TFLOPS. The L4 is roughly four times faster in raw single-precision math. For FP16, the L4 again delivers 30.29 TFLOPS at a 1:1 ratio, while the CMP 40HX delivers 15.21 TFLOPS at a 2:1 ratio. The L4's FP16 performance is double that of the CMP 40HX, and the 1:1 ratio means it does not sacrifice throughput for precision.

Pixel and texture rates follow the same pattern. The L4 achieves 163.2 GPixel/s and 489.6 GTexel/s, while the CMP 40HX achieves 105.6 GPixel/s and 237.6 GTexel/s. The L4 leads by 54.5% in pixel throughput and 106.1% in texture throughput.

The Verdict

The data is unambiguous: the NVIDIA L4 outperforms the NVIDIA CMP 40HX in every recorded benchmark and every compute metric. The L4 wins both head-to-head tests, achieving a 50.8% lead in OpenCL and a 55.8% lead in Vulkan. Its average benchmark score of 131072 is 53.1% higher than the CMP 40HX's 85637.

The L4 achieves this dominance while consuming only 72 W, compared to the CMP 40HX's 185 W. This means the L4 delivers vastly more performance per watt, a critical factor for server deployments where power density and cooling are primary constraints. The L4 requires no power connectors and fits in a single slot, while the CMP 40HX needs a dual-slot bracket, a single 8-pin connector, and a 450 W suggested power supply.

The CMP 40HX does have one advantage: memory bandwidth. Its 448.0 GB/s bandwidth is 49.3% higher than the L4's 300.1 GB/s, and its 256-bit bus is wider. However, this advantage does not translate into better benchmark scores, likely because the CMP 40HX's much lower compute throughput cannot utilize the bandwidth effectively.

The L4 also offers nearly three times the memory capacity at 24 GB versus 8 GB, which matters for large model inference and data-parallel workloads. The CMP 40HX's 8 GB capacity is limiting for modern deep learning models, while the L4 can accommodate substantially larger working sets.

The production status seals the decision. The L4 remains active, while the CMP 40HX is end-of-life. Choosing the L4 is both a performance decision and a longevity decision. The CMP 40HX was designed for mining, not general compute, and its Turing architecture lacks the efficiency and feature set of Ada Lovelace.

For any organization choosing between these two GPUs, the L4 is the clear pick based on the recorded data. It wins every benchmark, offers superior compute specifications, consumes less power, and remains in active production.

Specification Differences

The following specifications differ between the NVIDIA L4 and NVIDIA CMP 40HX:

  • Architecture: Ada Lovelace (L4) versus Turing (CMP 40HX)
  • Process node: 5 nm (L4) versus 12 nm (CMP 40HX), both TSMC
  • Transistors: 35,800 million (L4) versus 10,800 million (CMP 40HX)
  • Die size: 294 mm² (L4) versus 445 mm² (CMP 40HX)
  • Base clock: 795 MHz (L4) versus 1470 MHz (CMP 40HX)
  • Boost clock: 2040 MHz (L4) versus 1650 MHz (CMP 40HX)
  • Memory speed: 1563 MHz / 12.5 Gbps effective (L4) versus 1750 MHz / 14 Gbps effective (CMP 40HX)
  • Memory size: 24 GB (L4) versus 8 GB (CMP 40HX)
  • Memory bus width: 192 bit (L4) versus 256 bit (CMP 40HX)
  • Memory bandwidth: 300.1 GB/s (L4) versus 448.0 GB/s (CMP 40HX)
  • Shading units: 7424 (L4) versus 2304 (CMP 40HX)
  • Texture mapping units: 240 (L4) versus 144 (CMP 40HX)
  • Raster operation pipelines: 80 (L4) versus 64 (CMP 40HX)
  • Ray tracing cores: 60 (L4) versus 36 (CMP 40HX)
  • Tensor cores: 240 (L4) versus 288 (CMP 40HX)
  • Pixel rate: 163.2 GPixel/s (L4) versus 105.6 GPixel/s (CMP 40HX)
  • Texture rate: 489.6 GTexel/s (L4) versus 237.6 GTexel/s (CMP 40HX)
  • FP32 performance: 30.29 TFLOPS (L4) versus 7.603 TFLOPS (CMP 40HX)
  • FP16 performance: 30.29 TFLOPS 1:1 (L4) versus 15.21 TFLOPS 2:1 (CMP 40HX)
  • TDP: 72 W (L4) versus 185 W (CMP 40HX)
  • Slot width: Single-slot (L4) versus Dual-slot (CMP 40HX)
  • Power connectors: None (L4) versus 1x 8-pin (CMP 40HX)
  • Suggested PSU: 250 W (L4) versus 450 W (CMP 40HX)
  • Bus interface: PCIe 4.0 x16 (L4) versus PCIe 1.0 x4 (CMP 40HX)
  • Length: 169 mm (L4) versus 229 mm (CMP 40HX)
  • Height: 56 mm (L4) versus 111 mm (CMP 40HX)
  • Width: not listed (L4) versus 35 mm (CMP 40HX)
  • Production status: Active (L4) versus End-of-life (CMP 40HX)
  • Release date: 2023-03-20 (L4) versus 2021-02-24 (CMP 40HX)
  • Launch MSRP: not listed (L4) versus 699 USD (CMP 40HX)

Where Each One Wins

The NVIDIA L4 wins in every compute benchmark and most architectural metrics. Its strengths are compute throughput, memory capacity, power efficiency, and modern interface support. The L4 is the appropriate choice for general GPU compute, deep learning inference, AI workloads, and any server application requiring high FP32 or FP16 throughput. Its 24 GB memory capacity allows it to handle large model weights and datasets that would exceed the CMP 40HX's 8 GB limit. The 1:1 FP16 ratio means the L4 can process half-precision workloads without the throughput penalty seen on the CMP 40HX.

The L4's 72 W TDP and single-slot design make it suitable for dense server deployments where power and space are constrained. The lack of power connectors simplifies installation, and the 250 W suggested PSU is modest. The PCIe 4.0 x16 interface ensures high data transfer rates with modern host systems.

The NVIDIA CMP 40HX has a narrower set of advantages. Its 448.0 GB/s memory bandwidth is significantly higher than the L4's 300.1 GB/s, which could benefit workloads that are memory-bound rather than compute-bound, such as certain data processing or hash-based operations. The 256-bit memory bus provides higher parallel access to memory. The CMP 40HX also has a higher base clock at 1470 MHz, which may help in short-burst workloads that do not sustain boost clocks.

The CMP 40HX has more tensor cores (288 versus 240), but these are Turing-generation tensor cores with lower per-core throughput. The CMP 40HX was designed for mining workloads, as indicated by its generation label "Mining GPUs," and its PCIe 1.0 x4 interface reflects that design intent. For its original purpose, the higher memory bandwidth and lower cost may have been sufficient, but the data shows it lags far behind the L4 in all general compute benchmarks.

For users with workloads that specifically require maximum memory bandwidth and can tolerate lower compute throughput, the CMP 40HX retains a niche. However, the L4's 50.8% to 55.8% benchmark leads, combined with its active production status and far lower power draw, make it the dominant choice for nearly all modern compute applications. The CMP 40HX is end-of-life, which means driver support and availability will only decline over time, whereas the L4 remains an active product line.

DETAILED SPECIFICATIONS

SPECIFICATION
CMP 40HX
L4
Core Specs
Shading Units
2,304
7,424 +222.2%
Shaders
2,304
7,424 +222.2%
TMUs
144
240 +66.7%
ROPs
64
80 +25.0%
SM Count
36
60 +66.7%
Clocks
Base Clock
1470 MHz
795 MHz
Boost Clock
1650 MHz
2040 MHz
Memory Clock
1750 MHz 14 Gbps effective
1563 MHz 12.5 Gbps effective
Memory
Memory Size
8 GB
24 GB
VRAM (MB)
8,192
24,576 +200.0%
Memory Type
GDDR6
GDDR6
Memory Bus
256 bit
192 bit
Bandwidth
448.0 GB/s
300.1 GB/s
Cache
L1 Cache
64 KB (per SM)
128 KB (per SM)
L2 Cache
4 MB
48 MB
Performance
Pixel Rate
105.6 GPixel/s
163.2 GPixel/s
Texture Rate
237.6 GTexel/s
489.6 GTexel/s
FP32 (TFLOPS)
7.603 TFLOPS
30.29 TFLOPS
FP64 (TFLOPS)
237.6 GFLOPS (1:32)
473.3 GFLOPS (1:64)
FP16 (TFLOPS)
15.21 TFLOPS (2:1)
30.29 TFLOPS (1:1)
AI/RT
RT Cores
36
60 +66.7%
Tensor Cores
288
240 -16.7%
Power
TDP
185 W
72 W
TDP (W)
185
72 -61.1%
Suggested PSU
450 W
250 W
Power Connectors
1x 8-pin
None
Architecture
Architecture
Turing
Ada Lovelace
GPU Name
TU106
AD104
Generation
Mining GPUs
Server Ada (Lxx)
Process Size
12 nm
5 nm
Transistors
10,800 million
35,800 million
Die Size
445 mm²
294 mm²
Foundry
TSMC
TSMC
Density
24.3M / mm²
121.8M / mm²
API Support
DirectX
12 Ultimate (12_2)
12 Ultimate (12_2)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
3.0
3.0
CUDA
7.5
8.9
Shader Model
6.8
6.8
Physical
Slot Width
Dual-slot
Single-slot
Length
229 mm 9 inches
169 mm 6.7 inches
Height
111 mm 4.4 inches
56 mm 2.2 inches
Outputs
No outputs
No outputs
Bus Interface
PCIe 1.0 x4
PCIe 4.0 x16
Other
Launch Price
699 USD
Production
End-of-life
Active
Predecessor
Server Ampere
Successor
Server Hopper
View CMP 40HX Details View L4 Details