NVIDIA CMP 40HX vs NVIDIA L20 Comparison

NVIDIA
GEFORCE

NVIDIA CMP 40HX

CORE STATE TU106
VRAM 8 GB
CLOCK SPEED 1650 MHz
TDP 185 W
BUS WIDTH 256 bit
ARCHITECTURE Turing
nm
PROCESS 12 nm
LAUNCH DATE 2021
VS
NVIDIA
GEFORCE

L20

CORE STATE AD102
VRAM 48 GB
CLOCK SPEED 2520 MHz
TDP 275 W
BUS WIDTH 384 bit
ARCHITECTURE Ada Lovelace
nm
PROCESS 5 nm
LAUNCH DATE 2023

PERFORMANCE BENCHMARKS

geekbench_opencl
93,395
274,276
geekbench_vulkan
77,879
228,018

Analysis: NVIDIA CMP 40HX vs NVIDIA L20

Head-to-Head Benchmarks

The recorded benchmark data shows a complete sweep for the NVIDIA L20 across both tested workloads. In Geekbench OpenCL, the L20 scores 274,276 against 93,395 for the CMP 40HX, a delta of 193.7%. That is not a marginal lead; it is a near-tripling of compute throughput in this API. The Vulkan result is equally lopsided: 228,018 versus 77,879, a 192.8% advantage. The L20 wins both head-to-head tests, giving it 2 wins and the CMP 40HX 0 wins.

Context from the nearest rivals puts these numbers in perspective. The L20 sits at the 99th percentile among all GPUs in the database, with an average benchmark score of 251,147. Its closest competitors are the NVIDIA RTX 6000 Ada Generation at 287,237 (12.6% higher), the NVIDIA L40 at 284,111 (11.6% higher), the NVIDIA PG506-232 at 225,124 (11.6% lower), and the AMD Radeon PRO W7900D at 219,827 (14.2% lower). So the L20 is positioned just below the top-tier server Ada cards, but it still outruns the previous-generation Ampere server part and a modern AMD workstation flagship by double-digit percentages.

The CMP 40HX, by contrast, sits at the 93rd percentile with an average score of 85,637. Its nearest rivals are much closer in performance: the AMD Radeon PRO W7600 trails by only 1.7% (87,108), the NVIDIA Quadro GP100 trails by 2.1% (87,445), while the AMD Radeon PRO W6600 scores 81,995 (4.4% lower) and the AMD Radeon Pro Vega 64X scores 80,959 (5.8% lower). The CMP 40HX is essentially at parity with those mid-range workstation cards, not in the same performance class as the L20. The gap between the two cards in this comparison is roughly 193%, which dwarfs any delta seen within their respective rival clusters.

For builders deciding between these two, the data says the L20 is not merely faster; it is in a different tier entirely. The CMP 40HX was designed for a different purpose, and the benchmark results reflect that chasm.

FAQ

Q: Which card has the higher average benchmark score?

A: The NVIDIA L20 averages 251,147 across all recorded benchmarks, compared to 85,637 for the NVIDIA CMP 40HX. The L20 also ranks at the 99th percentile among all GPUs, while the CMP 40HX ranks at the 93rd percentile.

Q: How much faster is the L20 in OpenCL and Vulkan?

A: In Geekbench OpenCL, the L20 scores 274,276 versus 93,395 for the CMP 40HX, a 193.7% advantage. In Geekbench Vulkan, the L20 scores 228,018 versus 77,879, a 192.8% advantage.

Q: What are the memory capacities and types?

A: The L20 has 48 GB of GDDR6 on a 384-bit bus with 864.0 GB/s bandwidth. The CMP 40HX has 8 GB of GDDR6 on a 256-bit bus with 448.0 GB/s bandwidth.

Q: Which card has higher boost clocks?

A: The L20 boosts to 2520 MHz from a 1440 MHz base. The CMP 40HX boosts to 1650 MHz from a 1470 MHz base. The L20 also runs memory at 2250 MHz (18 Gbps effective), while the CMP 40HX runs memory at 1750 MHz (14 Gbps effective).

Q: What is the production status of each card?

A: The NVIDIA L20 is listed as Active production. The NVIDIA CMP 40HX is listed as End-of-life.

Q: Do both cards support the same APIs?

A: Yes, both support DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4. However, the L20 has 4x DisplayPort 1.4a outputs, while the CMP 40HX has no display outputs at all.

Architecture Differences

The two cards come from entirely different architectural generations. The NVIDIA L20 is built on Ada Lovelace, using the AD102 chip, while the CMP 40HX is built on Turing, using the TU106 chip. This is a two-generation jump, and the process node reflects it: the L20 is on a 5 nm process at TSMC, whereas the CMP 40HX is on a 12 nm process, also from TSMC. Transistor counts tell the story of scale: the L20 packs 76,300 million transistors on a 609 mm² die, yielding a density of 125.3 million transistors per square millimeter. The CMP 40HX has 10,800 million transistors on a 445 mm² die, a density of just 24.3 million per square millimeter. That is over a 5x difference in transistor density, which explains the massive compute gap.

The compute resources are similarly lopsided. The L20 has 11,776 shading units, 368 texture mapping units, and 128 ROPs. It also carries 92 ray tracing cores and 368 tensor cores. The CMP 40HX has 2,304 shading units, 144 TMUs, and 64 ROPs, with 36 ray tracing cores and 288 tensor cores. The L20 has more than 5x the shading units and more than 2.5x the TMUs. The tensor core count is closer (368 versus 288), but the L20's newer architecture delivers far higher per-core throughput.

Pixel and texture rates reinforce the hierarchy. The L20 achieves 322.6 GPixel/s and 927.4 GTexel/s. The CMP 40HX manages 105.6 GPixel/s and 237.6 GTexel/s. FP32 compute is 59.35 TFLOPS for the L20 versus 7.603 TFLOPS for the CMP 40HX. FP16 is where the architectures differ in approach: the L20 delivers 59.35 TFLOPS at a 1:1 ratio with FP32, while the CMP 40HX delivers 15.21 TFLOPS at a 2:1 ratio. The L20's FP16 is nearly 4x higher, and it does not sacrifice FP32 throughput to get there.

Another key distinction is the PCIe interface. The L20 uses PCIe 4.0 x16, which is a full-bandwidth modern connection. The CMP 40HX uses PCIe 1.0 x4, an extremely old and narrow interface that severely limits data transfer to and from the host system. This is a critical architectural difference for any workload that involves frequent host-device transfers, as the CMP 40HX will bottleneck long before its compute resources are saturated. The L20 also provides 4x DisplayPort 1.4a outputs, while the CMP 40HX has no display outputs, reflecting its mining-oriented design.

Specification Differences

The specification sheet shows a clear divide across nearly every field. The process node differs: 5 nm for the L20 versus 12 nm for the CMP 40HX. Transistors differ by a factor of roughly 7: 76,300 million versus 10,800 million. Die size is 609 mm² versus 445 mm². Clock speeds: the L20 has a 1440 MHz base and 2520 MHz boost; the CMP 40HX has a 1470 MHz base and 1650 MHz boost. Memory clocks are 2250 MHz (18 Gbps effective) versus 1750 MHz (14 Gbps effective).

Memory capacity is 48 GB versus 8 GB, bus width is 384-bit versus 256-bit, and bandwidth is 864.0 GB/s versus 448.0 GB/s. Shading units are 11,776 versus 2,304, TMUs are 368 versus 144, and ROPs are 128 versus 64. Ray tracing cores are 92 versus 36, and tensor cores are 368 versus 288. Pixel rate is 322.6 GPixel/s versus 105.6 GPixel/s, and texture rate is 927.4 GTexel/s versus 237.6 GTexel/s. FP32 is 59.35 TFLOPS versus 7.603 TFLOPS. FP16 is 59.35 TFLOPS versus 15.21 TFLOPS.

Power draw differs as well: the L20 has a TDP of 275 W with a suggested PSU of 600 W and a 1x 16-pin power connector. The CMP 40HX has a TDP of 185 W, a suggested PSU of 450 W, and a 1x 8-pin connector. The L20 is longer at 267 mm (10.5 inches) versus 229 mm (9 inches), but both share the same 111 mm height (4.4 inches). The CMP 40HX is 35 mm wide (1.4 inches), while width data for the L20 is not recorded. Both are dual-slot cards.

The L20 uses PCIe 4.0 x16, while the CMP 40HX uses PCIe 1.0 x4. Display outputs: the L20 has 4x DisplayPort 1.4a; the CMP 40HX has none. Release dates are also far apart: the L20 launched on 2023-11-15, while the CMP 40HX launched on 2021-02-24. The L20 lists a predecessor of Server Ampere and a successor of Server Hopper. The CMP 40HX lists neither. Production status is Active for the L20 and End-of-life for the CMP 40HX.

Where Each One Wins

The NVIDIA L20 wins in every recorded benchmark, so the use-case split is not about raw score leadership. The L20 is the clear choice for compute-heavy server workloads. Its 48 GB memory capacity and 864.0 GB/s bandwidth make it suitable for large models, rendering scenes, or data sets that require both high capacity and high throughput. The 1:1 FP16 ratio and 59.35 TFLOPS of FP32 compute make it a flexible workhorse for AI inference, scientific computing, and graphics workloads where precision matters. The 4x DisplayPort outputs also allow it to drive displays, so it can serve in visualization or workstation roles.

The NVIDIA CMP 40HX, despite losing all head-to-head tests, still has a niche. Its 185 W TDP and single 8-pin connector mean it is easier to power in systems with modest PSUs. Its 8 GB memory and 448.0 GB/s bandwidth are sufficient for lighter tasks, and its 93rd percentile ranking among all GPUs shows it is not a slouch. The lack of display outputs, however, means it is strictly a compute or mining card. Its PCIe 1.0 x4 interface is a serious limitation for any workload that streams data frequently, but for compute jobs where the data fits in local memory and the kernel runs long, that bottleneck is less impactful. The CMP 40HX is also an end-of-life product, so it is only relevant for existing deployments or the used market.

For a builder choosing between these two today, the L20 is the only sensible new purchase. The data shows a 193.7% lead in OpenCL and a 192.8% lead in Vulkan, along with superior memory, newer architecture, and active production status. The CMP 40HX only makes sense if the requirement is minimal power draw and the workload is already proven to fit within its 8 GB frame buffer and tolerate its slow host interface.

The Verdict

The benchmark data is unambiguous: the NVIDIA L20 outperforms the NVIDIA CMP 40HX by roughly 193% in both OpenCL and Vulkan. The L20 also holds a 99th percentile rank against the CMP 40HX's 93rd, and its nearest rivals are far more powerful cards. The L20 is a modern Ada Lovelace server GPU with 48 GB of memory, a 384-bit bus, 864.0 GB/s of bandwidth, and 59.35 TFLOPS of FP32 compute. The CMP 40HX is a Turing-based mining card with 8 GB of memory, a 256-bit bus, 448.0 GB/s of bandwidth, and 7.603 TFLOPS of FP32 compute.

Who should pick the L20? Anyone running AI inference, scientific simulation, large-scale rendering, or any workload that can use 48 GB of VRAM and high FP32 throughput. Its PCIe 4.0 x16 interface and display outputs also make it suitable for hybrid compute-visualization roles. The 275 W TDP is reasonable for a card of this class, and the suggested 600 W PSU is a common specification for modern workstations.

Who should pick the CMP 40HX? Only those with an existing deployment that already runs on this card, or a very specific need for a low-power (185 W) compute accelerator with no display output. Its PCIe 1.0 x4 interface is a severe bottleneck, and its end-of-life status means no future driver optimizations are guaranteed. The 8 GB memory is also a hard ceiling for modern data sets. The CMP 40HX is not a competitive choice against the L20 in any recorded metric. The verdict is straightforward: the L20 is the superior card in every measurable way, and the data supports that conclusion without qualification.

DETAILED SPECIFICATIONS

SPECIFICATION
CMP 40HX
L20
Core Specs
Shading Units
2,304
11,776 +411.1%
Shaders
2,304
11,776 +411.1%
TMUs
144
368 +155.6%
ROPs
64
128 +100.0%
SM Count
36
92 +155.6%
Clocks
Base Clock
1470 MHz
1440 MHz
Boost Clock
1650 MHz
2520 MHz
Memory Clock
1750 MHz 14 Gbps effective
2250 MHz 18 Gbps effective
Memory
Memory Size
8 GB
48 GB
VRAM (MB)
8,192
49,152 +500.0%
Memory Type
GDDR6
GDDR6
Memory Bus
256 bit
384 bit
Bandwidth
448.0 GB/s
864.0 GB/s
Cache
L1 Cache
64 KB (per SM)
128 KB (per SM)
L2 Cache
4 MB
96 MB
Performance
Pixel Rate
105.6 GPixel/s
322.6 GPixel/s
Texture Rate
237.6 GTexel/s
927.4 GTexel/s
FP32 (TFLOPS)
7.603 TFLOPS
59.35 TFLOPS
FP64 (TFLOPS)
237.6 GFLOPS (1:32)
927.4 GFLOPS (1:64)
FP16 (TFLOPS)
15.21 TFLOPS (2:1)
59.35 TFLOPS (1:1)
AI/RT
RT Cores
36
92 +155.6%
Tensor Cores
288
368 +27.8%
Power
TDP
185 W
275 W
TDP (W)
185
275 +48.6%
Suggested PSU
450 W
600 W
Power Connectors
1x 8-pin
1x 16-pin
Architecture
Architecture
Turing
Ada Lovelace
GPU Name
TU106
AD102
Generation
Mining GPUs
Server Ada (Lxx)
Process Size
12 nm
5 nm
Transistors
10,800 million
76,300 million
Die Size
445 mm²
609 mm²
Foundry
TSMC
TSMC
Density
24.3M / mm²
125.3M / mm²
API Support
DirectX
12 Ultimate (12_2)
12 Ultimate (12_2)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
3.0
3.0
CUDA
7.5
8.9
Shader Model
6.8
6.8
Physical
Slot Width
Dual-slot
Dual-slot
Length
229 mm 9 inches
267 mm 10.5 inches
Height
111 mm 4.4 inches
111 mm 4.4 inches
Outputs
No outputs
4x DisplayPort 1.4a
Bus Interface
PCIe 1.0 x4
PCIe 4.0 x16
Other
Launch Price
699 USD
Production
End-of-life
Active
Predecessor
Server Ampere
Successor
Server Hopper
View CMP 40HX Details View L20 Details