NVIDIA CMP 90HX vs NVIDIA L4 Comparison

NVIDIA
GEFORCE

NVIDIA CMP 90HX

CORE STATE GA102
VRAM 10 GB
CLOCK SPEED 1710 MHz
TDP 320 W
BUS WIDTH 320 bit
ARCHITECTURE Ampere
nm
PROCESS 8 nm
LAUNCH DATE 2021
VS
NVIDIA
GEFORCE

L4

CORE STATE AD104
VRAM 24 GB
CLOCK SPEED 2040 MHz
TDP 72 W
BUS WIDTH 192 bit
ARCHITECTURE Ada Lovelace
nm
PROCESS 5 nm
LAUNCH DATE 2023

PERFORMANCE BENCHMARKS

geekbench_opencl
69,000
140,838
geekbench_vulkan
N/A
121,306

Analysis: NVIDIA CMP 90HX vs NVIDIA L4

FAQ

Q: What are the average benchmark scores for the NVIDIA L4 and the NVIDIA CMP 90HX?

A: The NVIDIA L4 has an average benchmark score of 131,072, while the NVIDIA CMP 90HX has an average benchmark score of 69,000. The L4 sits in the 95th percentile of all GPUs, whereas the CMP 90HX is in the 90th percentile.

Q: Which GPU has a higher memory bandwidth, and what is the difference?

A: The NVIDIA CMP 90HX has a significantly higher memory bandwidth at 760.3 GB/s, compared to the L4's 300.1 GB/s. This is largely due to the CMP 90HX using a 320-bit memory bus with GDDR6X memory, while the L4 uses a 192-bit bus with GDDR6.

Q: How do the two GPUs compare in terms of power consumption?

A: The NVIDIA L4 has a TDP of 72 W, making it a low-power solution with a suggested PSU of 250 W. In contrast, the CMP 90HX has a TDP of 320 W and requires a suggested PSU of 700 W, a substantial difference in power requirements.

Q: Are these GPUs suitable for display output?

A: No, neither GPU has display outputs. Both the NVIDIA L4 and the NVIDIA CMP 90HX are designed for compute or mining workloads, not for driving displays directly.

Q: What is the architectural generation for each GPU?

A: The NVIDIA L4 is based on the Ada Lovelace architecture (chip AD104) and belongs to the Server Ada (Lxx) generation. The NVIDIA CMP 90HX is based on the Ampere architecture (chip GA102) and belongs to the Mining GPUs generation.

Q: What is the production status of each card?

A: The NVIDIA L4 is listed as "Active" in production, while the NVIDIA CMP 90HX is marked as "End-of-life". The L4 was released in March 2023, whereas the CMP 90HX was released in July 2021.

Architecture Differences

The NVIDIA L4 and the NVIDIA CMP 90HX represent two distinct architectural eras from NVIDIA. The L4 is built on the Ada Lovelace architecture, utilizing the AD104 chip manufactured on a 5 nm process at TSMC. This process node allows for a transistor density of 121.8 million transistors per square millimeter, totaling 35,800 million transistors on a 294 mm² die. The CMP 90HX, conversely, uses the older Ampere architecture with the GA102 chip, fabricated on Samsung's 8 nm process. This results in a transistor density of only 45.1 million per square millimeter, with 28,300 million transistors spread across a much larger 628 mm² die.

The difference in transistor density is a key factor in their capabilities. The L4's denser 5 nm process enables more compute units per area, contributing to higher shading unit counts: 7,424 shading units, 240 texture mapping units, and 80 ROPs. The CMP 90HX has 6,400 shading units, 200 TMUs, and 80 ROPs. Both cards feature ray tracing and tensor cores, with the L4 sporting 60 RT cores and 240 tensor cores, while the CMP 90HX has 50 RT cores and 200 tensor cores.

Clock speeds also differ significantly. The L4 has a base clock of 795 MHz and a boost clock of 2040 MHz, while the CMP 90HX runs at a base of 1500 MHz and a boost of 1710 MHz. Despite the L4's lower base clock, its higher boost clock and newer architecture allow it to achieve higher FP32 performance at 30.29 TFLOPS compared to 21.89 TFLOPS for the CMP 90HX. Both support DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4.

Another notable architectural difference is the memory subsystem. The L4 uses 24 GB of GDDR6 memory on a 192-bit bus, with a memory clock of 1563 MHz (12.5 Gbps effective), yielding 300.1 GB/s bandwidth. The CMP 90HX uses 10 GB of GDDR6X on a 320-bit bus, with a memory clock of 1188 MHz (19 Gbps effective), resulting in a much higher 760.3 GB/s bandwidth. This difference is crucial for memory-intensive workloads, though the L4's larger capacity may benefit certain large dataset applications.

Where Each One Wins

The benchmark data reveals a clear split in use cases between these two GPUs, based on their design philosophies and the single head-to-head benchmark available.

The NVIDIA L4 wins decisively in the only shared benchmark test, Geekbench OpenCL, with a score of 140,838 compared to 69,000 for the CMP 90HX. This represents a 104.1% advantage. The L4's architecture, with its higher shading unit count (7,424 vs 6,400) and newer Ada Lovelace design, makes it the superior choice for general compute workloads, AI inference, and tasks that benefit from FP32 throughput. Its 24 GB VRAM capacity is also a significant advantage for large models or datasets that exceed the CMP 90HX's 10 GB limit.

The CMP 90HX, on the other hand, shows its strength in memory bandwidth. Its 760.3 GB/s bandwidth is more than double the L4's 300.1 GB/s, making it potentially more suitable for workloads that are highly sensitive to memory throughput, such as certain mining algorithms or specific types of data processing that access memory sequentially. However, the data shows that in the OpenCL benchmark, this bandwidth advantage did not translate to a win, suggesting that compute efficiency matters more in general-purpose tasks.

For power-constrained environments, the L4 is the clear winner with its 72 W TDP versus the CMP 90HX's 320 W. The L4 offers a much more efficient solution, requiring only a 250 W suggested PSU compared to 700 W for the CMP 90HX. This makes the L4 easier to integrate into dense server environments where power and cooling are limited.

Specification Differences

The two GPUs differ across nearly every major specification category. The most striking contrast is in power consumption: the L4 consumes 72 W, while the CMP 90HX draws 320 W. This leads to different physical requirements, with the L4 being a single-slot card that requires no power connectors, while the CMP 90HX is a dual-slot card with two 8-pin power connectors.

Memory configuration is another major divergence. The L4 offers 24 GB of GDDR6 on a 192-bit bus, while the CMP 90HX provides 10 GB of GDDR6X on a 320-bit bus. The bandwidth figures reflect this: 300.1 GB/s for the L4 versus 760.3 GB/s for the CMP 90HX. The effective memory speeds also differ, with the L4 running at 1563 MHz (12.5 Gbps effective) and the CMP 90HX at 1188 MHz (19 Gbps effective).

The compute specifications show the L4's advantage in raw throughput. It has 7,424 shading units, 240 TMUs, 60 RT cores, and 240 tensor cores, versus 6,400 shading units, 200 TMUs, 50 RT cores, and 200 tensor cores for the CMP 90HX. The L4 also achieves higher pixel and texture rates at 163.2 GPixel/s and 489.6 GTexel/s, respectively, compared to 136.8 GPixel/s and 342.0 GTexel/s for the CMP 90HX. The FP32 performance is 30.29 TFLOPS for the L4 and 21.89 TFLOPS for the CMP 90HX, with both offering 1:1 FP16 to FP32 ratios.

Physical dimensions differ substantially as well. The L4 measures 169 mm in length and 56 mm in height, while the CMP 90HX is much larger at 285 mm in length and 112 mm in height. The bus interface also differs: the L4 uses PCIe 4.0 x16, while the CMP 90HX uses PCIe 1.0 x4, which is a significant bottleneck for data transfer in the latter. Both cards have no display outputs and differ in production status, with the L4 being active and the CMP 90HX end-of-life.

Head-to-Head Benchmarks

The only direct benchmark comparison in the database is the Geekbench OpenCL test, and it shows a dominant performance by the NVIDIA L4. The L4 scored 140,838 in this test, while the CMP 90HX managed 69,000. This results in a delta percentage of 104.1%, meaning the L4 is more than twice as fast as the CMP 90HX in this specific compute benchmark.

This performance gap is notable given the CMP 90HX's superior memory bandwidth. The CMP 90HX's 760.3 GB/s bandwidth should theoretically benefit memory-intensive workloads, but the OpenCL test appears to favor the L4's higher compute throughput and newer architecture. The L4's 30.29 TFLOPS FP32 performance versus 21.89 TFLOPS for the CMP 90HX likely plays a significant role, as OpenCL often stress-tests arithmetic and logic operations.

Looking at the nearest rivals provides context for these scores. The L4's average score of 131,072 places it within 0.7% of the NVIDIA GeForce RTX 3090 Ti (131,938) and 3.1% behind the NVIDIA RTX 4000 Ada Generation (135,218) and NVIDIA A10M (135,230). This positions the L4 as a high-end compute solution, despite its low 72 W TDP. The CMP 90HX's average score of 69,000 is nearly identical to the Intel Arc A770 (68,809, a 0.3% difference) and the AMD Radeon Instinct MI25 (68,562, a 0.6% difference). This places the CMP 90HX in a much lower performance tier, closer to mid-range or older workstation cards.

The percentile rankings reinforce this gap. The L4 sits in the 95th percentile of all GPUs, while the CMP 90HX is in the 90th percentile. Although both are above average, the L4's performance is significantly higher in absolute terms, and the gap in the OpenCL test is substantial.

The Verdict

Based on the recorded data, the NVIDIA L4 is the superior choice for nearly all compute-intensive applications. Its Geekbench OpenCL score of 140,838 is more than double that of the CMP 90HX's 69,000, and its 30.29 TFLOPS FP32 performance outpaces the CMP 90HX's 21.89 TFLOPS. The L4 also offers more VRAM (24 GB vs 10 GB), a higher shading unit count (7,424 vs 6,400), and a significantly lower power draw (72 W vs 320 W). Its active production status and newer architecture make it a more future-proof investment.

The CMP 90HX, however, is not without its strengths. Its 760.3 GB/s memory bandwidth is a clear advantage over the L4's 300.1 GB/s, which could be beneficial for specific memory-bound workloads. Its larger physical size and dual-slot design may also allow for better cooling in some chassis. However, its PCIe 1.0 x4 interface is a severe limitation for data transfer, and its end-of-life status suggests it is not a viable long-term solution.

For users prioritizing compute performance, power efficiency, and modern architecture, the NVIDIA L4 is the clear choice based on the data. For those with workloads that are uniquely sensitive to memory bandwidth and can tolerate higher power consumption and an older platform, the CMP 90HX remains a niche option. Overall, the benchmark results indicate that the L4 is the more capable and versatile GPU, while the CMP 90HX is a specialized product that has been surpassed by newer designs.

DETAILED SPECIFICATIONS

SPECIFICATION
CMP 90HX
L4
Core Specs
Shading Units
6,400
7,424 +16.0%
Shaders
6,400
7,424 +16.0%
TMUs
200
240 +20.0%
ROPs
80
80 0.0%
SM Count
50
60 +20.0%
Clocks
Base Clock
1500 MHz
795 MHz
Boost Clock
1710 MHz
2040 MHz
Memory Clock
1188 MHz 19 Gbps effective
1563 MHz 12.5 Gbps effective
Memory
Memory Size
10 GB
24 GB
VRAM (MB)
10,240
24,576 +140.0%
Memory Type
GDDR6X
GDDR6
Memory Bus
320 bit
192 bit
Bandwidth
760.3 GB/s
300.1 GB/s
Cache
L1 Cache
128 KB (per SM)
128 KB (per SM)
L2 Cache
5 MB
48 MB
Performance
Pixel Rate
136.8 GPixel/s
163.2 GPixel/s
Texture Rate
342.0 GTexel/s
489.6 GTexel/s
FP32 (TFLOPS)
21.89 TFLOPS
30.29 TFLOPS
FP64 (TFLOPS)
342.0 GFLOPS (1:64)
473.3 GFLOPS (1:64)
FP16 (TFLOPS)
21.89 TFLOPS (1:1)
30.29 TFLOPS (1:1)
AI/RT
RT Cores
50
60 +20.0%
Tensor Cores
200
240 +20.0%
Power
TDP
320 W
72 W
TDP (W)
320
72 -77.5%
Suggested PSU
700 W
250 W
Power Connectors
2x 8-pin
None
Architecture
Architecture
Ampere
Ada Lovelace
GPU Name
GA102
AD104
Generation
Mining GPUs
Server Ada (Lxx)
Process Size
8 nm
5 nm
Transistors
28,300 million
35,800 million
Die Size
628 mm²
294 mm²
Foundry
Samsung
TSMC
Density
45.1M / mm²
121.8M / mm²
API Support
DirectX
12 Ultimate (12_2)
12 Ultimate (12_2)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
3.0
3.0
CUDA
8.6
8.9
Shader Model
6.8
6.8
Physical
Slot Width
Dual-slot
Single-slot
Length
285 mm 11.2 inches
169 mm 6.7 inches
Height
112 mm 4.4 inches
56 mm 2.2 inches
Outputs
No outputs
No outputs
Bus Interface
PCIe 1.0 x4
PCIe 4.0 x16
Other
Production
End-of-life
Active
Predecessor
Server Ampere
Successor
Server Hopper
View CMP 90HX Details View L4 Details