NVIDIA CMP 40HX vs NVIDIA L40S Comparison
NVIDIA CMP 40HX
L40S
PERFORMANCE BENCHMARKS
Analysis: NVIDIA CMP 40HX vs NVIDIA L40S
Head-to-Head Benchmarks
The recorded data shows a decisive performance gap between the NVIDIA L40S and the NVIDIA CMP 40HX. In the Geekbench OpenCL test, the L40S scores 330,727 points, while the CMP 40HX manages 93,395 points. That is a delta of 254.1% in favor of the L40S. In the Geekbench Vulkan test, the L40S scores 260,799 points against the CMP 40HX's 77,879 points, a 234.9% advantage. The L40S wins both head-to-head tests, giving it a clean 2-0 record.
These results place the L40S in the 99th percentile of all GPUs in the database, while the CMP 40HX sits in the 93rd percentile. The gap between the two is not marginal; it is a multiple of the CMP 40HX's entire output. For context, the L40S's average benchmark score of 295,763 is roughly 3.5 times the CMP 40HX's average of 85,637. The Vulkan test shows a slightly narrower margin than OpenCL, but the CMP 40HX still trails by more than a factor of three.
Looking at the L40S's nearest rivals, it sits 3% above the NVIDIA RTX 6000 Ada Generation (287,237 average) and 4.1% above the NVIDIA L40 (284,111 average). It trails the AMD Instinct MI300X by 7% (317,994 average) and the NVIDIA H200 NVL by 11.7% (334,891 average). The CMP 40HX, by contrast, is nearly tied with the AMD Radeon PRO W7600 (87,108 average, 1.7% behind) and the NVIDIA Quadro GP100 (87,445 average, 2.1% behind). It leads the AMD Radeon PRO W6600 by 4.4% (81,995 average) and the AMD Radeon Pro Vega 64X by 5.8% (80,959 average). The CMP 40HX's closest competition is therefore in the workstation midrange, while the L40S operates in an entirely different tier.
The Verdict
The data does not present a balanced comparison. The L40S is the clear choice for any workload that relies on OpenCL or Vulkan compute performance. Its 254.1% OpenCL lead and 234.9% Vulkan lead over the CMP 40HX mean that any task which is bottlenecked by raw throughput will favor the L40S overwhelmingly. The CMP 40HX, with its 93rd percentile standing, is competitive only within its own performance class, and even there it is not the top performer: it trails the AMD Radeon PRO W7600 and NVIDIA Quadro GP100 by small margins.
For users whose software depends on OpenCL or Vulkan acceleration, the L40S is the only defensible pick from this pair. The CMP 40HX's 8 GB memory and 448.0 GB/s bandwidth are dwarfed by the L40S's 48 GB and 864.0 GB/s. The CMP 40HX also lacks display outputs entirely, which limits its usefulness in any interactive or visualization context. The L40S, by contrast, includes one HDMI 2.1 port and three DisplayPort 1.4a outputs.
The CMP 40HX does have one advantage worth noting: its 185 W TDP is lower than the L40S's 300 W, and its 450 W suggested PSU is below the L40S's 700 W recommendation. It is also physically shorter at 229 mm compared to 267 mm. But these are operational differences, not performance compensations. The benchmark data shows no test where the CMP 40HX wins. For any compute-oriented purchase decision based on the recorded measurements, the L40S is the answer.
Architecture Differences
The two GPUs come from different architectural generations. The L40S uses the AD102 chip built on Ada Lovelace architecture, fabricated on a 5 nm process at TSMC. It contains 76,300 million transistors on a 609 mm² die, giving a transistor density of 125.3 million per mm². The CMP 40HX uses the TU106 chip with Turing architecture, also fabricated at TSMC but on a 12 nm process. It has 10,800 million transistors on a 445 mm² die, with a density of 24.3 million per mm². The L40S packs roughly seven times the transistor count into a die that is only about 37% larger, a direct result of the denser manufacturing node.
The L40S has 18,176 shading units, 568 texture mapping units, 192 ROPs, 142 RT cores, and 568 tensor cores. The CMP 40HX has 2,304 shading units, 144 TMUs, 64 ROPs, 36 RT cores, and 288 tensor cores. The L40S leads in every category, with the shading unit gap being particularly large. The FP32 compute rates reflect this: the L40S delivers 91.61 TFLOPS, while the CMP 40HX delivers 7.603 TFLOPS. The FP16 numbers also differ in ratio. The L40S achieves 91.61 TFLOPS at a 1:1 ratio, while the CMP 40HX achieves 15.21 TFLOPS at a 2:1 ratio.
Memory architecture differs as well. The L40S uses 48 GB of GDDR6 on a 384-bit bus, with 2250 MHz memory clock and 18 Gbps effective speed, yielding 864.0 GB/s of bandwidth. The CMP 40HX uses 8 GB of GDDR6 on a 256-bit bus, with 1750 MHz memory clock and 14 Gbps effective speed, yielding 448.0 GB/s. The L40S has six times the capacity and nearly double the bandwidth. Clock speeds tell a different story: the CMP 40HX has a higher base clock at 1470 MHz versus 1110 MHz, but the L40S boosts to 2520 MHz versus 1650 MHz.
The PCIe interface also differs substantially. The L40S uses PCIe 4.0 x16, while the CMP 40HX uses PCIe 1.0 x4. This is a significant bottleneck for the CMP 40HX in any data transfer scenario that involves host communication. The CMP 40HX has no display outputs, while the L40S has one HDMI 2.1 and three DisplayPort 1.4a connectors. Both support DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4, so API compatibility is not a differentiator.
The L40S was released on October 12, 2022, and is marked end-of-life. The CMP 40HX was released on February 24, 2021, and is also end-of-life. The L40S belongs to the Server Ada generation, while the CMP 40HX belongs to the Mining GPUs generation. The L40S lists its predecessor as Server Ampere and successor as Server Hopper; the CMP 40HX has no listed predecessor or successor.
FAQ
Q: Which GPU has the higher average benchmark score?
A: The NVIDIA L40S has an average benchmark score of 295,763, which is 210,126 points higher than the CMP 40HX's 85,637.
Q: How much faster is the L40S in the OpenCL test?
A: The L40S scores 330,727 in Geekbench OpenCL versus 93,395 for the CMP 40HX, a 254.1% advantage.
Q: Does the CMP 40HX have any display outputs?
A: No, the CMP 40HX has no display outputs. The L40S has one HDMI 2.1 port and three DisplayPort 1.4a outputs.
Q: What is the memory capacity difference?
A: The L40S has 48 GB of GDDR6 memory, while the CMP 40HX has 8 GB of GDDR6 memory. The L40S also has a wider 384-bit bus and 864.0 GB/s bandwidth versus 448.0 GB/s.
Q: What is the launch MSRP of the CMP 40HX?
A: The CMP 40HX has a launch MSRP of 699 USD. The L40S has no recorded launch MSRP in the database.
Q: How do the two GPUs compare in their closest rival groups?
A: The L40S sits 3% above the RTX 6000 Ada Generation and 4.1% above the L40, while trailing the Instinct MI300X by 7% and the H200 NVL by 11.7%. The CMP 40HX is 1.7% below the Radeon PRO W7600 and 2.1% below the Quadro GP100, but 4.4% above the Radeon PRO W6600 and 5.8% above the Radeon Pro Vega 64X.
Where Each One Wins
The L40S wins in all measured compute benchmarks, so the practical question is where each GPU fits based on the recorded characteristics. The L40S is suited for heavy compute workloads that can use its 48 GB memory capacity, 864.0 GB/s bandwidth, and 91.61 TFLOPS FP32 throughput. Its 142 RT cores and 568 tensor cores make it viable for ray tracing and tensor-based tasks. The presence of display outputs also allows for direct visualization, which the CMP 40HX cannot offer. The 99th percentile standing means it competes with the top tier of the database, including the H200 NVL and Instinct MI300X, against which it trails by 11.7% and 7% respectively.
The CMP 40HX, with its 93rd percentile standing, is positioned in the midrange. Its 8 GB memory and 448.0 GB/s bandwidth limit it to smaller datasets. The 185 W TDP and 450 W suggested PSU make it a lighter power draw, and its 229 mm length fits shorter chassis. The lack of display outputs means it is intended for compute-only installations. Its nearest rivals, the Radeon PRO W7600 and Quadro GP100, are all within a 2.1% band, so the CMP 40HX is not a standout in its own class. It is a functional, low-power compute card for tasks that fit within its memory ceiling, but the data gives it no advantage over the L40S in any measured test.
Specification Differences
The process node differs: 5 nm for the L40S versus 12 nm for the CMP 40HX. Transistor count is 76,300 million versus 10,800 million, and die size is 609 mm² versus 445 mm². Transistor density is 125.3M per mm² versus 24.3M per mm². Base clocks are 1110 MHz versus 1470 MHz, and boost clocks are 2520 MHz versus 1650 MHz. Memory clock is 2250 MHz versus 1750 MHz, with effective speeds of 18 Gbps versus 14 Gbps. Memory capacity is 48 GB versus 8 GB, bus width is 384-bit versus 256-bit, and bandwidth is 864.0 GB/s versus 448.0 GB/s.
Shading units are 18,176 versus 2,304. TMUs are 568 versus 144. ROPs are 192 versus 64. RT cores are 142 versus 36. Tensor cores are 568 versus 288. Pixel rate is 483.8 GPixel/s versus 105.6 GPixel/s. Texture rate is 1,431.4 GTexel/s versus 237.6 GTexel/s. FP32 is 91.61 TFLOPS versus 7.603 TFLOPS. FP16 is 91.61 TFLOPS (1:1) versus 15.21 TFLOPS (2:1). TDP is 300 W versus 185 W. Power connectors are one 16-pin versus one 8-pin. Suggested PSU is 700 W versus 450 W. Bus interface is PCIe 4.0 x16 versus PCIe 1.0 x4. Display outputs are one HDMI 2.1 and three DisplayPort 1.4a versus none. Length is 267 mm versus 229 mm. Height is 111 mm for both. The CMP 40HX has a listed width of 35 mm; the L40S has no recorded width. Release dates are October 12, 2022, versus February 24, 2021. The L40S has predecessor and successor entries; the CMP 40HX has neither. The CMP 40HX has a launch MSRP of 699 USD; the L40S has none recorded.