GPU Comparison

NVIDIA
GEFORCE

NVIDIA L40

CORE STATE AD102
VRAM 48 GB
CLOCK SPEED 2490 MHz
TDP 300 W
BUS WIDTH 384 bit
ARCHITECTURE Ada Lovelace
nm
PROCESS 5 nm
LAUNCH DATE 2022
VS
NVIDIA
GEFORCE

L40S

CORE STATE AD102
VRAM 48 GB
CLOCK SPEED 2520 MHz
TDP 300 W
BUS WIDTH 384 bit
ARCHITECTURE Ada Lovelace
nm
PROCESS 5 nm
LAUNCH DATE 2022

PERFORMANCE BENCHMARKS

geekbench_opencl
330,926
330,727
geekbench_vulkan
237,295
260,799

Analysis: NVIDIA L40 vs NVIDIA L40S

The NVIDIA L40S and NVIDIA L40 are both server-class Ada Lovelace GPUs built on the same AD102 chip at TSMC's 5 nm node. They share the same 48 GB GDDR6 memory configuration, 384-bit bus, and 864.0 GB/s bandwidth, but the L40S runs higher core clocks and posts faster benchmark scores across the board. The data below breaks down where they differ, which one wins each test, and what that means for a builder choosing between them.

Head-to-Head Benchmarks

The head-to-head results are unambiguous: the L40S wins both benchmark tests, taking 2 wins to 0. In Geekbench OpenCL, the L40S scores 334,437 against the L40's 330,683, a 1.1% lead. The gap widens substantially in Geekbench Vulkan, where the L40S posts 250,769 versus 232,627, a 7.8% advantage. That Vulkan result is the single largest performance difference between the two cards, suggesting the L40S has a more meaningful edge in compute workloads that exercise the graphics and compute pipeline together.

Looking at the broader average benchmark scores, the L40S averages 292,603 across its tests, while the L40 averages 281,655. That places the L40S 3.9% ahead of the L40 according to the L40S's own rival listing, and 3.7% behind the L40S from the L40's perspective, the small rounding difference comes from how each card's delta is calculated. Against the wider field, the L40S sits at the 100th percentile of all GPUs, while the L40 is at the 99th. The L40S also leads the RTX 6000 Ada Generation (281,932) by 3.8% and the L20 (266,428) by 9.8%, but trails the H200 NVL (305,608) by 4.3%. The L40, by comparison, is effectively tied with the RTX 6000 Ada at 0.1% behind, leads the L20 by 5.7%, and trails the H200 NVL by 7.8% and the L40S by 3.7%.

What this means in practice: the L40S is the stronger card in every measured benchmark, and the gap is not trivial in Vulkan. The L40 is still a fast GPU in absolute terms, it beats the L20 by 5.7% and is within a hair of the RTX 6000 Ada, but it cannot match the L40S's clock-for-clock performance advantage.

Architecture Differences

Both cards are built on the same AD102 chip, the same Ada Lovelace architecture, and the same TSMC 5 nm process. The transistor count is identical at 76,300 million, the die size is the same 609 mm², and the transistor density is 125.3M per mm². All the core compute resources match exactly: 18,176 shading units, 568 TMUs, 192 ROPs, 142 RT cores, and 568 tensor cores. The memory subsystem is also identical, 48 GB of GDDR6 on a 384-bit bus, with a memory clock of 2250 MHz (18 Gbps effective) and 864.0 GB/s of bandwidth.

The only architectural difference is in clock speeds. The L40S has a base clock of 1110 MHz and a boost clock of 2520 MHz, while the L40 runs a base clock of 735 MHz and a boost of 2490 MHz. That 375 MHz base-clock gap and 30 MHz boost-clock gap account for the entire performance delta between the two cards. Because the core counts and memory are identical, the L40S's higher clocks translate directly into higher pixel, texture, and compute rates. The L40S achieves 483.8 GPixel/s and 1,431.4 GTexel/s, while the L40 manages 478.1 GPixel/s and 1,414.3 GTexel/s. In FP32 and FP16, the L40S delivers 91.61 TFLOPS (1:1) versus the L40's 90.52 TFLOPS (1:1).

Both cards use the same dual-slot cooler, the same 1x 16-pin power connector, and have the same 300 W TDP with a 700 W suggested PSU. They also share the same PCIe 4.0 x16 bus interface and identical physical dimensions: 267 mm (10.5 inches) long and 111 mm (4.4 inches) high. The display outputs differ, the L40S has 1x HDMI 2.1 and 3x DisplayPort 1.4a, while the L40 has 4x DisplayPort 1.4a, but both are end-of-life products released on the same date (2022-10-12), with the same predecessor (Server Ampere) and successor (Server Hopper).

Where Each One Wins

The L40S wins every benchmark in the head-to-head comparison, so the "where each one wins" section is largely a one-sided story. In Geekbench OpenCL, the L40S leads by 1.1%, and in Geekbench Vulkan it leads by 7.8%. If you are choosing based purely on compute or graphics performance, the L40S is the clear pick. Its 100th-percentile ranking versus all GPUs, combined with a 3.9% average-score lead over the L40, makes it the stronger accelerator for any workload that shows up in these tests.

The L40 does have one tangible advantage: its display output configuration. With 4x DisplayPort 1.4a, the L40 can drive four DisplayPort monitors directly, whereas the L40S offers only 1x HDMI 2.1 and 3x DisplayPort 1.4a. For a workstation that relies on four DisplayPort connections without an HDMI adapter, the L40 is the better fit. That is a niche use case, but it is a real one. The L40 also has a lower base clock (735 MHz vs 1110 MHz), which could theoretically result in different idle behavior, but the data does not include power or thermal measurements, so that cannot be confirmed.

Beyond those points, the L40's only other distinction is that it sits at the 99th percentile rather than the 100th, and it trails the L40S by 3.7% in average score. It is not a weak card, it beats the L20 by 5.7% and is essentially tied with the RTX 6000 Ada, but it is consistently slower than the L40S in every measured benchmark.

Specification Differences

The two cards differ in the following fields:

  • Base clock: L40S 1110 MHz vs L40 735 MHz
  • Boost clock: L40S 2520 MHz vs L40 2490 MHz
  • Pixel rate: L40S 483.8 GPixel/s vs L40 478.1 GPixel/s
  • Texture rate: L40S 1,431.4 GTexel/s vs L40 1,414.3 GTexel/s
  • FP32: L40S 91.61 TFLOPS vs L40 90.52 TFLOPS
  • FP16: L40S 91.61 TFLOPS (1:1) vs L40 90.52 TFLOPS (1:1)
  • Display outputs: L40S 1x HDMI 2.1 + 3x DisplayPort 1.4a vs L40 4x DisplayPort 1.4a
  • Average benchmark score: L40S 292,603 vs L40 281,655
  • Percentile vs all GPUs: L40S 100 vs L40 99
  • Head-to-head wins: L40S 2 vs L40 0

Everything else is identical: memory size, type, bus width, bandwidth, shading units, TMUs, ROPs, RT cores, tensor cores, TDP, slot width, power connector, suggested PSU, bus interface, dimensions, process node, foundry, transistor count, die size, and release date.

FAQ

Q: Which GPU is faster in the head-to-head benchmarks?

A: The L40S wins both tests. It scores 334,437 in Geekbench OpenCL (1.1% ahead of the L40's 330,683) and 250,769 in Geekbench Vulkan (7.8% ahead of the L40's 232,627). That gives the L40S 2 wins to 0.

Q: Do the two cards have the same memory configuration?

A: Yes. Both have 48 GB of GDDR6 on a 384-bit bus, with a memory clock of 2250 MHz (18 Gbps effective) and 864.0 GB/s of bandwidth.

Q: What clock speeds do they run?

A: The L40S has a base clock of 1110 MHz and a boost clock of 2520 MHz. The L40 has a base clock of 735 MHz and a boost clock of 2490 MHz. The memory clock is the same on both at 2250 MHz.

Q: Are they built on the same architecture?

A: Yes. Both use the AD102 chip on Ada Lovelace, fabricated by TSMC at 5 nm, with 76,300 million transistors on a 609 mm² die. The core counts are identical: 18,176 shading units, 568 TMUs, 192 ROPs, 142 RT cores, and 568 tensor cores.

Q: How do they compare to other GPUs in the same class?

A: The L40S average score of 292,603 is 3.8% above the RTX 6000 Ada Generation (281,932) and 9.8% above the L20 (266,428), but 4.3% below the H200 NVL (305,608). The L40 average score of 281,655 is 5.7% above the L20 but 7.8% below the H200 NVL and 3.7% below the L40S.

Q: Which card has more display outputs?

A: The L40 has 4x DisplayPort 1.4a. The L40S has 1x HDMI 2.1 and 3x DisplayPort 1.4a, so it offers one fewer DisplayPort connection but adds an HDMI port.

The Verdict

The data is clear: the L40S is the better performer in every measured benchmark. It wins both Geekbench OpenCL and Vulkan, holds a higher average score (292,603 vs 281,655), and sits at the 100th percentile of all GPUs versus the L40's 99th. The 7.8% Vulkan lead is the standout difference, and even the smaller 1.1% OpenCL edge is consistent. For any compute workload that relies on OpenCL or Vulkan, the L40S is the correct choice.

The L40 is not a bad card, it beats the L20 by 5.7% and is within 0.1% of the RTX 6000 Ada, but it is slower than the L40S in every category that matters for raw performance. Its only concrete advantage is the 4x DisplayPort 1.4a output, which matters if you need to connect four DisplayPort monitors without an HDMI adapter. If that is not a requirement, the L40S's higher clocks and faster scores make it the straightforward pick. Both cards are end-of-life, so availability and system compatibility will ultimately drive the decision, but on the numbers alone, the L40S wins.

DETAILED SPECIFICATIONS

SPECIFICATION
L40
L40S
Core Specs
Shading Units
18,176
18,176 0.0%
Shaders
18,176
18,176 0.0%
TMUs
568
568 0.0%
ROPs
192
192 0.0%
SM Count
142
142 0.0%
Clocks
Base Clock
735 MHz
1110 MHz
Boost Clock
2490 MHz
2520 MHz
Memory Clock
2250 MHz 18 Gbps effective
2250 MHz 18 Gbps effective
Memory
Memory Size
48 GB
48 GB
VRAM (MB)
49,152
49,152 0.0%
Memory Type
GDDR6
GDDR6
Memory Bus
384 bit
384 bit
Bandwidth
864.0 GB/s
864.0 GB/s
Cache
L1 Cache
128 KB (per SM)
128 KB (per SM)
L2 Cache
96 MB
48 MB
Performance
Pixel Rate
478.1 GPixel/s
483.8 GPixel/s
Texture Rate
1,414.3 GTexel/s
1,431.4 GTexel/s
FP32 (TFLOPS)
90.52 TFLOPS
91.61 TFLOPS
FP64 (TFLOPS)
1,414.3 GFLOPS (1:64)
1,431.4 GFLOPS (1:64)
FP16 (TFLOPS)
90.52 TFLOPS (1:1)
91.61 TFLOPS (1:1)
AI/RT
RT Cores
142
142 0.0%
Tensor Cores
568
568 0.0%
Power
TDP
300 W
300 W
TDP (W)
300
300 0.0%
Suggested PSU
700 W
700 W
Power Connectors
1x 16-pin
1x 16-pin
Architecture
Architecture
Ada Lovelace
Ada Lovelace
GPU Name
AD102
AD102
Generation
Server Ada (Lxx)
Server Ada (Lxx)
Process Size
5 nm
5 nm
Transistors
76,300 million
76,300 million
Die Size
609 mm²
609 mm²
Foundry
TSMC
TSMC
Density
125.3M / mm²
125.3M / mm²
API Support
DirectX
12 Ultimate (12_2)
12 Ultimate (12_2)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
3.0
3.0
CUDA
8.9
8.9
Shader Model
6.8
6.8
Physical
Slot Width
Dual-slot
Dual-slot
Length
267 mm 10.5 inches
267 mm 10.5 inches
Height
111 mm 4.4 inches
111 mm 4.4 inches
Outputs
4x DisplayPort 1.4a
1x HDMI 2.13x DisplayPort 1.4a
Bus Interface
PCIe 4.0 x16
PCIe 4.0 x16
Other
Production
End-of-life
End-of-life
Predecessor
Server Ampere
Server Ampere
Successor
Server Hopper
Server Hopper
View L40 Details View L40S Details