NVIDIA L20 vs NVIDIA L40S Comparison

NVIDIA
GEFORCE

NVIDIA L20

CORE STATE AD102
VRAM 48 GB
CLOCK SPEED 2520 MHz
TDP 275 W
BUS WIDTH 384 bit
ARCHITECTURE Ada Lovelace
nm
PROCESS 5 nm
LAUNCH DATE 2023
VS
NVIDIA
GEFORCE

L40S

CORE STATE AD102
VRAM 48 GB
CLOCK SPEED 2520 MHz
TDP 300 W
BUS WIDTH 384 bit
ARCHITECTURE Ada Lovelace
nm
PROCESS 5 nm
LAUNCH DATE 2022

PERFORMANCE BENCHMARKS

geekbench_opencl
274,276
330,727
geekbench_vulkan
228,018
260,799

Analysis: NVIDIA L20 vs NVIDIA L40S

The NVIDIA L40S and NVIDIA L20 are both server-grade Ada Lovelace GPUs built on the same AD102 chip and 5 nm TSMC process, but they are positioned very differently in performance and capability. The data shows a clear hierarchy: the L40S leads in every benchmark recorded, while the L20 is a more specialized option that trades raw compute for a different feature set. This analysis breaks down where each card wins, what the numbers mean for real workloads, and who should choose which.

Where Each One Wins

The L40S wins outright in both recorded benchmarks. In Geekbench OpenCL, it scores 330,727 against the L20’s 274,276, a 20.6% advantage. In Geekbench Vulkan, the L40S scores 260,799 versus 228,018, a 14.4% lead. These are not marginal differences; they represent a substantial gap in raw compute throughput. The L40S achieves this with 18,176 shading units, 568 tensor cores, and 91.61 TFLOPS of FP32 performance, while the L20 is configured with 11,776 shading units, 368 tensor cores, and 59.35 TFLOPS FP32. For any workload that scales with shader count or FP32 throughput — such as graphics rendering, simulation, or general compute — the L40S is the clear winner.

The L20’s wins are not in raw speed but in efficiency and power characteristics. It has a 275 W TDP versus the L40S’s 300 W, and its suggested PSU is 600 W versus 700 W. The L20 also boosts from a higher base clock of 1440 MHz versus 1110 MHz on the L40S, though both boost to 2520 MHz. This means the L20 can sustain a higher idle-to-boost ramp but ultimately delivers less peak performance. The L20 also features 4x DisplayPort 1.4a outputs, compared to the L40S’s 1x HDMI 2.1 and 3x DisplayPort 1.4a, making the L20 more flexible for multi-display configurations. Neither card has a launch MSRP listed, so no price comparison is possible from the data.

The Verdict

The data is unambiguous: the NVIDIA L40S is the better performer across every measured benchmark. It leads by 20.6% in OpenCL and 14.4% in Vulkan, with a higher average benchmark score of 295,763 versus 251,147. The L40S also sits in the 99th percentile of all GPUs, matching the L20’s percentile, but with a much higher absolute score. Its nearest rivals include the AMD Instinct MI300X (which outscores it by 7%) and the NVIDIA H200 NVL (which outscores it by 11.7%), but the L20’s rivals are all lower-tier: it beats the NVIDIA PG506-232 by 11.6% and the AMD Radeon PRO W7900D by 14.2%, but loses to the NVIDIA L40 by 11.6% and the RTX 6000 Ada Generation by 12.6%.

For buyers who need maximum compute throughput, the L40S is the obvious pick. It delivers 91.61 TFLOPS FP32 and 1,431.4 GTexel/s texture rate, which are the highest figures in this comparison. The L20, by contrast, is a lower-power part with the same memory configuration — 48 GB GDDR6 on a 384-bit bus with 864.0 GB/s bandwidth. If your workload is memory-bound and not shader-bound, the L20’s identical memory subsystem could be sufficient, but the benchmark data shows it still trails significantly in compute. The L20 is the choice only if you specifically need the extra DisplayPort outputs or the lower 275 W TDP, and you are willing to accept a 20.6% compute penalty.

Head-to-Head Benchmarks

The two recorded head-to-head benchmarks tell a consistent story. In Geekbench OpenCL, the L40S scores 330,727 against the L20’s 274,276. That 20.6% delta is the largest gap between the two cards in any test. This test is generally representative of general-purpose compute workloads, including physics simulation, data processing, and rendering tasks that rely on OpenCL. The L40S’s advantage here aligns with its 54% more shading units (18,176 vs 11,776) and 54% more tensor cores (568 vs 368). The FP32 throughput difference is 54% as well (91.61 vs 59.35 TFLOPS), though the actual benchmark delta is smaller, suggesting some workloads are not perfectly scaling with shader count.

In Geekbench Vulkan, the L40S scores 260,799 versus 228,018, a 14.4% lead. Vulkan is more graphics-oriented, and the L40S’s higher pixel rate (483.8 GPixel/s vs 322.6 GPixel/s) and texture rate (1,431.4 GTexel/s vs 927.4 GTexel/s) contribute to this win. The L40S also has 192 ROPs versus 128 on the L20, which directly impacts fill-rate-bound scenarios. The smaller delta in Vulkan compared to OpenCL suggests that the L20’s higher base clock (1440 MHz vs 1110 MHz) helps it stay closer in lighter workloads, but the L40S’s raw resource advantage still dominates. Both cards share the same boost clock of 2520 MHz, so the L40S’s superiority comes entirely from its wider configuration, not from higher frequency.

FAQ

Q: Which GPU has a higher average benchmark score?

A: The NVIDIA L40S has an average benchmark score of 295,763, while the NVIDIA L20 scores 251,147. That is a 44,616-point gap, or roughly 17.8% higher for the L40S.

Q: Do both cards have the same memory configuration?

A: Yes. Both the L40S and the L20 feature 48 GB of GDDR6 memory on a 384-bit bus with 864.0 GB/s bandwidth and 18 Gbps effective memory speed.

Q: How much faster is the L40S in Geekbench OpenCL?

A: The L40S scores 330,727 versus the L20’s 274,276, which is a 20.6% advantage. This is the largest delta between the two cards in any benchmark.

Q: Which card has a lower power draw?

A: The L20 has a 275 W TDP and a suggested PSU of 600 W, while the L40S has a 300 W TDP and a suggested PSU of 700 W. The L20 is the lower-power option.

Q: Are there any differences in display outputs?

A: Yes. The L20 has 4x DisplayPort 1.4a outputs, while the L40S has 1x HDMI 2.1 and 3x DisplayPort 1.4a. The L20 supports more simultaneous display connections.

Q: Which card has a higher base clock?

A: The L20 has a base clock of 1440 MHz, which is higher than the L40S’s 1110 MHz. Both cards share the same boost clock of 2520 MHz.

Architecture Differences

Both the L40S and L20 are built on the Ada Lovelace architecture using the same AD102 chip, fabricated on TSMC’s 5 nm process. They share the identical die size of 609 mm² and transistor count of 76,300 million, with a transistor density of 125.3M per mm². The architecture itself is identical — both support DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4. The key architectural difference is in how the AD102 chip is configured. The L40S enables 18,176 shading units, 568 TMUs, 192 ROPs, 142 RT cores, and 568 tensor cores. The L20 disables a significant portion of the chip, leaving 11,776 shading units, 368 TMUs, 128 ROPs, 92 RT cores, and 368 tensor cores. This is a 35% reduction in shader units and RT cores, and a 33% reduction in TMUs and tensor cores. The L20 is essentially a cut-down AD102, which explains its lower FP32 throughput of 59.35 TFLOPS versus 91.61 TFLOPS on the L40S. Both cards support FP16 at the same rate as FP32 (1:1 ratio), so the FP16 advantage for the L40S is proportionally identical.

Specification Differences

The specification differences between the two cards are substantial, though they share many fundamentals. Both are dual-slot cards measuring 267 mm in length and 111 mm in height, using a PCIe 4.0 x16 interface and a single 16-pin power connector. The process node, foundry, die size, and transistor count are identical. The memory subsystem is also the same: 48 GB GDDR6, 384-bit bus, 864.0 GB/s bandwidth, and 2250 MHz memory clock (18 Gbps effective). The boost clock is identical at 2520 MHz.

The differences start with the base clock: the L20 runs at 1440 MHz while the L40S runs at 1110 MHz. The L40S then pulls ahead in every compute metric — FP32 is 91.61 TFLOPS versus 59.35 TFLOPS, pixel rate is 483.8 GPixel/s versus 322.6 GPixel/s, and texture rate is 1,431.4 GTexel/s versus 927.4 GTexel/s. The L40S also has more of every execution unit: 18,176 shading units vs 11,776, 568 TMUs vs 368, 192 ROPs vs 128, 142 RT cores vs 92, and 568 tensor cores vs 368. The L20 draws 275 W versus the L40S’s 300 W, with a suggested PSU of 600 W versus 700 W. The display configuration differs as noted: the L20 has 4x DisplayPort 1.4a, while the L40S has 1x HDMI 2.1 and 3x DisplayPort 1.4a. Production status also differs — the L40S is end-of-life, while the L20 is still active. The L40S was released on 2022-10-12, while the L20 came later on 2023-11-15. Neither card has a launch MSRP listed in the data.

DETAILED SPECIFICATIONS

SPECIFICATION
L20
L40S
Core Specs
Shading Units
11,776
18,176 +54.3%
Shaders
11,776
18,176 +54.3%
TMUs
368
568 +54.3%
ROPs
128
192 +50.0%
SM Count
92
142 +54.3%
Clocks
Base Clock
1440 MHz
1110 MHz
Boost Clock
2520 MHz
2520 MHz
Memory Clock
2250 MHz 18 Gbps effective
2250 MHz 18 Gbps effective
Memory
Memory Size
48 GB
48 GB
VRAM (MB)
49,152
49,152 0.0%
Memory Type
GDDR6
GDDR6
Memory Bus
384 bit
384 bit
Bandwidth
864.0 GB/s
864.0 GB/s
Cache
L1 Cache
128 KB (per SM)
128 KB (per SM)
L2 Cache
96 MB
48 MB
Performance
Pixel Rate
322.6 GPixel/s
483.8 GPixel/s
Texture Rate
927.4 GTexel/s
1,431.4 GTexel/s
FP32 (TFLOPS)
59.35 TFLOPS
91.61 TFLOPS
FP64 (TFLOPS)
927.4 GFLOPS (1:64)
1,431.4 GFLOPS (1:64)
FP16 (TFLOPS)
59.35 TFLOPS (1:1)
91.61 TFLOPS (1:1)
AI/RT
RT Cores
92
142 +54.3%
Tensor Cores
368
568 +54.3%
Power
TDP
275 W
300 W
TDP (W)
275
300 +9.1%
Suggested PSU
600 W
700 W
Power Connectors
1x 16-pin
1x 16-pin
Architecture
Architecture
Ada Lovelace
Ada Lovelace
GPU Name
AD102
AD102
Generation
Server Ada (Lxx)
Server Ada (Lxx)
Process Size
5 nm
5 nm
Transistors
76,300 million
76,300 million
Die Size
609 mm²
609 mm²
Foundry
TSMC
TSMC
Density
125.3M / mm²
125.3M / mm²
API Support
DirectX
12 Ultimate (12_2)
12 Ultimate (12_2)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
3.0
3.0
CUDA
8.9
8.9
Shader Model
6.8
6.8
Physical
Slot Width
Dual-slot
Dual-slot
Length
267 mm 10.5 inches
267 mm 10.5 inches
Height
111 mm 4.4 inches
111 mm 4.4 inches
Outputs
4x DisplayPort 1.4a
1x HDMI 2.13x DisplayPort 1.4a
Bus Interface
PCIe 4.0 x16
PCIe 4.0 x16
Other
Production
Active
End-of-life
Predecessor
Server Ampere
Server Ampere
Successor
Server Hopper
Server Hopper
View L20 Details View L40S Details