NVIDIA L20 vs NVIDIA L40 Comparison

NVIDIA
GEFORCE

NVIDIA L20

CORE STATE AD102
VRAM 48 GB
CLOCK SPEED 2520 MHz
TDP 275 W
BUS WIDTH 384 bit
ARCHITECTURE Ada Lovelace
nm
PROCESS 5 nm
LAUNCH DATE 2023
VS
NVIDIA
GEFORCE

L40

CORE STATE AD102
VRAM 48 GB
CLOCK SPEED 2490 MHz
TDP 300 W
BUS WIDTH 384 bit
ARCHITECTURE Ada Lovelace
nm
PROCESS 5 nm
LAUNCH DATE 2022

PERFORMANCE BENCHMARKS

geekbench_opencl
274,276
330,926
geekbench_vulkan
228,018
237,295

Analysis: NVIDIA L20 vs NVIDIA L40

The NVIDIA L40 and NVIDIA L20 are both server-grade accelerators built on the Ada Lovelace architecture, sharing the same AD102 chip, a 5 nm TSMC process, and an identical 48 GB GDDR6 memory configuration with a 384-bit bus and 864.0 GB/s of bandwidth. Despite these fundamental similarities, the benchmark data reveals a clear performance hierarchy, with the L40 holding a substantial lead in compute throughput. This analysis walks through the head-to-head results, architectural differences, and use-case implications strictly from the provided data.

Head-to-Head Benchmarks

The only direct comparison available in the data is the Geekbench OpenCL test, where the NVIDIA L40 scores 330,683 against the L20’s 266,428. This translates to a 24.1% delta in favor of the L40, a decisive margin that underscores the L40’s superior raw compute capability. In the broader context of the L40’s nearest rivals, this score places it 5.7% ahead of the L20, while the L20’s own data shows it trailing the L40 by 5.4% from its perspective. These reciprocal deltaPct values confirm the consistency of the measurement.

The L40 also posts a Geekbench Vulkan score of 232,627, a metric not recorded for the L20, which further highlights its advantage in graphics-adjacent workloads. The average benchmark score for the L40 is 281,655, compared to 266,428 for the L20. This 15,227-point gap in average scores reinforces the OpenCL result, indicating that the L40’s advantage is not isolated to a single test but reflects a general performance lead. Looking at the L40’s position among its nearest rivals, it sits just 0.1% below the NVIDIA RTX 6000 Ada Generation (281,932) and 3.7% below the NVIDIA L40S (292,603), while remaining 7.8% behind the NVIDIA H200 NVL (305,608). The L20, by contrast, trails the RTX 6000 Ada Generation by 5.5%, the L40S by 8.9%, and the H200 NVL by 12.8%.

These numbers indicate that the L40 occupies a higher performance tier than the L20, with the 24.1% delta in the head-to-head test being the single most important data point. The L20’s performance, while still in the 99th percentile of all GPUs, is measurably closer to the lower end of that elite group, whereas the L40’s scores push it toward the upper boundary. The data shows no benchmark in which the L20 wins; the L40 wins the sole head-to-head test and holds a decisive advantage in average score.

Architecture Differences

Both cards share the same foundational architecture: Ada Lovelace, built on the AD102 chip, manufactured by TSMC on a 5 nm process with 76,300 million transistors and a die size of 609 mm². The transistor density of 125.3M / mm² is identical, as are the memory specifications—48 GB GDDR6, 384-bit bus, 864.0 GB/s bandwidth, and 18 Gbps effective memory clock. The core configuration, however, diverges significantly.

The L40 features 18,176 shading units, 568 texture mapping units (TMUs), and 192 raster operation units (ROPs). It also carries 142 ray tracing cores and 568 tensor cores. The L20, by contrast, is equipped with 11,776 shading units, 368 TMUs, 128 ROPs, 92 ray tracing cores, and 368 tensor cores. This represents a substantial reduction in every compute unit category for the L20—roughly 35% fewer shading units, TMUs, and tensor cores, and about 35% fewer ray tracing cores as well.

Clock speeds tell a complementary story. The L20 has a higher base clock of 1440 MHz compared to the L40’s 735 MHz, and a slightly higher boost clock of 2520 MHz versus 2490 MHz. Despite this clock advantage, the L20 cannot compensate for its reduced core count. The pixel rate for the L40 is 478.1 GPixel/s versus 322.6 GPixel/s for the L20, and the texture rate is 1,414.3 GTexel/s versus 927.4 GTexel/s. The FP32 and FP16 throughput figures are equally lopsided: the L40 delivers 90.52 TFLOPS in both precision modes (1:1 ratio), while the L20 delivers 59.35 TFLOPS in both.

The power envelope differs modestly: the L40 has a 300 W TDP with a suggested PSU of 700 W, while the L20 draws 275 W with a 600 W suggested PSU. Both are dual-slot cards, use a single 16-pin power connector, and share identical physical dimensions of 267 mm length and 111 mm height. The production status differs—the L40 is end-of-life, while the L20 remains active. The L40 was released on 2022-10-12, and the L20 on 2023-11-15. Both share the same generation (Server Ada Lxx), predecessor (Server Ampere), successor (Server Hopper), and API support (DirectX 12 Ultimate 12_2, OpenGL 4.6, Vulkan 1.4).

FAQ

Q: Which GPU has a higher FP32 compute throughput?

A: The NVIDIA L40 delivers 90.52 TFLOPS in FP32, which is significantly higher than the L20’s 59.35 TFLOPS. This represents a 24.1% advantage in the head-to-head Geekbench OpenCL test, consistent with the raw compute specification difference.

Q: Do both cards have the same memory configuration?

A: Yes. Both the L40 and L20 feature 48 GB of GDDR6 memory on a 384-bit bus, with 864.0 GB/s of bandwidth and an 18 Gbps effective memory clock. Memory is not a differentiating factor between these two accelerators.

Q: Why does the L20 have a higher base clock but lower performance?

A: The L20 runs a 1440 MHz base clock and 2520 MHz boost clock, compared to the L40’s 735 MHz base and 2490 MHz boost. However, the L40 has 18,176 shading units versus 11,776 on the L20, along with more TMUs, ROPs, ray tracing cores, and tensor cores. The core count deficit overwhelms the clock speed advantage, resulting in lower overall throughput.

Q: What is the production status of each card?

A: The NVIDIA L40 is marked as end-of-life, while the NVIDIA L20 is listed as active. The L40 was released on 2022-10-12, and the L20 on 2023-11-15.

Q: How does each card compare to the RTX 6000 Ada Generation?

A: The L40’s average benchmark score of 281,655 is just 0.1% below the RTX 6000 Ada Generation’s 281,932. The L20’s average of 266,428 is 5.5% below the same rival, indicating the L40 is effectively on par with the RTX 6000 Ada Generation while the L20 trails it more noticeably.

Q: Do the cards differ in power consumption?

A: Yes. The L40 has a 300 W TDP with a suggested PSU of 700 W, while the L20 has a 275 W TDP with a suggested PSU of 600 W. Both use a single 16-pin power connector and are dual-slot designs.

The Verdict

The data points to a clear, unambiguous conclusion: the NVIDIA L40 is the superior performer in every measured metric of compute capability. Its 24.1% lead over the L20 in the Geekbench OpenCL test is mirrored by its higher average benchmark score (281,655 vs. 266,428) and its closer proximity to higher-tier rivals like the RTX 6000 Ada Generation (-0.1%) and L40S (-3.7%). The L20, while still a top-1% GPU, sits further from those same rivals—5.5% behind the RTX 6000 Ada Generation and 8.9% behind the L40S.

From a specification standpoint, the L40’s advantage is rooted in its larger core configuration: 18,176 shading units, 568 TMUs, 192 ROPs, 142 ray tracing cores, and 568 tensor cores. The L20’s higher base and boost clocks do not offset this deficit. The L40’s 90.52 TFLOPS FP32 throughput versus the L20’s 59.35 TFLOPS is a decisive gap that will translate directly to faster execution in compute-heavy workloads.

The L20, however, is not without merit. It is an active product, while the L40 is end-of-life, and it draws 25 W less power with a 100 W lower suggested PSU requirement. For scenarios where power draw is a constraint, these differences are relevant. The L20’s lower core count also means it may be more efficient per unit of compute in certain power-limited configurations, but the data does not provide efficiency metrics to confirm this. The verdict, based strictly on the numbers, is that the L40 wins on raw performance and benchmark scores; the L20 wins on availability and power envelope.

Specification Differences

The following fields differ between the two cards:

  • Base Clock: L40 at 735 MHz; L20 at 1440 MHz.
  • Boost Clock: L40 at 2490 MHz; L20 at 2520 MHz.
  • Shading Units: L40 at 18,176; L20 at 11,776.
  • TMUs: L40 at 568; L20 at 368.
  • ROPs: L40 at 192; L20 at 128.
  • Ray Tracing Cores: L40 at 142; L20 at 92.
  • Tensor Cores: L40 at 568; L20 at 368.
  • Pixel Rate: L40 at 478.1 GPixel/s; L20 at 322.6 GPixel/s.
  • Texture Rate: L40 at 1,414.3 GTexel/s; L20 at 927.4 GTexel/s.
  • FP32 / FP16: L40 at 90.52 TFLOPS; L20 at 59.35 TFLOPS.
  • TDP: L40 at 300 W; L20 at 275 W.
  • Suggested PSU: L40 at 700 W; L20 at 600 W.
  • Production Status: L40 end-of-life; L20 active.
  • Release Date: L40 on 2022-10-12; L20 on 2023-11-15.

All other fields—chip, architecture, process node, foundry, transistor count, die size, memory size/type/bus width/bandwidth, slot width, power connectors, bus interface, display outputs, APIs, dimensions, generation, predecessor, and successor—are identical.

Where Each One Wins

The NVIDIA L40 wins decisively in raw compute performance. Its 90.52 TFLOPS FP32 throughput, 24.1% head-to-head benchmark lead, and higher pixel and texture rates make it the clear choice for workloads that demand maximum processing power. The data shows the L40 performing at a level comparable to the RTX 6000 Ada Generation, making it suitable for the most demanding AI training, scientific simulation, or rendering tasks where every extra TFLOPS matters. Its higher core counts across every category—shading units, TMUs, ROPs, ray tracing cores, and tensor cores—mean it will handle parallel workloads with greater efficiency.

The NVIDIA L20 wins on operational flexibility. As an active product, it remains available for purchase, whereas the L40 is end-of-life. Its lower TDP of 275 W (versus 300 W) and lower suggested PSU of 600 W (versus 700 W) make it easier to integrate into existing server infrastructure with tighter power budgets. The higher base clock of 1440 MHz suggests it may respond more quickly to short bursts of activity, though the L40’s boost clock is nearly identical. For users prioritizing long-term availability and lower power draw over peak performance, the L20 is the viable option. For those who need the highest benchmark scores and compute throughput, the L40 is the only choice based on this data.

DETAILED SPECIFICATIONS

SPECIFICATION
L20
L40
Core Specs
Shading Units
11,776
18,176 +54.3%
Shaders
11,776
18,176 +54.3%
TMUs
368
568 +54.3%
ROPs
128
192 +50.0%
SM Count
92
142 +54.3%
Clocks
Base Clock
1440 MHz
735 MHz
Boost Clock
2520 MHz
2490 MHz
Memory Clock
2250 MHz 18 Gbps effective
2250 MHz 18 Gbps effective
Memory
Memory Size
48 GB
48 GB
VRAM (MB)
49,152
49,152 0.0%
Memory Type
GDDR6
GDDR6
Memory Bus
384 bit
384 bit
Bandwidth
864.0 GB/s
864.0 GB/s
Cache
L1 Cache
128 KB (per SM)
128 KB (per SM)
L2 Cache
96 MB
96 MB
Performance
Pixel Rate
322.6 GPixel/s
478.1 GPixel/s
Texture Rate
927.4 GTexel/s
1,414.3 GTexel/s
FP32 (TFLOPS)
59.35 TFLOPS
90.52 TFLOPS
FP64 (TFLOPS)
927.4 GFLOPS (1:64)
1,414.3 GFLOPS (1:64)
FP16 (TFLOPS)
59.35 TFLOPS (1:1)
90.52 TFLOPS (1:1)
AI/RT
RT Cores
92
142 +54.3%
Tensor Cores
368
568 +54.3%
Power
TDP
275 W
300 W
TDP (W)
275
300 +9.1%
Suggested PSU
600 W
700 W
Power Connectors
1x 16-pin
1x 16-pin
Architecture
Architecture
Ada Lovelace
Ada Lovelace
GPU Name
AD102
AD102
Generation
Server Ada (Lxx)
Server Ada (Lxx)
Process Size
5 nm
5 nm
Transistors
76,300 million
76,300 million
Die Size
609 mm²
609 mm²
Foundry
TSMC
TSMC
Density
125.3M / mm²
125.3M / mm²
API Support
DirectX
12 Ultimate (12_2)
12 Ultimate (12_2)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
3.0
3.0
CUDA
8.9
8.9
Shader Model
6.8
6.8
Physical
Slot Width
Dual-slot
Dual-slot
Length
267 mm 10.5 inches
267 mm 10.5 inches
Height
111 mm 4.4 inches
111 mm 4.4 inches
Outputs
4x DisplayPort 1.4a
4x DisplayPort 1.4a
Bus Interface
PCIe 4.0 x16
PCIe 4.0 x16
Other
Production
Active
End-of-life
Predecessor
Server Ampere
Server Ampere
Successor
Server Hopper
Server Hopper
View L20 Details View L40 Details