NVIDIA L4 vs NVIDIA L40S Comparison

NVIDIA
GEFORCE

NVIDIA L4

CORE STATE AD104
VRAM 24 GB
CLOCK SPEED 2040 MHz
TDP 72 W
BUS WIDTH 192 bit
ARCHITECTURE Ada Lovelace
nm
PROCESS 5 nm
LAUNCH DATE 2023
VS
NVIDIA
GEFORCE

L40S

CORE STATE AD102
VRAM 48 GB
CLOCK SPEED 2520 MHz
TDP 300 W
BUS WIDTH 384 bit
ARCHITECTURE Ada Lovelace
nm
PROCESS 5 nm
LAUNCH DATE 2022

PERFORMANCE BENCHMARKS

geekbench_opencl
140,838
330,727
geekbench_vulkan
121,306
260,799

Analysis: NVIDIA L4 vs NVIDIA L40S

Head-to-Head Benchmarks

The benchmark data delivers a decisive outcome: the NVIDIA L40S wins both recorded head-to-head tests outright, with no victories recorded for the NVIDIA L4. The margin is substantial in both OpenCL and Vulkan workloads, reinforcing the L40S as the dominant compute part in this pairing.

In the Geekbench OpenCL test, the L40S scores 330,727 against the L4’s 140,838. That is a delta of 134.8%, meaning the L40S more than doubles the L4’s raw throughput in this API. OpenCL is a common proxy for general-purpose compute, so this gap reflects a fundamental difference in execution resources, not a niche optimization. The L4 falls behind by a factor of 2.35, a margin that will be felt across rendering, simulation, and data-parallel workloads.

The Vulkan test tells a similar story, though the gap narrows slightly. The L40S posts 260,799, while the L4 manages 121,306. The delta here is 115%, meaning the L40S is 2.15 times faster. Vulkan’s lower-level access to the hardware can sometimes favor efficiency over raw size, but even so, the L4 cannot close the chasm. The L40S’s advantage in Vulkan is 139,493 points, which is larger than the L4’s entire Vulkan score. That is a striking metric: the L40S’s winning margin alone exceeds the L4’s total output.

Looking at the broader database context, the L40S sits at the 99th percentile among all GPUs, with an average benchmark score of 295,763. The L4, by contrast, is at the 95th percentile, with an average score of 131,072. The L40S’s average is 2.26 times the L4’s average, consistent with the head-to-head deltas. In the L40S’s nearest rival list, it trails the NVIDIA H200 NVL by 11.7% (334,891), but it leads the AMD Instinct MI300X by 7% (317,994), the NVIDIA L40 by 4.1% (284,111), and the NVIDIA RTX 6000 Ada Generation by 3% (287,237). The L4, meanwhile, sits in a much tighter cluster: it is essentially tied with the GeForce RTX 3090 Ti (131,938, a 0.7% deficit), the RTX 4000 Ada Generation (135,218, a 3.1% deficit), the A10M (135,230, a 3.1% deficit), and the Radeon PRO W6800 (135,396, a 3.2% deficit). The L4’s nearest rivals are all within a few percentage points, indicating that it competes in a crowded mid-range tier, whereas the L40S operates near the top of the stack.

The wins tally is 2-0 in favor of the L40S. No benchmark in the database shows the L4 ahead. This is a clean sweep, and the magnitude of each victory leaves little room for workload-specific reversals. The data suggests that any task which scales with shading units, texture throughput, or memory bandwidth will favor the L40S by a wide margin.

FAQ

Q: Which GPU has the higher average benchmark score?

A: The NVIDIA L40S, with an average score of 295,763, compared to the L4’s 131,072. The L40S is 2.26 times higher.

Q: How large is the L40S’s lead in the OpenCL test?

A: The L40S scores 330,727 versus 140,838 for the L4, a delta of 134.8%. The L40S is more than double the L4’s OpenCL performance.

Q: Does the L4 win any benchmark in the head-to-head data?

A: No. The L40S wins both recorded tests (OpenCL and Vulkan), with the L4 recording zero wins.

Q: Where does each GPU rank relative to all other GPUs in the database?

A: The L40S is in the 99th percentile, while the L4 is in the 95th percentile. The L40S is among the top 1% of all GPUs; the L4 is in the top 5%.

Q: How does the L4 compare to its nearest rivals?

A: The L4 is nearly tied with the GeForce RTX 3090 Ti (0.7% behind), the RTX 4000 Ada Generation (3.1% behind), the A10M (3.1% behind), and the Radeon PRO W6800 (3.2% behind). All of these are within a 3.2% band.

Q: How does the L40S compare to its nearest rivals?

A: The L40S leads the RTX 6000 Ada Generation by 3%, the L40 by 4.1%, and the AMD Instinct MI300X by 7%. It trails only the H200 NVL, which is 11.7% ahead.

Architecture Differences

Both GPUs are built on NVIDIA’s Ada Lovelace architecture and use TSMC’s 5 nm process node, but they are fundamentally different chips. The L40S uses the AD102 die, while the L4 uses the AD104 die. This is the root cause of the performance gap. AD102 is NVIDIA’s largest server-class chip in this generation, and AD104 is a much smaller, lower-power variant.

The transistor counts reflect this divergence. The L40S packs 76,300 million transistors on a 609 mm² die, yielding a transistor density of 125.3 million per mm². The L4 has 35,800 million transistors on a 294 mm² die, with a density of 121.8 million per mm². The L40S has more than double the transistors and more than double the die area. The density figures are nearly identical, so the performance difference is not from process efficiency but from raw scale.

The execution resources tell the same story. The L40S has 18,176 shading units, 568 texture mapping units (TMUs), and 192 raster output units (ROPs). The L4 has 7,424 shading units, 240 TMUs, and 80 ROPs. In every category, the L40S has roughly 2.4 to 2.5 times the hardware. Ray tracing cores follow the same pattern: 142 on the L40S versus 60 on the L4. Tensor cores, which are critical for AI inference and training, number 568 on the L40S versus 240 on the L4.

The memory subsystem also differs at the architectural level. The L40S uses a 384-bit memory bus, while the L4 uses a 192-bit bus. This directly affects bandwidth, which we will cover in the specification section, but the bus width difference is a structural feature of the two chips. The L40S is designed for memory-hungry workloads, and the L4 is not.

Both GPUs support the same API feature set: DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4. This means software compatibility is identical, and any differences in benchmark scores come down to raw hardware throughput, not API support. Both are also PCIe 4.0 x16 cards, so bus interface does not differentiate them.

The production status differs. The L40S is listed as end-of-life, while the L4 is active. The L40S was released on October 12, 2022, and the L4 on March 20, 2023. Both belong to the Server Ada (Lxx) generation and share the same predecessor (Server Ampere) and successor (Server Hopper) lineage.

Specification Differences

The clock speeds are one area where the smaller chip actually has a lower ceiling. The L40S has a base clock of 1110 MHz and a boost clock of 2520 MHz. The L4 has a base clock of 795 MHz and a boost clock of 2040 MHz. The L40S runs 480 MHz higher at boost, which compounds its architectural advantage.

Memory capacity and bandwidth are starkly different. The L40S has 48 GB of GDDR6 memory on a 384-bit bus, delivering 864.0 GB/s of bandwidth. The L4 has 24 GB of GDDR6 memory on a 192-bit bus, delivering 300.1 GB/s. The L40S has twice the capacity and 2.88 times the bandwidth. The memory clock also differs: the L40S runs at 2250 MHz (18 Gbps effective), while the L4 runs at 1563 MHz (12.5 Gbps effective).

Compute throughput figures amplify the gap. The L40S achieves 91.61 TFLOPS for FP32 and FP16 (1:1 ratio). The L4 achieves 30.29 TFLOPS for both. The L40S is 3.02 times faster in raw floating-point throughput. Pixel rate is 483.8 GPixel/s on the L40S versus 163.2 GPixel/s on the L4. Texture rate is 1,431.4 GTexel/s on the L40S versus 489.6 GTexel/s on the L4.

Power and physical design also diverge significantly. The L40S has a TDP of 300 W, requiring a single 16-pin power connector and a 700 W suggested PSU. It is a dual-slot card. The L4 has a TDP of 72 W, requires no power connector, and has a 250 W suggested PSU. It is a single-slot card. This makes the L4 far easier to integrate into dense servers, but the power budget explains why it cannot match the L40S’s output.

Dimensions reinforce the form-factor difference. The L40S is 267 mm long (10.5 inches) and 111 mm high (4.4 inches). The L4 is 169 mm long (6.7 inches) and 56 mm high (2.2 inches). The L40S is nearly twice as long and twice as tall.

Display outputs differ as well. The L40S has one HDMI 2.1 and three DisplayPort 1.4a outputs. The L4 has no display outputs, making it a pure compute accelerator. Neither card has a launch MSRP recorded in the database.

The Verdict

The data supports a clear verdict: the NVIDIA L40S is the superior performer in every measured benchmark. It wins the OpenCL test by 134.8% and the Vulkan test by 115%. Its average benchmark score of 295,763 places it in the 99th percentile, while the L4’s 131,072 places it in the 95th percentile. There is no recorded scenario where the L4 is faster.

The L40S’s advantages are not marginal. It has 2.45 times the shading units, 2.37 times the TMUs, 2.4 times the ROPs, 2.37 times the ray tracing cores, and 2.37 times the tensor cores. It has double the memory capacity and 2.88 times the bandwidth. Its FP32 throughput is 3.02 times higher. Every architectural and specification difference favors the L40S, and the benchmark results reflect that consistency.

However, the L4 is not without merit. Its 72 W TDP and single-slot form factor make it a low-power, space-efficient option. It requires no external power connector and only a 250 W PSU, which means it can be deployed in systems where the L40S would be physically or electrically impractical. The L4 is also an active product, while the L40S is end-of-life, so the L4 has a longer expected availability for new deployments.

The verdict for buyers is straightforward: if raw compute performance is the priority, the L40S is the only choice. If power and space constraints dominate, the L4 is the only choice. There is no middle ground where the L4’s performance is competitive with the L40S.

Where Each One Wins

The L40S wins every workload that benefits from high compute throughput. This includes large-scale rendering, complex simulation, deep learning training, and high-resolution inference. Its 48 GB memory capacity and 864.0 GB/s bandwidth make it suitable for datasets that exceed 24 GB, which the L4 cannot accommodate. The 91.61 TFLOPS FP32 throughput is nearly three times the L4’s, so any task that is compute-bound will see a proportional speedup. The 99th percentile ranking confirms that the L40S competes with the top accelerators in the database, trailing only the H200 NVL among its nearest rivals.

The L4 wins in deployment flexibility. Its 72 W TDP means it can be powered without additional connectors, and its 250 W PSU requirement is less than half of the L40S’s 700 W. The single-slot, 169 mm length design allows for dense packing in servers where the L40S’s dual-slot, 267 mm footprint would not fit. The L4’s 95th percentile ranking still places it above the vast majority of GPUs, and its nearest rivals (RTX 3090 Ti, RTX 4000 Ada, A10M, Radeon PRO W6800) are all within 3.2%, so it is a competitive mid-range option. For inference workloads that fit within 24 GB and do not require extreme throughput, the L4 can deliver adequate performance at a fraction of the power draw. It is also the only one of the two with an active production status, making it the more practical choice for new system designs that require long-term availability.

In short, the L40S is the performance king, and the L4 is the efficiency specialist. The benchmark data does not suggest any scenario where the L4 outperforms the L40S. It only suggests scenarios where the L4’s lower power and smaller size make it the only feasible option.

DETAILED SPECIFICATIONS

SPECIFICATION
L4
L40S
Core Specs
Shading Units
7,424
18,176 +144.8%
Shaders
7,424
18,176 +144.8%
TMUs
240
568 +136.7%
ROPs
80
192 +140.0%
SM Count
60
142 +136.7%
Clocks
Base Clock
795 MHz
1110 MHz
Boost Clock
2040 MHz
2520 MHz
Memory Clock
1563 MHz 12.5 Gbps effective
2250 MHz 18 Gbps effective
Memory
Memory Size
24 GB
48 GB
VRAM (MB)
24,576
49,152 +100.0%
Memory Type
GDDR6
GDDR6
Memory Bus
192 bit
384 bit
Bandwidth
300.1 GB/s
864.0 GB/s
Cache
L1 Cache
128 KB (per SM)
128 KB (per SM)
L2 Cache
48 MB
48 MB
Performance
Pixel Rate
163.2 GPixel/s
483.8 GPixel/s
Texture Rate
489.6 GTexel/s
1,431.4 GTexel/s
FP32 (TFLOPS)
30.29 TFLOPS
91.61 TFLOPS
FP64 (TFLOPS)
473.3 GFLOPS (1:64)
1,431.4 GFLOPS (1:64)
FP16 (TFLOPS)
30.29 TFLOPS (1:1)
91.61 TFLOPS (1:1)
AI/RT
RT Cores
60
142 +136.7%
Tensor Cores
240
568 +136.7%
Power
TDP
72 W
300 W
TDP (W)
72
300 +316.7%
Suggested PSU
250 W
700 W
Power Connectors
None
1x 16-pin
Architecture
Architecture
Ada Lovelace
Ada Lovelace
GPU Name
AD104
AD102
Generation
Server Ada (Lxx)
Server Ada (Lxx)
Process Size
5 nm
5 nm
Transistors
35,800 million
76,300 million
Die Size
294 mm²
609 mm²
Foundry
TSMC
TSMC
Density
121.8M / mm²
125.3M / mm²
API Support
DirectX
12 Ultimate (12_2)
12 Ultimate (12_2)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
3.0
3.0
CUDA
8.9
8.9
Shader Model
6.8
6.8
Physical
Slot Width
Single-slot
Dual-slot
Length
169 mm 6.7 inches
267 mm 10.5 inches
Height
56 mm 2.2 inches
111 mm 4.4 inches
Outputs
No outputs
1x HDMI 2.13x DisplayPort 1.4a
Bus Interface
PCIe 4.0 x16
PCIe 4.0 x16
Other
Production
Active
End-of-life
Predecessor
Server Ampere
Server Ampere
Successor
Server Hopper
Server Hopper
View L4 Details View L40S Details