NVIDIA L4 vs NVIDIA L40S Comparison
NVIDIA L4
L40S
PERFORMANCE BENCHMARKS
Analysis: NVIDIA L4 vs NVIDIA L40S
Head-to-Head Benchmarks
The benchmark data delivers a decisive outcome: the NVIDIA L40S wins both recorded head-to-head tests outright, with no victories recorded for the NVIDIA L4. The margin is substantial in both OpenCL and Vulkan workloads, reinforcing the L40S as the dominant compute part in this pairing.
In the Geekbench OpenCL test, the L40S scores 330,727 against the L4’s 140,838. That is a delta of 134.8%, meaning the L40S more than doubles the L4’s raw throughput in this API. OpenCL is a common proxy for general-purpose compute, so this gap reflects a fundamental difference in execution resources, not a niche optimization. The L4 falls behind by a factor of 2.35, a margin that will be felt across rendering, simulation, and data-parallel workloads.
The Vulkan test tells a similar story, though the gap narrows slightly. The L40S posts 260,799, while the L4 manages 121,306. The delta here is 115%, meaning the L40S is 2.15 times faster. Vulkan’s lower-level access to the hardware can sometimes favor efficiency over raw size, but even so, the L4 cannot close the chasm. The L40S’s advantage in Vulkan is 139,493 points, which is larger than the L4’s entire Vulkan score. That is a striking metric: the L40S’s winning margin alone exceeds the L4’s total output.
Looking at the broader database context, the L40S sits at the 99th percentile among all GPUs, with an average benchmark score of 295,763. The L4, by contrast, is at the 95th percentile, with an average score of 131,072. The L40S’s average is 2.26 times the L4’s average, consistent with the head-to-head deltas. In the L40S’s nearest rival list, it trails the NVIDIA H200 NVL by 11.7% (334,891), but it leads the AMD Instinct MI300X by 7% (317,994), the NVIDIA L40 by 4.1% (284,111), and the NVIDIA RTX 6000 Ada Generation by 3% (287,237). The L4, meanwhile, sits in a much tighter cluster: it is essentially tied with the GeForce RTX 3090 Ti (131,938, a 0.7% deficit), the RTX 4000 Ada Generation (135,218, a 3.1% deficit), the A10M (135,230, a 3.1% deficit), and the Radeon PRO W6800 (135,396, a 3.2% deficit). The L4’s nearest rivals are all within a few percentage points, indicating that it competes in a crowded mid-range tier, whereas the L40S operates near the top of the stack.
The wins tally is 2-0 in favor of the L40S. No benchmark in the database shows the L4 ahead. This is a clean sweep, and the magnitude of each victory leaves little room for workload-specific reversals. The data suggests that any task which scales with shading units, texture throughput, or memory bandwidth will favor the L40S by a wide margin.
FAQ
Q: Which GPU has the higher average benchmark score?
A: The NVIDIA L40S, with an average score of 295,763, compared to the L4’s 131,072. The L40S is 2.26 times higher.
Q: How large is the L40S’s lead in the OpenCL test?
A: The L40S scores 330,727 versus 140,838 for the L4, a delta of 134.8%. The L40S is more than double the L4’s OpenCL performance.
Q: Does the L4 win any benchmark in the head-to-head data?
A: No. The L40S wins both recorded tests (OpenCL and Vulkan), with the L4 recording zero wins.
Q: Where does each GPU rank relative to all other GPUs in the database?
A: The L40S is in the 99th percentile, while the L4 is in the 95th percentile. The L40S is among the top 1% of all GPUs; the L4 is in the top 5%.
Q: How does the L4 compare to its nearest rivals?
A: The L4 is nearly tied with the GeForce RTX 3090 Ti (0.7% behind), the RTX 4000 Ada Generation (3.1% behind), the A10M (3.1% behind), and the Radeon PRO W6800 (3.2% behind). All of these are within a 3.2% band.
Q: How does the L40S compare to its nearest rivals?
A: The L40S leads the RTX 6000 Ada Generation by 3%, the L40 by 4.1%, and the AMD Instinct MI300X by 7%. It trails only the H200 NVL, which is 11.7% ahead.
Architecture Differences
Both GPUs are built on NVIDIA’s Ada Lovelace architecture and use TSMC’s 5 nm process node, but they are fundamentally different chips. The L40S uses the AD102 die, while the L4 uses the AD104 die. This is the root cause of the performance gap. AD102 is NVIDIA’s largest server-class chip in this generation, and AD104 is a much smaller, lower-power variant.
The transistor counts reflect this divergence. The L40S packs 76,300 million transistors on a 609 mm² die, yielding a transistor density of 125.3 million per mm². The L4 has 35,800 million transistors on a 294 mm² die, with a density of 121.8 million per mm². The L40S has more than double the transistors and more than double the die area. The density figures are nearly identical, so the performance difference is not from process efficiency but from raw scale.
The execution resources tell the same story. The L40S has 18,176 shading units, 568 texture mapping units (TMUs), and 192 raster output units (ROPs). The L4 has 7,424 shading units, 240 TMUs, and 80 ROPs. In every category, the L40S has roughly 2.4 to 2.5 times the hardware. Ray tracing cores follow the same pattern: 142 on the L40S versus 60 on the L4. Tensor cores, which are critical for AI inference and training, number 568 on the L40S versus 240 on the L4.
The memory subsystem also differs at the architectural level. The L40S uses a 384-bit memory bus, while the L4 uses a 192-bit bus. This directly affects bandwidth, which we will cover in the specification section, but the bus width difference is a structural feature of the two chips. The L40S is designed for memory-hungry workloads, and the L4 is not.
Both GPUs support the same API feature set: DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4. This means software compatibility is identical, and any differences in benchmark scores come down to raw hardware throughput, not API support. Both are also PCIe 4.0 x16 cards, so bus interface does not differentiate them.
The production status differs. The L40S is listed as end-of-life, while the L4 is active. The L40S was released on October 12, 2022, and the L4 on March 20, 2023. Both belong to the Server Ada (Lxx) generation and share the same predecessor (Server Ampere) and successor (Server Hopper) lineage.
Specification Differences
The clock speeds are one area where the smaller chip actually has a lower ceiling. The L40S has a base clock of 1110 MHz and a boost clock of 2520 MHz. The L4 has a base clock of 795 MHz and a boost clock of 2040 MHz. The L40S runs 480 MHz higher at boost, which compounds its architectural advantage.
Memory capacity and bandwidth are starkly different. The L40S has 48 GB of GDDR6 memory on a 384-bit bus, delivering 864.0 GB/s of bandwidth. The L4 has 24 GB of GDDR6 memory on a 192-bit bus, delivering 300.1 GB/s. The L40S has twice the capacity and 2.88 times the bandwidth. The memory clock also differs: the L40S runs at 2250 MHz (18 Gbps effective), while the L4 runs at 1563 MHz (12.5 Gbps effective).
Compute throughput figures amplify the gap. The L40S achieves 91.61 TFLOPS for FP32 and FP16 (1:1 ratio). The L4 achieves 30.29 TFLOPS for both. The L40S is 3.02 times faster in raw floating-point throughput. Pixel rate is 483.8 GPixel/s on the L40S versus 163.2 GPixel/s on the L4. Texture rate is 1,431.4 GTexel/s on the L40S versus 489.6 GTexel/s on the L4.
Power and physical design also diverge significantly. The L40S has a TDP of 300 W, requiring a single 16-pin power connector and a 700 W suggested PSU. It is a dual-slot card. The L4 has a TDP of 72 W, requires no power connector, and has a 250 W suggested PSU. It is a single-slot card. This makes the L4 far easier to integrate into dense servers, but the power budget explains why it cannot match the L40S’s output.
Dimensions reinforce the form-factor difference. The L40S is 267 mm long (10.5 inches) and 111 mm high (4.4 inches). The L4 is 169 mm long (6.7 inches) and 56 mm high (2.2 inches). The L40S is nearly twice as long and twice as tall.
Display outputs differ as well. The L40S has one HDMI 2.1 and three DisplayPort 1.4a outputs. The L4 has no display outputs, making it a pure compute accelerator. Neither card has a launch MSRP recorded in the database.
The Verdict
The data supports a clear verdict: the NVIDIA L40S is the superior performer in every measured benchmark. It wins the OpenCL test by 134.8% and the Vulkan test by 115%. Its average benchmark score of 295,763 places it in the 99th percentile, while the L4’s 131,072 places it in the 95th percentile. There is no recorded scenario where the L4 is faster.
The L40S’s advantages are not marginal. It has 2.45 times the shading units, 2.37 times the TMUs, 2.4 times the ROPs, 2.37 times the ray tracing cores, and 2.37 times the tensor cores. It has double the memory capacity and 2.88 times the bandwidth. Its FP32 throughput is 3.02 times higher. Every architectural and specification difference favors the L40S, and the benchmark results reflect that consistency.
However, the L4 is not without merit. Its 72 W TDP and single-slot form factor make it a low-power, space-efficient option. It requires no external power connector and only a 250 W PSU, which means it can be deployed in systems where the L40S would be physically or electrically impractical. The L4 is also an active product, while the L40S is end-of-life, so the L4 has a longer expected availability for new deployments.
The verdict for buyers is straightforward: if raw compute performance is the priority, the L40S is the only choice. If power and space constraints dominate, the L4 is the only choice. There is no middle ground where the L4’s performance is competitive with the L40S.
Where Each One Wins
The L40S wins every workload that benefits from high compute throughput. This includes large-scale rendering, complex simulation, deep learning training, and high-resolution inference. Its 48 GB memory capacity and 864.0 GB/s bandwidth make it suitable for datasets that exceed 24 GB, which the L4 cannot accommodate. The 91.61 TFLOPS FP32 throughput is nearly three times the L4’s, so any task that is compute-bound will see a proportional speedup. The 99th percentile ranking confirms that the L40S competes with the top accelerators in the database, trailing only the H200 NVL among its nearest rivals.
The L4 wins in deployment flexibility. Its 72 W TDP means it can be powered without additional connectors, and its 250 W PSU requirement is less than half of the L40S’s 700 W. The single-slot, 169 mm length design allows for dense packing in servers where the L40S’s dual-slot, 267 mm footprint would not fit. The L4’s 95th percentile ranking still places it above the vast majority of GPUs, and its nearest rivals (RTX 3090 Ti, RTX 4000 Ada, A10M, Radeon PRO W6800) are all within 3.2%, so it is a competitive mid-range option. For inference workloads that fit within 24 GB and do not require extreme throughput, the L4 can deliver adequate performance at a fraction of the power draw. It is also the only one of the two with an active production status, making it the more practical choice for new system designs that require long-term availability.
In short, the L40S is the performance king, and the L4 is the efficiency specialist. The benchmark data does not suggest any scenario where the L4 outperforms the L40S. It only suggests scenarios where the L4’s lower power and smaller size make it the only feasible option.