NVIDIA L4 vs NVIDIA L40 Comparison
NVIDIA L4
L40
PERFORMANCE BENCHMARKS
Analysis: NVIDIA L4 vs NVIDIA L40
Head-to-Head Benchmarks
The database records two benchmark comparisons between the NVIDIA L40 and the NVIDIA L4, and in both tests the L40 takes a commanding lead. In Geekbench OpenCL, the L40 scores 330,926 against the L4’s 140,838, a difference of 135%. That is more than double the L4’s output in this compute-oriented workload. The Vulkan result narrows slightly in relative terms but remains decisive: the L40 posts 237,295 versus 121,306, a 95.6% advantage. Both numbers place the L40 in a different performance tier entirely, with the L4 trailing by roughly half in each scenario.
The L40’s OpenCL result is particularly striking when placed against its own nearest rivals. The database shows the L40 averaging 284,111 across all recorded benchmarks, which puts it 1.1% ahead of the NVIDIA RTX 6000 Ada Generation (287,237) and 3.9% behind the NVIDIA L40S (295,763). Against the NVIDIA L20 (251,147), the L40 leads by 13.1%. The L4, by contrast, averages 131,072, sitting 0.7% behind the GeForce RTX 3090 Ti (131,938) and 3.1% behind both the RTX 4000 Ada Generation (135,218) and the A10M (135,230). In practical terms, the L4 competes with high-end consumer and workstation cards from the previous generation, while the L40 sits among the fastest accelerators in the database’s records.
The gap between the two is not uniform across every workload type, however. The OpenCL delta of 135% is larger than the Vulkan delta of 95.6%, suggesting the L40’s advantage grows in compute-heavy scenarios that stress raw throughput. Vulkan, which often favors geometry and rasterization pipelines, still shows the L40 ahead by nearly double, but the margin is less extreme. This pattern indicates that the L40’s strength is not merely clock speed but architectural resources that scale with parallel compute demands.
The Verdict
The data points to a clear split in intended use cases. The L40 wins both recorded benchmarks outright, with win counts of 2 for the L40 and 0 for the L4. Any workload that depends on OpenCL or Vulkan performance will see a substantial benefit from choosing the L40. The L40 also holds the 99th percentile position among all GPUs in the database, while the L4 sits at the 95th percentile. That four-percentile gap matters: the L40 is near the top of the entire field, whereas the L4 is merely above average.
For users who need maximum compute throughput in server environments, the L40 is the clear choice from these measurements. Its OpenCL score alone is 2.35 times the L4’s, and its Vulkan score is 1.96 times higher. The L4, however, is not without merit. It delivers roughly 42.6% of the L40’s OpenCL performance and 51.1% of its Vulkan performance, which may be entirely adequate for lighter inference or rendering tasks. The L4 also draws far less power, is a single-slot card, and requires no external power connectors, making it far easier to deploy in dense server chassis where space and thermal envelopes are tight.
The verdict from the benchmark data is straightforward: the L40 dominates in raw performance, while the L4 wins on deployment flexibility and efficiency. Neither card is a substitute for the other; they serve different tiers of the same server market.
Architecture Differences
Both GPUs use the Ada Lovelace architecture and are built on TSMC’s 5 nm process, but they are fundamentally different chips. The L40 uses the AD102 die, the largest Ada Lovelace chip, while the L4 uses the AD104, a smaller and significantly cut-down variant. The AD102 contains 76,300 million transistors on a 609 mm² die, yielding a transistor density of 125.3 million per square millimeter. The AD104, by comparison, has 35,800 million transistors on a 294 mm² die, with a density of 121.8 million per square millimeter. The L40’s die is more than twice the size and packs more than twice the transistors, which explains its massive compute advantage.
The L40’s execution resources are proportionally larger across the board. It has 18,176 shading units, 568 texture mapping units, and 192 ROPs, versus the L4’s 7,424 shading units, 240 TMUs, and 80 ROPs. Ray tracing cores follow the same pattern: the L40 has 142, the L4 has 60. Tensor cores, which are critical for AI and deep learning workloads, number 568 on the L40 and 240 on the L4. These are not minor differences; they represent a wholesale reduction in compute capacity for the L4.
The memory subsystem also diverges sharply. The L40 uses a 384-bit memory bus, while the L4 uses a 192-bit bus. This is paired with different memory configurations: the L40 offers 48 GB of GDDR6, the L4 offers 24 GB. The L40’s memory bandwidth is 864.0 GB/s, nearly three times the L4’s 300.1 GB/s. Effective memory speed differs as well, with the L40 running at 18 Gbps and the L4 at 12.5 Gbps. The L40’s larger bus and faster memory make it far better suited to data sets that exceed the L4’s capacity or that require sustained high-bandwidth access.
Both chips support the same API feature set: DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4. The architectural differences are therefore not about feature support but about scale. The L40 is the full Ada Lovelace server implementation, while the L4 is a deliberately reduced variant for lower-power, lower-cost deployments.
Specification Differences
The two cards differ in nearly every measurable specification. Clock speeds are the exception where the L4 actually runs higher at base: the L4 has a base clock of 795 MHz versus the L40’s 735 MHz. Boost clocks reverse that, with the L40 reaching 2,490 MHz and the L4 peaking at 2,040 MHz. Memory clocks also favor the L40, which runs at 2,250 MHz (18 Gbps effective) versus the L4’s 1,563 MHz (12.5 Gbps effective).
The L40’s compute rates dwarf the L4’s. Pixel rate on the L40 is 478.1 GPixel/s against 163.2 GPixel/s on the L4. Texture rate is 1,414.3 GTexel/s versus 489.6 GTexel/s. FP32 throughput is 90.52 TFLOPS on the L40 and 30.29 TFLOPS on the L4, a 3x difference. FP16 performance is identical to FP32 on both cards, listed as a 1:1 ratio, so the same 3x gap applies to half-precision workloads.
Power and physical specifications show the L4’s design goals. The L40 has a 300 W TDP and requires a 700 W suggested power supply, along with a single 16-pin power connector. It is a dual-slot card measuring 267 mm in length and 111 mm in height. The L4, by contrast, has a 72 W TDP and a 250 W suggested power supply. It needs no external power connectors at all, is a single-slot card, and is much smaller at 169 mm long and 56 mm high. Both use a PCIe 4.0 x16 interface.
Display outputs differ as well: the L40 has four DisplayPort 1.4a outputs, while the L4 has no display outputs at all. This confirms the L4 is intended purely as a compute or inference accelerator, not a workstation graphics card. The L40, despite being a server part, retains display capabilities.
Production status also separates them. The L40 is listed as end-of-life, while the L4 is active. Release dates place the L40 in October 2022 and the L4 in March 2023. Both list their predecessor as Server Ampere and successor as Server Hopper, indicating they occupy the same generational slot in NVIDIA’s server lineup.
FAQ
Q: Which GPU is faster in OpenCL?
A: The NVIDIA L40 scores 330,926 in Geekbench OpenCL, which is 135% higher than the NVIDIA L4’s score of 140,838.
Q: How much of the L40’s Vulkan performance does the L4 achieve?
A: The L4 scores 121,306 in Geekbench Vulkan, which is roughly 51.1% of the L40’s 237,295, a 95.6% deficit.
Q: What is the memory capacity difference?
A: The L40 has 48 GB of GDDR6 memory, while the L4 has 24 GB. The L40 also has a 384-bit memory bus versus the L4’s 192-bit bus.
Q: Does the L4 support the same APIs as the L40?
A: Yes, both cards support DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4.
Q: Which card requires an external power connector?
A: The L40 requires a single 16-pin power connector and a 700 W suggested power supply. The L4 has no power connectors and only needs a 250 W suggested power supply.
Q: How do the two cards compare in transistor count?
A: The L40’s AD102 chip has 76,300 million transistors, while the L4’s AD104 chip has 35,800 million transistors.
Where Each One Wins
The L40 wins in every performance category recorded in the database. Its OpenCL and Vulkan scores are both substantially higher, and its compute resources are three times larger in shading units, TMUs, and FP32 throughput. For workloads that are compute-bound, such as large-scale AI training, scientific simulation, or high-resolution rendering, the L40 is the only reasonable choice among these two. Its 48 GB memory and 864.0 GB/s bandwidth also allow it to handle data sets that would exceed the L4’s 24 GB capacity and 300.1 GB/s bandwidth. The L40’s 99th percentile standing among all GPUs reinforces its position as a top-tier accelerator.
The L4 wins on deployment efficiency and physical footprint. It consumes 72 W versus the L40’s 300 W, meaning a server can house multiple L4s within the power budget of a single L40. It is a single-slot card with no power connectors, which simplifies installation in dense chassis. The L4 is also active in production, while the L40 is end-of-life, so new system builds may have an easier time sourcing the L4. Its 95th percentile rank is still strong, and for lighter inference tasks or edge deployments where space and power are constrained, the L4’s 30.29 TFLOPS of FP32 performance is far from trivial.
The use-case split is clear: choose the L40 for maximum throughput and memory capacity, choose the L4 for efficiency and density. The benchmark data does not show any workload where the L4 surpasses the L40, but it does show that the L4 occupies a viable lower-power tier that the L40 cannot match in physical terms.