NVIDIA A100 PCIe 40 GB vs NVIDIA L40S Comparison
NVIDIA A100 PCIe 40 GB
L40S
PERFORMANCE BENCHMARKS
Analysis: NVIDIA A100 PCIe 40 GB vs NVIDIA L40S
The NVIDIA L40S and NVIDIA A100 PCIe 40 GB are two server-class accelerators from different architectural generations, and the benchmark data shows a decisive performance gap between them. Based on the available Geekbench results, the L40S leads in both compute workloads tested, with a clear margin that reflects its newer design and higher specification envelope. This analysis walks through the head-to-head numbers, the architectural reasons for the gap, and what the data means for potential use cases.
Head-to-Head Benchmarks
The two available benchmark tests, Geekbench OpenCL and Geekbench Vulkan, both go to the NVIDIA L40S. In the OpenCL test, the L40S scores 330,727 against the A100’s 178,627, a delta of 85.1%. That is not a small edge; it is nearly double the raw score. The Vulkan test shows a similar pattern, with the L40S at 260,799 and the A100 at 146,380, a 78.2% advantage for the newer card. In both cases, the L40S wins outright, and the data records 2 wins for the L40S and 0 for the A100.
Context from the nearest rivals makes these numbers more meaningful. The L40S has an average benchmark score of 295,763, which places it in the 99th percentile of all GPUs. Its closest competitors are the NVIDIA RTX 6000 Ada Generation (287,237, 3% lower), the NVIDIA L40 (284,111, 4.1% lower), and the AMD Instinct MI300X (317,994, 7% higher). The L40S is clearly a top-tier performer, sitting just below the MI300X and above its Ada siblings. The A100, by contrast, averages 162,504, putting it in the 97th percentile. Its nearest rivals are the AMD Radeon Pro W6800X (160,671, 1.1% lower), the AMD Radeon PRO W7800 (164,894, 1.4% higher), the NVIDIA RTX A5500 (165,217, 1.6% higher), and the NVIDIA RTX 4500 Ada Generation (166,094, 2.2% higher). The A100 is competitive within its own segment, but that segment is far below the L40S.
The deltas between the two cards are consistent across both tests, which suggests the performance gap is structural rather than workload-specific. The OpenCL margin of 85.1% is slightly larger than the Vulkan margin of 78.2%, but both are massive. The L40S does not just edge out the A100; it outperforms it by a wide margin in every measured scenario. For any workload that relies on these APIs, the L40S is the clear choice based on raw benchmark scores.
Architecture Differences
The performance gap traces directly to the underlying architectures. The L40S uses the AD102 chip on the Ada Lovelace architecture, built on a 5 nm process at TSMC. The A100 uses the GA100 chip on the Ampere architecture, built on a 7 nm process at the same foundry. The process node difference alone explains part of the efficiency and clock speed advantage, but the chip designs diverge further.
The L40S packs 76,300 million transistors into a 609 mm² die, giving a transistor density of 125.3M per mm². The A100 has 54,200 million transistors on a much larger 826 mm² die, resulting in a lower density of 65.6M per mm². The L40S is a smaller, denser chip, which allows for higher clock speeds. Its base clock is 1110 MHz and boost clock is 2520 MHz, while the A100 runs at 765 MHz base and 1410 MHz boost. That clock advantage is substantial and feeds directly into compute throughput.
The compute resources differ even more sharply. The L40S has 18,176 shading units, 568 TMUs, and 192 ROPs. The A100 has 6,912 shading units, 432 TMUs, and 160 ROPs. The L40S also has 142 RT cores and 568 tensor cores, while the A100 lists no RT cores and 432 tensor cores. The pixel rate for the L40S is 483.8 GPixel/s versus 225.6 GPixel/s for the A100. Texture rate is 1,431.4 GTexel/s versus 609.1 GTexel/s. FP32 performance is 91.61 TFLOPS for the L40S versus 19.49 TFLOPS for the A100. The L40S also posts 91.61 TFLOPS FP16 with a 1:1 ratio, while the A100 reaches 77.97 TFLOPS FP16 but at a 4:1 ratio, meaning its FP16 throughput is achieved through a different execution path.
Memory is another major divider. The L40S has 48 GB of GDDR6 on a 384-bit bus, delivering 864.0 GB/s bandwidth with an effective memory clock of 18 Gbps. The A100 has 40 GB of HBM2e on a 5120-bit bus, delivering 1.56 TB/s bandwidth with a much lower effective clock of 2.4 Gbps. The A100’s HBM2e gives it a bandwidth advantage, but the L40S compensates with a higher memory clock and larger capacity. The bus interface is PCIe 4.0 x16 on both, so interconnect bandwidth is identical.
The L40S also includes display outputs (1x HDMI 2.1 and 3x DisplayPort 1.4a), while the A100 has no outputs, indicating the A100 is a compute-only card. The L40S supports DirectX 12 Ultimate, OpenGL 4.6, and Vulkan 1.4, while the A100 lists no API support in the data. Power draw is 300 W for the L40S versus 250 W for the A100, with suggested PSUs of 700 W and 600 W respectively. Both are dual-slot cards with the same physical dimensions (267 mm length, 111 mm height), and both use different power connectors (1x 16-pin for the L40S, 8-pin EPS for the A100).
The Verdict
The data is unambiguous. The NVIDIA L40S wins both benchmark tests by margins of 85.1% and 78.2%, and its average benchmark score of 295,763 is roughly 82% higher than the A100’s 162,504. The L40S sits in the 99th percentile of all GPUs, while the A100 sits in the 97th, but that percentile difference understates the gap: the L40S is closer to the AMD Instinct MI300X (only 7% behind) than it is to the A100. The A100’s nearest rivals are all within a few percent of its score, meaning it is a solid mid-to-high performer, but it is not in the same class as the L40S.
For workloads that depend on Geekbench OpenCL or Vulkan performance, the L40S is the superior choice. Its higher shading unit count, tensor core count, clock speeds, and FP32 throughput all contribute to its dominance. The A100 does have a bandwidth advantage with its HBM2e memory (1.56 TB/s versus 864.0 GB/s), which could matter for memory-bound tasks, but the benchmark data does not reflect that advantage in the tests available. The A100 also has a lower TDP (250 W versus 300 W), which might be relevant for power-constrained environments, but the performance cost is severe.
The L40S is also the newer product, released in October 2022 versus June 2020 for the A100, and it is the successor to the Server Ampere generation that the A100 belongs to. Both are end-of-life, but the L40S represents the next step in NVIDIA’s server lineup. The A100’s only clear wins are lower power draw and higher memory bandwidth, but neither translates into a benchmark victory here. For anyone choosing between these two based on the data, the L40S is the performance pick.
FAQ
Q: How much faster is the NVIDIA L40S than the A100 in the OpenCL benchmark?
A: The L40S scores 330,727 versus the A100’s 178,627, which is an 85.1% advantage.
Q: Does the A100 win any benchmark in the head-to-head data?
A: No. The data records 2 wins for the L40S and 0 wins for the A100 across the available Geekbench OpenCL and Vulkan tests.
Q: What is the difference in FP32 performance between the two cards?
A: The L40S delivers 91.61 TFLOPS FP32, while the A100 delivers 19.49 TFLOPS, making the L40S roughly 4.7 times higher.
Q: Which card has more memory bandwidth?
A: The A100 has 1.56 TB/s bandwidth from its HBM2e memory, while the L40S has 864.0 GB/s from GDDR6.
Q: Are both cards the same physical size?
A: Yes, both are dual-slot cards with a length of 267 mm and a height of 111 mm.
Q: What is the transistor count for each chip?
A: The L40S has 76,300 million transistors on a 609 mm² die, while the A100 has 54,200 million transistors on an 826 mm² die.
Where Each One Wins
The L40S wins every benchmark test in the data, so its strengths are broad. It dominates in shading units (18,176 versus 6,912), texture rate (1,431.4 GTexel/s versus 609.1 GTexel/s), and pixel rate (483.8 GPixel/s versus 225.6 GPixel/s). It also has a massive FP32 advantage (91.61 TFLOPS versus 19.49 TFLOPS) and a tensor core count of 568 versus 432. The L40S is the clear choice for any workload that leverages these compute resources, particularly those that rely on the Ada Lovelace architecture’s features like RT cores (142 on the L40S, none listed on the A100).
The A100’s wins are narrower and not reflected in the benchmark scores. It has higher memory bandwidth (1.56 TB/s versus 864.0 GB/s), which could benefit memory-intensive applications such as large model inference or data processing that saturates bandwidth. It also draws less power (250 W versus 300 W) and has a lower suggested PSU (600 W versus 700 W), making it easier to fit into existing power budgets. The A100’s HBM2e memory on a 5120-bit bus is a different memory architecture than the L40S’s GDDR6 on 384-bit, and for certain workloads that bandwidth could be the deciding factor. However, the A100 has no display outputs and no API support listed, while the L40S has full display connectivity and modern API support, so the L40S is also more versatile for tasks that require visual output or current graphics APIs.
Specification Differences
The two cards differ across nearly every major specification. The L40S uses the AD102 chip on Ada Lovelace, while the A100 uses GA100 on Ampere. The process node is 5 nm for the L40S versus 7 nm for the A100. Transistor count is 76,300 million versus 54,200 million, and die size is 609 mm² versus 826 mm². Base clocks are 1110 MHz versus 765 MHz, and boost clocks are 2520 MHz versus 1410 MHz. Memory size is 48 GB versus 40 GB, with the L40S using GDDR6 and the A100 using HBM2e. Bus width is 384 bit versus 5120 bit, and bandwidth is 864.0 GB/s versus 1.56 TB/s.
Shading units are 18,176 versus 6,912, TMUs are 568 versus 432, and ROPs are 192 versus 160. The L40S has 142 RT cores while the A100 lists none. Tensor cores are 568 versus 432. Pixel rate is 483.8 GPixel/s versus 225.6 GPixel/s, and texture rate is 1,431.4 GTexel/s versus 609.1 GTexel/s. FP32 is 91.61 TFLOPS versus 19.49 TFLOPS, and FP16 is 91.61 TFLOPS (1:1) versus 77.97 TFLOPS (4:1). TDP is 300 W versus 250 W, with power connectors of 1x 16-pin versus 8-pin EPS. The L40S has display outputs (1x HDMI 2.1, 3x DisplayPort 1.4a) while the A100 has none. API support is present on the L40S (DirectX 12 Ultimate, OpenGL 4.6, Vulkan 1.4) and absent on the A100. Release dates differ (October 2022 versus June 2020), and the L40S’s predecessor is Server Ampere while the A100’s is Tesla Turing.