NVIDIA A100 PCIe 80 GB vs NVIDIA L40S Comparison
NVIDIA A100 PCIe 80 GB
L40S
PERFORMANCE BENCHMARKS
Analysis: NVIDIA A100 PCIe 80 GB vs NVIDIA L40S
# NVIDIA L40S vs NVIDIA A100 PCIe 80 GB
The NVIDIA L40S and NVIDIA A100 PCIe 80 GB represent two distinct generations of server accelerators, with the L40S built on the newer Ada Lovelace architecture and the A100 on the older Ampere design. The benchmark data shows a decisive performance gap: the L40S achieves an average benchmark score of 295,763 versus 207,124 for the A100, placing both in the 99th percentile of all GPUs but separating them by a substantial margin in raw compute output.
Head-to-Head Benchmarks
The only directly comparable benchmark in the dataset is Geekbench OpenCL, where the L40S scores 330,727 against the A100's 207,124. That represents a 59.7% advantage for the L40S—a massive lead that underscores the generational leap between the two architectures. To put this in context, the L40S sits 3% ahead of the NVIDIA RTX 6000 Ada Generation (287,237) and 4.1% ahead of the NVIDIA L40 (284,111) in average score. The A100, by comparison, trails the AMD Instinct MI300X by 5.8% (which scores 219,827) and the NVIDIA PG506-232 by 8% (225,124), while leading the RTX 6000D by 5.7% (195,964) and the Tesla V100S PCIe 32 GB by 6.5% (194,415).
The 59.7% delta in OpenCL is not a marginal improvement; it is a wholesale shift in compute capability. The L40S also shows strength in Vulkan, scoring 260,799, a test the A100 does not appear in the dataset for. When examining the nearest rivals for each card, the L40S is positioned among much faster accelerators—it trails the AMD Instinct MI300X by 7% (317,994) and the NVIDIA H200 NVL by 11.7% (334,891)—whereas the A100's competitors are all in a tighter, lower band. This suggests the L40S competes in a higher performance tier altogether.
Where Each One Wins
The L40S wins the only head-to-head benchmark available, and it does so by a wide margin. Its 91.61 TFLOPS of FP32 performance dwarfs the A100's 19.49 TFLOPS—a 4.7× difference in single-precision throughput. For workloads that rely heavily on FP32, such as traditional rendering, simulation, or certain AI inference paths, the L40S is clearly the superior choice. The L40S also offers 91.61 TFLOPS of FP16 performance at a 1:1 ratio, meaning it does not sacrifice half-precision throughput relative to FP32. The A100, in contrast, delivers 77.97 TFLOPS of FP16 but at a 4:1 ratio, indicating its FP16 output requires specialized tensor operations to reach that figure.
The A100's strengths lie elsewhere. It offers 80 GB of HBM2e memory versus the L40S's 48 GB of GDDR6, and its memory bandwidth of 1.94 TB/s is more than double the L40S's 864.0 GB/s. For workloads that are memory-capacity-bound—such as very large model inference or datasets that exceed 48 GB—the A100 has a clear advantage. The A100 also uses a 5120-bit memory bus, which is over 13× wider than the L40S's 384-bit bus, reflecting its design focus on memory-intensive compute rather than graphics-oriented tasks.
In terms of raw pixel and texture throughput, the L40S wins decisively: 483.8 GPixel/s versus 225.6 GPixel/s, and 1,431.4 GTexel/s versus 609.1 GTexel/s. The L40S also has far more shading units (18,176 vs 6,912), TMUs (568 vs 432), and ROPs (192 vs 160). The L40S includes 142 RT cores while the A100 has none listed, making the L40S the only option for ray-traced workloads. The A100 does match the L40S on TDP (300 W each) and suggested PSU (700 W each), but the performance-per-watt picture heavily favors the L40S given its higher output at the same power draw.
Architecture Differences
The L40S is built on TSMC's 5 nm process node with 76,300 million transistors packed into a 609 mm² die, yielding a transistor density of 125.3M per mm². The A100 uses TSMC's 7 nm node with 54,200 million transistors on a larger 826 mm² die, resulting in a density of just 65.6M per mm². The L40S's smaller, denser process gives it a fundamental advantage in both performance and efficiency—it achieves 91.61 TFLOPS FP32 at 300 W, while the A100 manages only 19.49 TFLOPS FP32 at the same power envelope.
The memory subsystems could not be more different. The L40S uses 48 GB of GDDR6 at 2250 MHz (18 Gbps effective) across a 384-bit bus, producing 864.0 GB/s of bandwidth. The A100 uses 80 GB of HBM2e at 1512 MHz (3 Gbps effective) across a 5120-bit bus, producing 1.94 TB/s. This is a fundamental design split: the L40S prioritizes capacity per dollar and clock speed, while the A100 prioritizes raw bandwidth and total capacity. The L40S's clock speeds are also much higher: 1110 MHz base and 2520 MHz boost versus 1065 MHz base and 1410 MHz boost for the A100.
The L40S is built on the Ada Lovelace architecture (chip AD102) and belongs to the Server Ada (Lxx) generation, while the A100 uses the Ampere architecture (chip GA100) from the Server Ampere (Axx) generation. The L40S supports DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4, while the A100 lists no API support at all—it is a compute-only accelerator with no display outputs. The L40S has 1x HDMI 2.1 and 3x DisplayPort 1.4a outputs, whereas the A100 has no outputs whatsoever. The L40S uses a 1x 16-pin power connector, while the A100 uses an 8-pin EPS connector. Both are dual-slot cards measuring 267 mm in length and 111 mm in height.
The L40S has 568 tensor cores, the A100 has 432. The L40S has 568 TMUs, the A100 has 432. The L40S has 142 RT cores, the A100 has none. These are not minor differences—they reflect entirely different design philosophies. The L40S is a versatile accelerator that can handle graphics, ray tracing, and compute, while the A100 is a pure compute engine optimized for massive parallel workloads with no graphics capability.
The Verdict
From the data, the NVIDIA L40S is the clear performance winner for most workloads. It leads the A100 by 59.7% in the only shared benchmark, offers 4.7× higher FP32 throughput, and provides 91.61 TFLOPS of FP16 at a 1:1 ratio. It also includes ray tracing cores, display outputs, and modern API support, making it far more versatile. The L40S's nearest rivals are all significantly faster than the A100's nearest rivals, confirming that it belongs in a higher performance class.
The A100 PCIe 80 GB, however, retains one critical advantage: memory. With 80 GB of HBM2e and 1.94 TB/s of bandwidth, it offers 66.7% more memory capacity and 2.25× the bandwidth of the L40S. For workloads that require loading very large models or datasets into GPU memory—and where FP32 throughput is less critical than capacity—the A100 remains the practical choice. Its 77.97 TFLOPS of FP16 (at 4:1) is also respectable, though the L40S matches that with 91.61 TFLOPS at 1:1.
For most users, the L40S is the better accelerator: it is faster, more feature-rich, and more efficient in FP32 and FP16 compute. For users whose primary constraint is memory capacity or bandwidth, the A100's 80 GB HBM2e configuration makes it the only option that fits. The choice comes down to whether the workload is compute-bound (choose L40S) or memory-bound (choose A100).
FAQ
Q: Which GPU has a higher average benchmark score?
A: The NVIDIA L40S has an average benchmark score of 295,763, while the NVIDIA A100 PCIe 80 GB has an average score of 207,124—a difference of 88,639 points.
Q: How much faster is the L40S in Geekbench OpenCL?
A: The L40S scores 330,727 in Geekbench OpenCL versus 207,124 for the A100, giving the L40S a 59.7% advantage.
Q: Does the A100 have ray tracing cores?
A: No, the A100 lists no RT cores, while the L40S has 142 RT cores.
Q: Which GPU has more memory bandwidth?
A: The A100 PCIe 80 GB has 1.94 TB/s of bandwidth, more than double the L40S's 864.0 GB/s.
Q: What is the difference in FP32 performance?
A: The L40S delivers 91.61 TFLOPS of FP32, while the A100 delivers 19.49 TFLOPS—the L40S is roughly 4.7× faster in single-precision compute.
Q: Do both cards support display outputs?
A: No, the L40S has 1x HDMI 2.1 and 3x DisplayPort 1.4a outputs, while the A100 has no display outputs.
Specification Differences
| Specification | NVIDIA L40S | NVIDIA A100 PCIe 80 GB |
|---|---|---|
| Architecture | Ada Lovelace | Ampere |
| Chip | AD102 | GA100 |
| Process Node | 5 nm | 7 nm |
| Transistors | 76,300 million | 54,200 million |
| Die Size | 609 mm² | 826 mm² |
| Transistor Density | 125.3M / mm² | 65.6M / mm² |
| Base Clock | 1110 MHz | 1065 MHz |
| Boost Clock | 2520 MHz | 1410 MHz |
| Memory Size | 48 GB | 80 GB |
| Memory Type | GDDR6 | HBM2e |
| Memory Bus Width | 384 bit | 5120 bit |
| Memory Bandwidth | 864.0 GB/s | 1.94 TB/s |
| Shading Units | 18,176 | 6,912 |
| TMUs | 568 | 432 |
| ROPs | 192 | 160 |
| RT Cores | 142 | None |
| Tensor Cores | 568 | 432 |
| Pixel Rate | 483.8 GPixel/s | 225.6 GPixel/s |
| Texture Rate | 1,431.4 GTexel/s | 609.1 GTexel/s |
| FP32 Performance | 91.61 TFLOPS | 19.49 TFLOPS |
| FP16 Performance | 91.61 TFLOPS (1:1) | 77.97 TFLOPS (4:1) |
| Power Connectors | 1x 16-pin | 8-pin EPS |
| Display Outputs | 1x HDMI 2.1, 3x DisplayPort 1.4a | No outputs |
| DirectX Support | 12 Ultimate (12_2) | None |
| OpenGL Support | 4.6 | None |
| Vulkan Support | 1.4 | None |
| Release Date | 2022-10-12 | 2021-06-27 |