NVIDIA A100 PCIe 80 GB vs NVIDIA L40S Comparison

NVIDIA
GEFORCE

NVIDIA A100 PCIe 80 GB

CORE STATE GA100
VRAM 80 GB
CLOCK SPEED 1410 MHz
TDP 300 W
BUS WIDTH 5120 bit
ARCHITECTURE Ampere
nm
PROCESS 7 nm
LAUNCH DATE 2021
VS
NVIDIA
GEFORCE

L40S

CORE STATE AD102
VRAM 48 GB
CLOCK SPEED 2520 MHz
TDP 300 W
BUS WIDTH 384 bit
ARCHITECTURE Ada Lovelace
nm
PROCESS 5 nm
LAUNCH DATE 2022

PERFORMANCE BENCHMARKS

geekbench_opencl
207,124
330,727
geekbench_vulkan
N/A
260,799

Analysis: NVIDIA A100 PCIe 80 GB vs NVIDIA L40S

# NVIDIA L40S vs NVIDIA A100 PCIe 80 GB

The NVIDIA L40S and NVIDIA A100 PCIe 80 GB represent two distinct generations of server accelerators, with the L40S built on the newer Ada Lovelace architecture and the A100 on the older Ampere design. The benchmark data shows a decisive performance gap: the L40S achieves an average benchmark score of 295,763 versus 207,124 for the A100, placing both in the 99th percentile of all GPUs but separating them by a substantial margin in raw compute output.

Head-to-Head Benchmarks

The only directly comparable benchmark in the dataset is Geekbench OpenCL, where the L40S scores 330,727 against the A100's 207,124. That represents a 59.7% advantage for the L40S—a massive lead that underscores the generational leap between the two architectures. To put this in context, the L40S sits 3% ahead of the NVIDIA RTX 6000 Ada Generation (287,237) and 4.1% ahead of the NVIDIA L40 (284,111) in average score. The A100, by comparison, trails the AMD Instinct MI300X by 5.8% (which scores 219,827) and the NVIDIA PG506-232 by 8% (225,124), while leading the RTX 6000D by 5.7% (195,964) and the Tesla V100S PCIe 32 GB by 6.5% (194,415).

The 59.7% delta in OpenCL is not a marginal improvement; it is a wholesale shift in compute capability. The L40S also shows strength in Vulkan, scoring 260,799, a test the A100 does not appear in the dataset for. When examining the nearest rivals for each card, the L40S is positioned among much faster accelerators—it trails the AMD Instinct MI300X by 7% (317,994) and the NVIDIA H200 NVL by 11.7% (334,891)—whereas the A100's competitors are all in a tighter, lower band. This suggests the L40S competes in a higher performance tier altogether.

Where Each One Wins

The L40S wins the only head-to-head benchmark available, and it does so by a wide margin. Its 91.61 TFLOPS of FP32 performance dwarfs the A100's 19.49 TFLOPS—a 4.7× difference in single-precision throughput. For workloads that rely heavily on FP32, such as traditional rendering, simulation, or certain AI inference paths, the L40S is clearly the superior choice. The L40S also offers 91.61 TFLOPS of FP16 performance at a 1:1 ratio, meaning it does not sacrifice half-precision throughput relative to FP32. The A100, in contrast, delivers 77.97 TFLOPS of FP16 but at a 4:1 ratio, indicating its FP16 output requires specialized tensor operations to reach that figure.

The A100's strengths lie elsewhere. It offers 80 GB of HBM2e memory versus the L40S's 48 GB of GDDR6, and its memory bandwidth of 1.94 TB/s is more than double the L40S's 864.0 GB/s. For workloads that are memory-capacity-bound—such as very large model inference or datasets that exceed 48 GB—the A100 has a clear advantage. The A100 also uses a 5120-bit memory bus, which is over 13× wider than the L40S's 384-bit bus, reflecting its design focus on memory-intensive compute rather than graphics-oriented tasks.

In terms of raw pixel and texture throughput, the L40S wins decisively: 483.8 GPixel/s versus 225.6 GPixel/s, and 1,431.4 GTexel/s versus 609.1 GTexel/s. The L40S also has far more shading units (18,176 vs 6,912), TMUs (568 vs 432), and ROPs (192 vs 160). The L40S includes 142 RT cores while the A100 has none listed, making the L40S the only option for ray-traced workloads. The A100 does match the L40S on TDP (300 W each) and suggested PSU (700 W each), but the performance-per-watt picture heavily favors the L40S given its higher output at the same power draw.

Architecture Differences

The L40S is built on TSMC's 5 nm process node with 76,300 million transistors packed into a 609 mm² die, yielding a transistor density of 125.3M per mm². The A100 uses TSMC's 7 nm node with 54,200 million transistors on a larger 826 mm² die, resulting in a density of just 65.6M per mm². The L40S's smaller, denser process gives it a fundamental advantage in both performance and efficiency—it achieves 91.61 TFLOPS FP32 at 300 W, while the A100 manages only 19.49 TFLOPS FP32 at the same power envelope.

The memory subsystems could not be more different. The L40S uses 48 GB of GDDR6 at 2250 MHz (18 Gbps effective) across a 384-bit bus, producing 864.0 GB/s of bandwidth. The A100 uses 80 GB of HBM2e at 1512 MHz (3 Gbps effective) across a 5120-bit bus, producing 1.94 TB/s. This is a fundamental design split: the L40S prioritizes capacity per dollar and clock speed, while the A100 prioritizes raw bandwidth and total capacity. The L40S's clock speeds are also much higher: 1110 MHz base and 2520 MHz boost versus 1065 MHz base and 1410 MHz boost for the A100.

The L40S is built on the Ada Lovelace architecture (chip AD102) and belongs to the Server Ada (Lxx) generation, while the A100 uses the Ampere architecture (chip GA100) from the Server Ampere (Axx) generation. The L40S supports DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4, while the A100 lists no API support at all—it is a compute-only accelerator with no display outputs. The L40S has 1x HDMI 2.1 and 3x DisplayPort 1.4a outputs, whereas the A100 has no outputs whatsoever. The L40S uses a 1x 16-pin power connector, while the A100 uses an 8-pin EPS connector. Both are dual-slot cards measuring 267 mm in length and 111 mm in height.

The L40S has 568 tensor cores, the A100 has 432. The L40S has 568 TMUs, the A100 has 432. The L40S has 142 RT cores, the A100 has none. These are not minor differences—they reflect entirely different design philosophies. The L40S is a versatile accelerator that can handle graphics, ray tracing, and compute, while the A100 is a pure compute engine optimized for massive parallel workloads with no graphics capability.

The Verdict

From the data, the NVIDIA L40S is the clear performance winner for most workloads. It leads the A100 by 59.7% in the only shared benchmark, offers 4.7× higher FP32 throughput, and provides 91.61 TFLOPS of FP16 at a 1:1 ratio. It also includes ray tracing cores, display outputs, and modern API support, making it far more versatile. The L40S's nearest rivals are all significantly faster than the A100's nearest rivals, confirming that it belongs in a higher performance class.

The A100 PCIe 80 GB, however, retains one critical advantage: memory. With 80 GB of HBM2e and 1.94 TB/s of bandwidth, it offers 66.7% more memory capacity and 2.25× the bandwidth of the L40S. For workloads that require loading very large models or datasets into GPU memory—and where FP32 throughput is less critical than capacity—the A100 remains the practical choice. Its 77.97 TFLOPS of FP16 (at 4:1) is also respectable, though the L40S matches that with 91.61 TFLOPS at 1:1.

For most users, the L40S is the better accelerator: it is faster, more feature-rich, and more efficient in FP32 and FP16 compute. For users whose primary constraint is memory capacity or bandwidth, the A100's 80 GB HBM2e configuration makes it the only option that fits. The choice comes down to whether the workload is compute-bound (choose L40S) or memory-bound (choose A100).

FAQ

Q: Which GPU has a higher average benchmark score?

A: The NVIDIA L40S has an average benchmark score of 295,763, while the NVIDIA A100 PCIe 80 GB has an average score of 207,124—a difference of 88,639 points.

Q: How much faster is the L40S in Geekbench OpenCL?

A: The L40S scores 330,727 in Geekbench OpenCL versus 207,124 for the A100, giving the L40S a 59.7% advantage.

Q: Does the A100 have ray tracing cores?

A: No, the A100 lists no RT cores, while the L40S has 142 RT cores.

Q: Which GPU has more memory bandwidth?

A: The A100 PCIe 80 GB has 1.94 TB/s of bandwidth, more than double the L40S's 864.0 GB/s.

Q: What is the difference in FP32 performance?

A: The L40S delivers 91.61 TFLOPS of FP32, while the A100 delivers 19.49 TFLOPS—the L40S is roughly 4.7× faster in single-precision compute.

Q: Do both cards support display outputs?

A: No, the L40S has 1x HDMI 2.1 and 3x DisplayPort 1.4a outputs, while the A100 has no display outputs.

Specification Differences

| Specification | NVIDIA L40S | NVIDIA A100 PCIe 80 GB |

|---|---|---|

| Architecture | Ada Lovelace | Ampere |

| Chip | AD102 | GA100 |

| Process Node | 5 nm | 7 nm |

| Transistors | 76,300 million | 54,200 million |

| Die Size | 609 mm² | 826 mm² |

| Transistor Density | 125.3M / mm² | 65.6M / mm² |

| Base Clock | 1110 MHz | 1065 MHz |

| Boost Clock | 2520 MHz | 1410 MHz |

| Memory Size | 48 GB | 80 GB |

| Memory Type | GDDR6 | HBM2e |

| Memory Bus Width | 384 bit | 5120 bit |

| Memory Bandwidth | 864.0 GB/s | 1.94 TB/s |

| Shading Units | 18,176 | 6,912 |

| TMUs | 568 | 432 |

| ROPs | 192 | 160 |

| RT Cores | 142 | None |

| Tensor Cores | 568 | 432 |

| Pixel Rate | 483.8 GPixel/s | 225.6 GPixel/s |

| Texture Rate | 1,431.4 GTexel/s | 609.1 GTexel/s |

| FP32 Performance | 91.61 TFLOPS | 19.49 TFLOPS |

| FP16 Performance | 91.61 TFLOPS (1:1) | 77.97 TFLOPS (4:1) |

| Power Connectors | 1x 16-pin | 8-pin EPS |

| Display Outputs | 1x HDMI 2.1, 3x DisplayPort 1.4a | No outputs |

| DirectX Support | 12 Ultimate (12_2) | None |

| OpenGL Support | 4.6 | None |

| Vulkan Support | 1.4 | None |

| Release Date | 2022-10-12 | 2021-06-27 |

DETAILED SPECIFICATIONS

SPECIFICATION
A100 PCIe 80 GB
L40S
Core Specs
Shading Units
6,912
18,176 +163.0%
Shaders
6,912
18,176 +163.0%
TMUs
432
568 +31.5%
ROPs
160
192 +20.0%
SM Count
108
142 +31.5%
Clocks
Base Clock
1065 MHz
1110 MHz
Boost Clock
1410 MHz
2520 MHz
Memory Clock
1512 MHz 3 Gbps effective
2250 MHz 18 Gbps effective
Memory
Memory Size
80 GB
48 GB
VRAM (MB)
81,920
49,152 -40.0%
Memory Type
HBM2e
GDDR6
Memory Bus
5120 bit
384 bit
Bandwidth
1.94 TB/s
864.0 GB/s
Cache
L1 Cache
192 KB (per SM)
128 KB (per SM)
L2 Cache
80 MB
48 MB
Performance
Pixel Rate
225.6 GPixel/s
483.8 GPixel/s
Texture Rate
609.1 GTexel/s
1,431.4 GTexel/s
FP32 (TFLOPS)
19.49 TFLOPS
91.61 TFLOPS
FP64 (TFLOPS)
9.746 TFLOPS (1:2)
1,431.4 GFLOPS (1:64)
FP16 (TFLOPS)
77.97 TFLOPS (4:1)
91.61 TFLOPS (1:1)
AI/RT
RT Cores
—
142
Tensor Cores
432
568 +31.5%
BF16
311.84 TFLOPS (16:1)
—
TF32
155.92 TFLOPs (8:1)
—
Power
TDP
300 W
300 W
TDP (W)
300
300 0.0%
Suggested PSU
700 W
700 W
Power Connectors
8-pin EPS
1x 16-pin
Architecture
Architecture
Ampere
Ada Lovelace
GPU Name
GA100
AD102
Generation
Server Ampere (Axx)
Server Ada (Lxx)
Process Size
7 nm
5 nm
Transistors
54,200 million
76,300 million
Die Size
826 mm²
609 mm²
Foundry
TSMC
TSMC
Density
65.6M / mm²
125.3M / mm²
API Support
DirectX
—
12 Ultimate (12_2)
OpenGL
—
4.6
Vulkan
—
1.4
OpenCL
3.0
3.0
CUDA
8.0
8.9
Shader Model
—
6.8
Physical
Slot Width
Dual-slot
Dual-slot
Length
267 mm 10.5 inches
267 mm 10.5 inches
Height
111 mm 4.4 inches
111 mm 4.4 inches
Outputs
No outputs
1x HDMI 2.13x DisplayPort 1.4a
Bus Interface
PCIe 4.0 x16
PCIe 4.0 x16
Other
Production
End-of-life
End-of-life
Predecessor
Tesla Turing
Server Ampere
Successor
Server Ada
Server Hopper
View A100 PCIe 80 GB Details View L40S Details