NVIDIA A100 SXM4 40 GB vs NVIDIA L40 Comparison

NVIDIA
GEFORCE

NVIDIA A100 SXM4 40 GB

CORE STATE GA100
VRAM 40 GB
CLOCK SPEED 1410 MHz
TDP 400 W
BUS WIDTH 5120 bit
ARCHITECTURE Ampere
nm
PROCESS 7 nm
LAUNCH DATE 2020
VS
NVIDIA
GEFORCE

L40

CORE STATE AD102
VRAM 48 GB
CLOCK SPEED 2490 MHz
TDP 300 W
BUS WIDTH 384 bit
ARCHITECTURE Ada Lovelace
nm
PROCESS 5 nm
LAUNCH DATE 2022

PERFORMANCE BENCHMARKS

geekbench_opencl
201,096
330,926
geekbench_vulkan
173,198
237,295

Analysis: NVIDIA A100 SXM4 40 GB vs NVIDIA L40

The NVIDIA L40 and NVIDIA A100 SXM4 40 GB represent two distinct generations of server acceleration, and the benchmark data shows a decisive performance gap between them. The L40, built on the Ada Lovelace architecture, consistently outperforms the older Ampere-based A100 in the available tests, though the A100’s specialized memory subsystem and established server credentials make it a relevant comparison point for certain workloads. The data reveals that the L40 wins both head-to-head benchmarks, with the OpenCL test showing a particularly large margin.

Where Each One Wins

The L40 is the clear winner in raw compute throughput as measured by the two Geekbench tests. In the geekbench_opencl test, the L40 scores 330,926 against the A100’s 201,096, a delta of 64.6 percent. This suggests the L40 is substantially better suited for general-purpose compute tasks that leverage OpenCL, such as scientific simulation, data processing, and certain rendering workloads. The L40’s advantage extends to graphics-oriented tasks as well, with a geekbench_vulkan score of 237,295 versus 173,198 for the A100, representing a 37 percent lead. Vulkan support is notable because the A100 has no exposed graphics APIs at all—its DirectX, OpenGL, and Vulkan fields are null—so the L40 is the only one of the two that can handle real-time graphics or display output, reinforced by its four DisplayPort 1.4a outputs versus the A100’s none.

The A100 SXM4 40 GB does not win either benchmark, but its strengths lie elsewhere in the specification sheet. Its memory configuration is fundamentally different: 40 GB of HBM2e on a 5120-bit bus yields 1.56 TB/s of bandwidth, compared to the L40’s 48 GB of GDDR6 on a 384-bit bus delivering 864.0 GB/s. That 1.56 TB/s figure is a significant advantage for memory-bound workloads, even if the benchmark scores do not capture it. The A100 also has a higher base clock (1095 MHz vs 735 MHz) and a lower boost clock (1410 MHz vs 2490 MHz), suggesting it is tuned for sustained, power-limited operation rather than burst performance. The data indicates the L40 is the winner for compute-heavy tasks that fit within its memory capacity, while the A100’s bandwidth edge could make it preferable for specific data-intensive applications not represented in these two tests.

Architecture Differences

The two GPUs come from different architectural generations and process nodes. The L40 uses the AD102 chip on TSMC’s 5 nm process, packing 76,300 million transistors into a 609 mm² die, yielding a transistor density of 125.3M per mm². The A100 uses the GA100 chip on TSMC’s 7 nm process, with 54,200 million transistors on a much larger 826 mm² die, resulting in a lower density of 65.6M per mm². The L40’s newer process allows for significantly more transistors in a smaller area, which explains its higher compute throughput.

The shader and core configurations diverge sharply. The L40 has 18,176 shading units, 568 TMUs, and 192 ROPs, while the A100 has 6,912 shading units, 432 TMUs, and 160 ROPs. The L40 also includes 142 RT cores and 568 tensor cores, whereas the A100 lists 432 tensor cores and no RT cores at all. This means the L40 is equipped for ray tracing workloads that the A100 cannot accelerate at the hardware level. The L40’s FP32 throughput is 90.52 TFLOPS, while the A100 manages 19.49 TFLOPS in FP32—a 4.6x difference. For FP16, the L40 again shows 90.52 TFLOPS with a 1:1 ratio, while the A100 reaches 77.97 TFLOPS but at a 4:1 ratio, indicating the A100’s FP16 performance is achieved through a different execution path.

Memory technology is a major differentiator. The L40 uses GDDR6 with a 384-bit bus, while the A100 uses HBM2e with a 5120-bit bus. The A100’s bandwidth advantage (1.56 TB/s vs 864.0 GB/s) is a direct consequence of the wider bus and HBM2e design, which is optimized for high-bandwidth access patterns common in AI training and inference. The L40 compensates with higher effective memory clock speeds—18 Gbps effective versus 2.4 Gbps effective—but the bus width difference dominates. Power profiles also differ: the L40 is rated at 300 W TDP with a dual-slot design and a single 16-pin connector, while the A100 is a 400 W SXM module with no power connectors, designed for direct board integration.

Head-to-Head Benchmarks

The geekbench_opencl test shows the L40 at 330,926 versus the A100 at 201,096, a 64.6 percent advantage. This is the largest delta between the two in any benchmark, and it highlights the L40’s architectural efficiency. The L40’s 90.52 FP32 TFLOPS is more than 4.6 times the A100’s 19.49 TFLOPS, which likely drives this substantial OpenCL lead. The A100’s FP16 performance (77.97 TFLOPS) is closer, but the OpenCL test appears to favor FP32 or mixed-precision workloads where the L40’s raw shader count dominates.

In geekbench_vulkan, the L40 scores 237,295 against 173,198 for the A100, a 37 percent margin. This test is less one-sided because Vulkan is a graphics API, and the A100 has no Vulkan support listed—its API fields are all null. The fact that the A100 still produces a Vulkan score suggests the test may run through a compatibility layer, but the L40’s native support and 18,176 shading units give it a clear edge. The L40 also has a pixel rate of 478.1 GPixel/s and a texture rate of 1,414.3 GTexel/s, compared to the A100’s 225.6 GPixel/s and 609.1 GTexel/s, meaning the L40 should excel in rasterization and texture-heavy workloads.

The average benchmark score reinforces the L40’s dominance. The L40’s average is 284,111, placing it in the 99th percentile of all GPUs, while the A100’s average is 187,147, in the 98th percentile. The L40’s nearest rivals include the RTX 6000 Ada Generation (287,237, 1.1 percent behind) and the L40S (295,763, 3.9 percent ahead), showing it sits in a competitive tier. The A100’s nearest rivals are closer in relative terms: the RTX 5000 Ada Generation is 1.3 percent behind, and the A100 SXM4 80 GB is 1.9 percent ahead, indicating the A100 is near the bottom of its performance tier despite its high percentile ranking.

The Verdict

The data is unambiguous for general compute and graphics: the L40 is the superior choice. It wins both benchmarks by 64.6 percent and 37 percent, respectively, and its architecture provides features—RT cores, Vulkan support, DisplayPort outputs—that the A100 lacks entirely. The L40’s 48 GB of memory is also larger than the A100’s 40 GB, and while the A100 has higher bandwidth (1.56 TB/s vs 864.0 GB/s), the L40’s compute advantage likely outweighs that for most workloads. The L40 is the pick for any task involving FP32 compute, ray tracing, or graphics output.

The A100 SXM4 40 GB should be considered only for scenarios where its HBM2e bandwidth is critical. The 1.56 TB/s figure is nearly double the L40’s, and for memory-bound operations—such as large matrix multiplications or data shuffling in AI training—that bandwidth can be more important than raw FLOPs. The A100 also has a lower boost clock (1410 MHz vs 2490 MHz), which may indicate better power efficiency under sustained load, though its 400 W TDP is higher than the L40’s 300 W. Users with existing SXM4 infrastructure might prefer the A100 for compatibility, but the benchmark scores suggest they would sacrifice significant performance. The L40 is the better GPU for most users; the A100’s niche is narrow and specific.

FAQ

Q: Which GPU has higher FP32 performance?

A: The L40 delivers 90.52 TFLOPS FP32, while the A100 SXM4 40 GB provides 19.49 TFLOPS, making the L40 more than 4.6 times faster in this metric.

Q: Does the A100 support graphics APIs like Vulkan?

A: No. The A100’s DirectX, OpenGL, and Vulkan fields are all null, and it has no display outputs. The L40 supports DirectX 12 Ultimate, OpenGL 4.6, and Vulkan 1.4, with four DisplayPort 1.4a outputs.

Q: Which GPU has more memory bandwidth?

A: The A100 SXM4 40 GB has 1.56 TB/s of bandwidth from its HBM2e memory on a 5120-bit bus, compared to the L40’s 864.0 GB/s from GDDR6 on a 384-bit bus.

Q: What is the transistor count difference between the two?

A: The L40 has 76,300 million transistors on TSMC’s 5 nm process, while the A100 has 54,200 million transistors on TSMC’s 7 nm process. The L40’s die is smaller (609 mm² vs 826 mm²) but denser.

Q: Are there any benchmarks where the A100 wins?

A: No. In the two head-to-head tests (geekbench_opencl and geekbench_vulkan), the L40 wins both, with winsA set to 2 and winsB to 0.

Q: How do the two compare in ray tracing capability?

A: The L40 includes 142 RT cores, while the A100 has none listed. This makes the L40 the only one of the two capable of hardware-accelerated ray tracing.

Specification Differences

The following fields differ between the NVIDIA L40 and NVIDIA A100 SXM4 40 GB:

  • Chip: AD102 vs GA100
  • Architecture: Ada Lovelace vs Ampere
  • Generation: Server Ada (Lxx) vs Server Ampere (Axx)
  • Process Node: 5 nm vs 7 nm
  • Transistors: 76,300 million vs 54,200 million
  • Die Size: 609 mm² vs 826 mm²
  • Transistor Density: 125.3M / mm² vs 65.6M / mm²
  • Base Clock: 735 MHz vs 1095 MHz
  • Boost Clock: 2490 MHz vs 1410 MHz
  • Memory Clock: 2250 MHz (18 Gbps effective) vs 1215 MHz (2.4 Gbps effective)
  • Memory Size: 48 GB vs 40 GB
  • Memory Type: GDDR6 vs HBM2e
  • Memory Bus Width: 384 bit vs 5120 bit
  • Memory Bandwidth: 864.0 GB/s vs 1.56 TB/s
  • Shading Units: 18176 vs 6912
  • TMUs: 568 vs 432
  • ROPs: 192 vs 160
  • RT Cores: 142 vs null
  • Tensor Cores: 568 vs 432
  • Pixel Rate: 478.1 GPixel/s vs 225.6 GPixel/s
  • Texture Rate: 1,414.3 GTexel/s vs 609.1 GTexel/s
  • FP32: 90.52 TFLOPS vs 19.49 TFLOPS
  • FP16: 90.52 TFLOPS (1:1) vs 77.97 TFLOPS (4:1)
  • TDP: 300 W vs 400 W
  • Slot Width: Dual-slot vs SXM Module
  • Power Connectors: 1x 16-pin vs None
  • Suggested PSU: 700 W vs 800 W
  • Display Outputs: 4x DisplayPort 1.4a vs No outputs
  • DirectX: 12 Ultimate (12_2) vs null
  • OpenGL: 4.6 vs null
  • Vulkan: 1.4 vs null
  • Dimensions: 267 mm (10.5 inches) length, 111 mm (4.4 inches) height vs null
  • Release Date: 2022-10-12 vs 2020-05-13
  • Predecessor: Server Ampere vs Tesla Turing
  • Successor: Server Hopper vs Server Ada
  • Benchmark Scores: OpenCL 330,926 vs 201,096; Vulkan 237,295 vs 173,198
  • Percentile: 99 vs 98
  • Average Benchmark Score: 284,111 vs 187,147

DETAILED SPECIFICATIONS

SPECIFICATION
A100 SXM4 40 GB
L40
Core Specs
Shading Units
6,912
18,176 +163.0%
Shaders
6,912
18,176 +163.0%
TMUs
432
568 +31.5%
ROPs
160
192 +20.0%
SM Count
108
142 +31.5%
Clocks
Base Clock
1095 MHz
735 MHz
Boost Clock
1410 MHz
2490 MHz
Memory Clock
1215 MHz 2.4 Gbps effective
2250 MHz 18 Gbps effective
Memory
Memory Size
40 GB
48 GB
VRAM (MB)
40,960
49,152 +20.0%
Memory Type
HBM2e
GDDR6
Memory Bus
5120 bit
384 bit
Bandwidth
1.56 TB/s
864.0 GB/s
Cache
L1 Cache
192 KB (per SM)
128 KB (per SM)
L2 Cache
40 MB
96 MB
Performance
Pixel Rate
225.6 GPixel/s
478.1 GPixel/s
Texture Rate
609.1 GTexel/s
1,414.3 GTexel/s
FP32 (TFLOPS)
19.49 TFLOPS
90.52 TFLOPS
FP64 (TFLOPS)
9.746 TFLOPS (1:2)
1,414.3 GFLOPS (1:64)
FP16 (TFLOPS)
77.97 TFLOPS (4:1)
90.52 TFLOPS (1:1)
AI/RT
RT Cores
142
Tensor Cores
432
568 +31.5%
BF16
311.84 TFLOPS (16:1)
TF32
155.92 TFLOPs (8:1)
Power
TDP
400 W
300 W
TDP (W)
400
300 -25.0%
Suggested PSU
800 W
700 W
Power Connectors
None
1x 16-pin
Architecture
Architecture
Ampere
Ada Lovelace
GPU Name
GA100
AD102
Generation
Server Ampere (Axx)
Server Ada (Lxx)
Process Size
7 nm
5 nm
Transistors
54,200 million
76,300 million
Die Size
826 mm²
609 mm²
Foundry
TSMC
TSMC
Density
65.6M / mm²
125.3M / mm²
API Support
DirectX
12 Ultimate (12_2)
OpenGL
4.6
Vulkan
1.4
OpenCL
3.0
3.0
CUDA
8.0
8.9
Shader Model
6.8
Physical
Slot Width
SXM Module
Dual-slot
Length
267 mm 10.5 inches
Height
111 mm 4.4 inches
Outputs
No outputs
4x DisplayPort 1.4a
Bus Interface
PCIe 4.0 x16
PCIe 4.0 x16
Other
Production
End-of-life
End-of-life
Predecessor
Tesla Turing
Server Ampere
Successor
Server Ada
Server Hopper
View A100 SXM4 40 GB Details View L40 Details