GPU Comparison

NVIDIA
GEFORCE

NVIDIA A100 PCIe 40 GB

CORE STATE GA100
VRAM 40 GB
CLOCK SPEED 1410 MHz
TDP 250 W
BUS WIDTH 5120 bit
ARCHITECTURE Ampere
nm
PROCESS 7 nm
LAUNCH DATE 2020
VS
NVIDIA
GEFORCE

L40

CORE STATE AD102
VRAM 48 GB
CLOCK SPEED 2490 MHz
TDP 300 W
BUS WIDTH 384 bit
ARCHITECTURE Ada Lovelace
nm
PROCESS 5 nm
LAUNCH DATE 2022

PERFORMANCE BENCHMARKS

geekbench_opencl
178,627
330,926
geekbench_vulkan
146,380
237,295

Analysis: NVIDIA A100 PCIe 40 GB vs NVIDIA L40

The NVIDIA L40 and NVIDIA A100 PCIe 40 GB represent two distinct generations of NVIDIA's server accelerator lineup, with the data showing a clear performance hierarchy between them. The L40, built on the Ada Lovelace architecture, dominates the A100 across all recorded benchmark metrics, though the A100 retains relevance through its massive memory bandwidth and specialized compute characteristics.

Head-to-Head Benchmarks

The benchmark results paint a decisive picture: the NVIDIA L40 wins both recorded head-to-head comparisons, with a 2-0 sweep over the A100 PCIe 40 GB. In Geekbench OpenCL, the L40 scores 330,926 against the A100's 178,627, representing a massive 85.3% performance advantage. This is not a marginal lead, it is a near-doubling of raw compute throughput in a general-purpose GPU compute workload.

The Vulkan results tell a similar story, though with a slightly narrower gap. The L40 posts 237,295 points versus the A100's 146,380, a 62.1% delta in favor of the Ada Lovelace card. Both benchmarks consistently place the L40 ahead by a wide margin, indicating that its architectural advantages translate across different API environments.

Looking at the broader benchmark context, the L40's average benchmark score of 284,111 places it in the 99th percentile of all GPUs, while the A100's 162,504 average sits in the 97th percentile. The L40's nearest rivals include the NVIDIA RTX 6000 Ada Generation at 287,237 (1.1% higher) and the NVIDIA L40S at 295,763 (3.9% higher), showing that the L40 is competitive within its own generation. The A100's nearest rivals, the AMD Radeon Pro W6800X at 160,671 and the AMD Radeon PRO W7800 at 164,894, are separated by only 1-2%, indicating the A100 sits in a tightly contested mid-range tier by modern standards.

The delta between the two cards in OpenCL (85.3%) is substantially larger than the gap between the L40 and its closest competitor, the RTX 6000 Ada Generation, which trails by just 1.1%. This suggests the L40 is not merely ahead of the A100, it operates in an entirely different performance class.

Architecture Differences

The architectural divide between these two accelerators is generational and profound. The L40 uses the AD102 chip fabricated on TSMC's 5 nm process, while the A100 employs the GA100 chip on a 7 nm node. The process shrink enables the L40 to pack 76,300 million transistors into a 609 mm² die, achieving a transistor density of 125.3 million per square millimeter. The A100, by contrast, contains 54,200 million transistors across a larger 826 mm² die, with a density of just 65.6 million per square millimeter.

The L40's Ada Lovelace architecture brings fundamental compute advantages. It features 18,176 shading units, 568 TMUs, and 192 ROPs, compared to the A100's 6,912 shading units, 432 TMUs, and 160 ROPs. The L40 also includes 142 dedicated RT cores and 568 tensor cores, while the A100 lists tensor cores (432 of them) but no RT core specification. This makes the L40 a more complete accelerator for graphics and ray tracing workloads, whereas the A100 is positioned purely as a compute-focused server part.

Clock speeds further separate the pair. The L40 boosts to 2490 MHz from a 735 MHz base, while the A100 reaches only 1410 MHz from a 765 MHz base. This clock advantage compounds the L40's massive core count advantage, explaining the 90.52 TFLOPS FP32 throughput versus the A100's 19.49 TFLOPS. The FP16 comparison is more nuanced: the L40 delivers 90.52 TFLOPS at a 1:1 ratio, while the A100 offers 77.97 TFLOPS at a 4:1 ratio, meaning the A100's FP16 figure relies on specialized tensor operations rather than native throughput.

Memory architecture diverges sharply as well. The L40 uses 48 GB of GDDR6 on a 384-bit bus, delivering 864.0 GB/s of bandwidth. The A100 counters with 40 GB of HBM2e on a massive 5120-bit bus, achieving 1.56 TB/s, an 80% bandwidth advantage despite lower memory clocks. The A100's pixel rate of 225.6 GPixel/s and texture rate of 609.1 GTexel/s trail the L40's 478.1 GPixel/s and 1,414.3 GTexel/s respectively, reinforcing the L40's rasterization superiority.

Where Each One Wins

The L40 wins decisively in raw compute throughput, graphics performance, and modern API support. Its 90.52 TFLOPS FP32 performance is 4.6 times the A100's 19.49 TFLOPS, making it the clear choice for workloads that rely on single-precision floating-point math. The L40 also supports DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4, while the A100 lists no DirectX, OpenGL, or Vulkan support at all, a critical differentiator for any visualization or rendering task.

The L40's display outputs (4x DisplayPort 1.4a) versus the A100's complete lack of display outputs further cement its role as a graphics-capable accelerator. For AI inference and training, the L40's 568 tensor cores and higher FP16 throughput (90.52 TFLOPS at 1:1) provide substantial compute headroom, though the A100's 77.97 TFLOPS FP16 at 4:1 remains respectable.

The A100's wins are narrower but real. Its HBM2e memory subsystem with 1.56 TB/s bandwidth exceeds the L40's 864.0 GB/s by roughly 80%, making it superior for memory-bandwidth-bound workloads such as large sparse matrices or certain scientific simulations. The A100 also draws less power, 250 W versus 300 W, and uses a standard 8-pin EPS connector instead of the L40's 16-pin connector, potentially simplifying deployment in existing infrastructure. The A100's 40 GB of HBM2e memory, while smaller than the L40's 48 GB, offers higher per-byte bandwidth efficiency due to the 5120-bit bus.

For pure compute density, the L40's 5 nm process and higher clocks make it the efficiency leader in terms of performance per watt, but the A100's lower absolute power draw (250 W vs 300 W) and lower suggested PSU requirement (600 W vs 700 W) make it easier to integrate into power-constrained environments.

Specification Differences

The two cards diverge on nearly every measurable specification. The L40's 5 nm process node contrasts with the A100's 7 nm, and the transistor counts differ by over 22 billion (76,300 million vs 54,200 million). Die size actually favors the A100 at 826 mm² versus 609 mm², but the L40's superior density (125.3M/mm² vs 65.6M/mm²) shows the process advantage.

Clock speeds: the L40 boosts to 2490 MHz, the A100 to 1410 MHz. Memory: 48 GB GDDR6 at 864.0 GB/s versus 40 GB HBM2e at 1.56 TB/s. Core counts: 18,176 shading units, 568 TMUs, 192 ROPs, 142 RT cores, and 568 tensor cores for the L40; 6,912 shading units, 432 TMUs, 160 ROPs, and 432 tensor cores for the A100, with no RT cores listed.

Pixel rate: 478.1 GPixel/s versus 225.6 GPixel/s. Texture rate: 1,414.3 GTexel/s versus 609.1 GTexel/s. FP32: 90.52 TFLOPS versus 19.49 TFLOPS. FP16: 90.52 TFLOPS (1:1) versus 77.97 TFLOPS (4:1). Power: 300 W versus 250 W, with different connectors (16-pin vs 8-pin EPS). Display outputs: 4x DisplayPort 1.4a versus none. API support: the L40 lists DirectX 12 Ultimate, OpenGL 4.6, and Vulkan 1.4; the A100 lists none.

The L40 was released on 2022-10-12, succeeding Server Ampere and preceding Server Hopper. The A100 launched on 2020-06-21, succeeding Tesla Turing and preceding Server Ada. Both are end-of-life, dual-slot cards with identical physical dimensions (267 mm length, 111 mm height) and PCIe 4.0 x16 interfaces.

FAQ

Q: Which card is faster in OpenCL and Vulkan benchmarks?

A: The NVIDIA L40 wins both head-to-head tests. It scores 330,926 in Geekbench OpenCL versus 178,627 for the A100 (85.3% higher), and 237,295 in Vulkan versus 146,380 (62.1% higher).

Q: What is the memory capacity and bandwidth difference?

A: The L40 has 48 GB of GDDR6 with 864.0 GB/s bandwidth, while the A100 has 40 GB of HBM2e with 1.56 TB/s bandwidth. The A100 offers roughly 80% more bandwidth despite having less capacity.

Q: Does the A100 support graphics APIs?

A: No. The A100 lists no DirectX, OpenGL, or Vulkan support and has no display outputs. The L40 supports DirectX 12 Ultimate, OpenGL 4.6, and Vulkan 1.4, with 4x DisplayPort 1.4a outputs.

Q: How do the FP32 compute capabilities compare?

A: The L40 delivers 90.52 TFLOPS FP32, which is 4.6 times the A100's 19.49 TFLOPS. The L40's advantage comes from its 18,176 shading units and 2490 MHz boost clock versus 6,912 shading units and 1410 MHz.

Q: What are the power requirements for each card?

A: The L40 has a 300 W TDP with a 700 W suggested PSU and a 16-pin connector. The A100 has a 250 W TDP with a 600 W suggested PSU and an 8-pin EPS connector.

Q: Which card has better tensor core performance?

A: The L40 has 568 tensor cores and achieves 90.52 TFLOPS FP16 at a 1:1 ratio. The A100 has 432 tensor cores and achieves 77.97 TFLOPS FP16 at a 4:1 ratio, meaning the L40's FP16 performance is native while the A100's relies on tensor operations.

DETAILED SPECIFICATIONS

SPECIFICATION
A100 PCIe 40 GB
L40
Core Specs
Shading Units
6,912
18,176 +163.0%
Shaders
6,912
18,176 +163.0%
TMUs
432
568 +31.5%
ROPs
160
192 +20.0%
SM Count
108
142 +31.5%
Clocks
Base Clock
765 MHz
735 MHz
Boost Clock
1410 MHz
2490 MHz
Memory Clock
1215 MHz 2.4 Gbps effective
2250 MHz 18 Gbps effective
Memory
Memory Size
40 GB
48 GB
VRAM (MB)
40,960
49,152 +20.0%
Memory Type
HBM2e
GDDR6
Memory Bus
5120 bit
384 bit
Bandwidth
1.56 TB/s
864.0 GB/s
Cache
L1 Cache
192 KB (per SM)
128 KB (per SM)
L2 Cache
40 MB
96 MB
Performance
Pixel Rate
225.6 GPixel/s
478.1 GPixel/s
Texture Rate
609.1 GTexel/s
1,414.3 GTexel/s
FP32 (TFLOPS)
19.49 TFLOPS
90.52 TFLOPS
FP64 (TFLOPS)
9.746 TFLOPS (1:2)
1,414.3 GFLOPS (1:64)
FP16 (TFLOPS)
77.97 TFLOPS (4:1)
90.52 TFLOPS (1:1)
AI/RT
RT Cores
142
Tensor Cores
432
568 +31.5%
BF16
311.84 TFLOPS (16:1)
TF32
155.92 TFLOPs (8:1)
Power
TDP
250 W
300 W
TDP (W)
250
300 +20.0%
Suggested PSU
600 W
700 W
Power Connectors
8-pin EPS
1x 16-pin
Architecture
Architecture
Ampere
Ada Lovelace
GPU Name
GA100
AD102
Generation
Server Ampere (Axx)
Server Ada (Lxx)
Process Size
7 nm
5 nm
Transistors
54,200 million
76,300 million
Die Size
826 mm²
609 mm²
Foundry
TSMC
TSMC
Density
65.6M / mm²
125.3M / mm²
API Support
DirectX
12 Ultimate (12_2)
OpenGL
4.6
Vulkan
1.4
OpenCL
3.0
3.0
CUDA
8.0
8.9
Shader Model
6.8
Physical
Slot Width
Dual-slot
Dual-slot
Length
267 mm 10.5 inches
267 mm 10.5 inches
Height
111 mm 4.4 inches
111 mm 4.4 inches
Outputs
No outputs
4x DisplayPort 1.4a
Bus Interface
PCIe 4.0 x16
PCIe 4.0 x16
Other
Production
End-of-life
End-of-life
Predecessor
Tesla Turing
Server Ampere
Successor
Server Ada
Server Hopper
View A100 PCIe 40 GB Details View L40 Details