NVIDIA A10G vs NVIDIA L40S Comparison

NVIDIA
GEFORCE

NVIDIA A10G

CORE STATE GA102
VRAM 24 GB
CLOCK SPEED 1710 MHz
TDP 150 W
BUS WIDTH 384 bit
ARCHITECTURE Ampere
nm
PROCESS 8 nm
LAUNCH DATE 2021
VS
NVIDIA
GEFORCE

L40S

CORE STATE AD102
VRAM 48 GB
CLOCK SPEED 2520 MHz
TDP 300 W
BUS WIDTH 384 bit
ARCHITECTURE Ada Lovelace
nm
PROCESS 5 nm
LAUNCH DATE 2022

PERFORMANCE BENCHMARKS

geekbench_opencl
158,063
330,727
geekbench_vulkan
145,863
260,799

Analysis: NVIDIA A10G vs NVIDIA L40S

# NVIDIA L40S vs NVIDIA A10G

The NVIDIA L40S and NVIDIA A10G represent two distinct generations of server-grade accelerators, and the benchmark data reflects a substantial generational gap. The L40S, built on the Ada Lovelace architecture, delivers an average benchmark score of 295,763, placing it in the 99th percentile of all GPUs. The A10G, based on the older Ampere architecture, achieves an average score of 151,963, sitting in the 97th percentile. The head-to-head results are unambiguous: the L40S wins both benchmark tests, with a 109.2% advantage in Geekbench OpenCL and a 78.8% lead in Geekbench Vulkan. This is not a close contest; it is a decisive generational leap.

The Verdict

The data paints a clear picture for different use cases. For workloads that demand maximum compute throughput, the NVIDIA L40S is the only rational choice between these two. Its OpenCL score of 330,727 more than doubles the A10G's 158,063, and its Vulkan score of 260,799 dwarfs the A10G's 145,863. The L40S also offers 48 GB of memory versus 24 GB, double the FP32 performance at 91.61 TFLOPS versus 31.52 TFLOPS, and nearly double the memory bandwidth at 864.0 GB/s versus 600.2 GB/s. Any application that is compute-bound or memory-bandwidth-bound will see massive gains on the L40S.

However, the A10G is not without its merits. Its 150 W TDP is exactly half of the L40S's 300 W, and it is a single-slot card versus the L40S's dual-slot design. For dense server deployments where power and physical space are constrained, the A10G offers a compelling efficiency profile. It also uses an 8-pin EPS power connector rather than the L40S's 16-pin connector, which may simplify integration into existing infrastructure. The A10G's suggested PSU of 450 W versus the L40S's 700 W further underscores its lower system-level demands.

The verdict from the data: choose the L40S for raw performance, larger memory capacity, and future-proofing. Choose the A10G for power-constrained environments, single-slot density, and workloads that do not require the L40S's extreme compute headroom. The A10G's benchmark scores, while lower, still place it in the 97th percentile, so it remains a capable accelerator for less demanding tasks.

Architecture Differences

The architectural divide between these two GPUs is fundamental. The L40S uses the AD102 chip built on TSMC's 5 nm process, packing 76,300 million transistors into a 609 mm² die. The A10G uses the GA102 chip on Samsung's 8 nm process, with 28,300 million transistors on a slightly larger 628 mm² die. This is a stark contrast: the L40S achieves a transistor density of 125.3 million per mm², nearly three times the A10G's 45.1 million per mm². The node advantage is the primary driver of the L40S's superior performance-per-watt and raw throughput.

The compute resources differ dramatically. The L40S fields 18,176 shading units, 568 TMUs, 192 ROPs, 142 RT cores, and 568 tensor cores. The A10G has 9,216 shading units, 288 TMUs, 96 ROPs, 72 RT cores, and 288 tensor cores — almost exactly half of the L40S's resources across every category. This doubling pattern extends to pixel rate (483.8 GPixel/s vs 164.2 GPixel/s) and texture rate (1,431.4 GTexel/s vs 492.5 GTexel/s). Both GPUs support DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4, but the L40S's Ada Lovelace architecture brings architectural improvements beyond raw counts.

Clock speeds tell an interesting story. The A10G has a higher base clock (1320 MHz vs 1110 MHz) but a much lower boost clock (1710 MHz vs 2520 MHz). The L40S's boost clock advantage of 810 MHz is substantial, and it explains how the L40S doubles the A10G's FP32 output despite the A10G starting at a higher base frequency. Memory clocks also favor the L40S: 2250 MHz (18 Gbps effective) versus 1563 MHz (12.5 Gbps effective), contributing to its bandwidth advantage.

Where Each One Wins

The L40S wins decisively in every benchmark and every compute metric. Its 91.61 TFLOPS FP32 performance is nearly triple the A10G's 31.52 TFLOPS. For FP16 workloads, both GPUs offer 1:1 ratios with their FP32 figures, meaning the L40S also triples FP16 performance. This makes the L40S the clear choice for AI inference, scientific computing, and any workload that stresses floating-point throughput. The 48 GB memory capacity is double the A10G's 24 GB, enabling larger datasets and models to reside in GPU memory without spilling to system RAM.

The A10G's wins are not in performance but in physical and power characteristics. At 150 W TDP, it consumes half the power of the L40S, and its single-slot design allows for denser server configurations. The A10G uses an 8-pin EPS connector, which is more common in server power supplies than the L40S's 16-pin connector. The A10G also has no display outputs, while the L40S includes 1x HDMI 2.1 and 3x DisplayPort 1.4a — but for server workloads, display outputs are rarely relevant. The A10G's lower power draw means less heat generation, which can be critical in tightly packed racks.

The release timeline also matters. The A10G launched in April 2021, while the L40S arrived in October 2022. Both are now end-of-life, but the L40S's later release means it benefits from a newer architecture generation. The A10G's predecessor is Tesla Turing, while the L40S's predecessor is Server Ampere — meaning the L40S represents a newer tier in NVIDIA's server lineup.

FAQ

Q: How much faster is the NVIDIA L40S than the A10G in OpenCL?

A: The L40S scores 330,727 in Geekbench OpenCL versus the A10G's 158,063, representing a 109.2% advantage — the L40S is more than twice as fast.

Q: What is the power consumption difference between these two GPUs?

A: The L40S has a 300 W TDP, while the A10G has a 150 W TDP. The A10G consumes exactly half the power, and its suggested PSU is 450 W versus 700 W for the L40S.

Q: Which GPU has more memory and bandwidth?

A: The L40S has 48 GB of GDDR6 memory on a 384-bit bus, delivering 864.0 GB/s bandwidth. The A10G has 24 GB of GDDR6 on the same 384-bit bus but only achieves 600.2 GB/s due to lower memory clocks.

Q: Are these GPUs still in production?

A: No. Both the NVIDIA L40S and NVIDIA A10G are listed as end-of-life products. The L40S was released in October 2022, and the A10G was released in April 2021.

Q: How do these GPUs compare to their nearest rivals?

A: The L40S is 3% ahead of the NVIDIA RTX 6000 Ada Generation and 4.1% ahead of the NVIDIA L40, but 7% behind the AMD Instinct MI300X and 11.7% behind the NVIDIA H200 NVL. The A10G is 1.1% ahead of the Tesla V100 PCIe 32 GB and 9.3% ahead of the AMD Instinct MI100, but 5.4% behind the AMD Radeon Pro W6800X and 6.5% behind the A100 PCIe 40 GB.

Q: What are the physical form factor differences?

A: The L40S is dual-slot with a 16-pin power connector, while the A10G is single-slot with an 8-pin EPS connector. Both are 267 mm long, with the L40S at 111 mm height and the A10G at 112 mm height.

Head-to-Head Benchmarks

The benchmark results are one-sided but revealing. In Geekbench OpenCL, the L40S scores 330,727 against the A10G's 158,063, yielding a 109.2% delta. This is the largest gap between the two GPUs in any test, and it reflects the L40S's massive advantage in shading units (18,176 vs 9,216), texture units (568 vs 288), and FP32 throughput (91.61 TFLOPS vs 31.52 TFLOPS). The OpenCL test likely stresses raw compute and memory bandwidth, where the L40S's 864.0 GB/s versus 600.2 GB/s provides additional headroom.

In Geekbench Vulkan, the L40S scores 260,799 versus the A10G's 145,863, a 78.8% advantage. The smaller delta compared to OpenCL suggests that Vulkan workloads may be more sensitive to factors where the gap is narrower, such as clock speeds or memory latency. Still, a 78.8% lead is massive. The L40S's higher boost clock (2520 MHz vs 1710 MHz) and more than double the ROPs (192 vs 96) likely contribute to its Vulkan dominance, as Vulkan tests often stress rasterization and pixel throughput.

The L40S's average benchmark score of 295,763 versus the A10G's 151,963 represents a 94.6% overall advantage. The L40S also sits at the 99th percentile of all GPUs, while the A10G is at the 97th percentile. This percentile difference is notable: both are high-performing accelerators, but the L40S is in the top 1% of all GPUs ever benchmarked, while the A10G is in the top 3%. The L40S's nearest rivals include the H200 NVL (11.7% faster) and the MI300X (7% faster), while the A10G's nearest rivals include the A100 PCIe 40 GB (6.5% faster) and the Radeon Pro W6800X (5.4% faster).

Specification Differences

The specification sheet reveals a consistent pattern of doubling. The L40S uses the AD102 chip on TSMC's 5 nm process, while the A10G uses the GA102 chip on Samsung's 8 nm process. Transistor counts are 76,300 million versus 28,300 million, and die sizes are 609 mm² versus 628 mm². The L40S achieves 125.3M transistors per mm², while the A10G manages only 45.1M.

Memory configuration: 48 GB GDDR6 at 2250 MHz (18 Gbps effective) with 864.0 GB/s bandwidth versus 24 GB GDDR6 at 1563 MHz (12.5 Gbps effective) with 600.2 GB/s bandwidth. Both use a 384-bit bus. Compute resources: 18,176 shading units, 568 TMUs, 192 ROPs, 142 RT cores, and 568 tensor cores for the L40S; 9,216 shading units, 288 TMUs, 96 ROPs, 72 RT cores, and 288 tensor cores for the A10G.

Clock speeds: L40S runs at 1110 MHz base and 2520 MHz boost; A10G runs at 1320 MHz base and 1710 MHz boost. Pixel rate is 483.8 GPixel/s versus 164.2 GPixel/s, and texture rate is 1,431.4 GTexel/s versus 492.5 GTexel/s. FP32 and FP16 performance are both 91.61 TFLOPS for the L40S and 31.52 TFLOPS for the A10G.

Power and physical: L40S has 300 W TDP, dual-slot, 16-pin power connector, 700 W suggested PSU. A10G has 150 W TDP, single-slot, 8-pin EPS connector, 450 W suggested PSU. Both are 267 mm long, with heights of 111 mm and 112 mm respectively. The L40S has display outputs (1x HDMI 2.1, 3x DisplayPort 1.4a), while the A10G has none. The L40S was released in October 2022, the A10G in April 2021. Both support identical API levels: DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4.

DETAILED SPECIFICATIONS

SPECIFICATION
A10G
L40S
Core Specs
Shading Units
9,216
18,176 +97.2%
Shaders
9,216
18,176 +97.2%
TMUs
288
568 +97.2%
ROPs
96
192 +100.0%
SM Count
72
142 +97.2%
Clocks
Base Clock
1320 MHz
1110 MHz
Boost Clock
1710 MHz
2520 MHz
Memory Clock
1563 MHz 12.5 Gbps effective
2250 MHz 18 Gbps effective
Memory
Memory Size
24 GB
48 GB
VRAM (MB)
24,576
49,152 +100.0%
Memory Type
GDDR6
GDDR6
Memory Bus
384 bit
384 bit
Bandwidth
600.2 GB/s
864.0 GB/s
Cache
L1 Cache
128 KB (per SM)
128 KB (per SM)
L2 Cache
6 MB
48 MB
Performance
Pixel Rate
164.2 GPixel/s
483.8 GPixel/s
Texture Rate
492.5 GTexel/s
1,431.4 GTexel/s
FP32 (TFLOPS)
31.52 TFLOPS
91.61 TFLOPS
FP64 (TFLOPS)
985.0 GFLOPS (1:32)
1,431.4 GFLOPS (1:64)
FP16 (TFLOPS)
31.52 TFLOPS (1:1)
91.61 TFLOPS (1:1)
AI/RT
RT Cores
72
142 +97.2%
Tensor Cores
288
568 +97.2%
Power
TDP
150 W
300 W
TDP (W)
150
300 +100.0%
Suggested PSU
450 W
700 W
Power Connectors
8-pin EPS
1x 16-pin
Architecture
Architecture
Ampere
Ada Lovelace
GPU Name
GA102
AD102
Generation
Server Ampere (Axx)
Server Ada (Lxx)
Process Size
8 nm
5 nm
Transistors
28,300 million
76,300 million
Die Size
628 mm²
609 mm²
Foundry
Samsung
TSMC
Density
45.1M / mm²
125.3M / mm²
API Support
DirectX
12 Ultimate (12_2)
12 Ultimate (12_2)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
3.0
3.0
CUDA
8.6
8.9
Shader Model
6.8
6.8
Physical
Slot Width
Single-slot
Dual-slot
Length
267 mm 10.5 inches
267 mm 10.5 inches
Height
112 mm 4.4 inches
111 mm 4.4 inches
Outputs
No outputs
1x HDMI 2.13x DisplayPort 1.4a
Bus Interface
PCIe 4.0 x16
PCIe 4.0 x16
Other
Production
End-of-life
End-of-life
Predecessor
Tesla Turing
Server Ampere
Successor
Server Ada
Server Hopper
View A10G Details View L40S Details