GPU Comparison

NVIDIA
GEFORCE

NVIDIA A10G

CORE STATE GA102
VRAM 24 GB
CLOCK SPEED 1710 MHz
TDP 150 W
BUS WIDTH 384 bit
ARCHITECTURE Ampere
nm
PROCESS 8 nm
LAUNCH DATE 2021
VS
NVIDIA
GEFORCE

L40

CORE STATE AD102
VRAM 48 GB
CLOCK SPEED 2490 MHz
TDP 300 W
BUS WIDTH 384 bit
ARCHITECTURE Ada Lovelace
nm
PROCESS 5 nm
LAUNCH DATE 2022

PERFORMANCE BENCHMARKS

geekbench_opencl
158,063
330,926
geekbench_vulkan
145,863
237,295

Analysis: NVIDIA A10G vs NVIDIA L40

The NVIDIA L40 and NVIDIA A10G are both end-of-life server accelerators from NVIDIA, but they belong to different architectural generations and target very different performance tiers. The benchmark data is unambiguous: the L40 wins both head-to-head tests decisively, with an average benchmark score of 284,111 compared to the A10G’s 151,963. This places the L40 in the 99th percentile of all GPUs, while the A10G sits in the 97th percentile. The performance gap is not marginal; it represents a generational leap in compute capability, memory capacity, and raw throughput that fundamentally changes what each card is suitable for in a server environment.

Where Each One Wins

The L40 wins every benchmark category in the head-to-head comparison, making it the clear choice for compute-intensive workloads that demand maximum throughput. Its advantage is most pronounced in OpenCL compute tasks, where it scores 330,926 against the A10G’s 158,063, a staggering 109.4% improvement. This indicates that the L40 is built for heavy parallel processing, AI inference, and rendering tasks that can utilize its massive shading unit count and tensor core array. The data suggests the L40 is the appropriate solution for workloads where time-to-solution is critical, such as large-scale neural network training, scientific simulation, or high-resolution 3D rendering.

The A10G, while losing both head-to-head tests, still holds its own as a capable server accelerator in its own right. Its average benchmark score of 151,963 places it just 1.1% ahead of the NVIDIA Tesla V100 PCIe 32 GB and only 6.5% behind the NVIDIA A100 PCIe 40 GB. This positioning reveals that the A10G is not a weak performer, it competes effectively with the previous generation’s flagship accelerators. Where the A10G wins is in efficiency and form factor rather than raw performance. Its 150 W TDP and single-slot design make it suitable for dense server deployments where power and space are constrained. The data shows that for workloads that do not require the L40’s extreme compute capacity, the A10G provides a more efficient alternative without sacrificing compatibility with PCIe 4.0 infrastructure.

The L40’s wins are consistent across both API types, which suggests its advantage is architectural rather than workload-specific. In Vulkan, the L40 scores 237,295 versus the A10G’s 145,863, a 62.7% lead. This smaller but still substantial gap indicates that even in graphics-oriented tasks, the L40’s Ada Lovelace architecture delivers significant improvements over the A10G’s Ampere design.

Architecture Differences

The architectural divide between these two GPUs is stark. The L40 uses the AD102 chip built on a 5 nm process at TSMC, packing 76,300 million transistors into a 609 mm² die. This results in a transistor density of 125.3M per mm², nearly three times higher than the A10G’s density. The A10G, in contrast, uses the GA102 chip on Samsung’s 8 nm process, with 28,300 million transistors on a slightly larger 628 mm² die, yielding a density of just 45.1M per mm². This process advantage gives the L40 a fundamental edge in power efficiency and transistor budget.

The L40’s compute resources dwarf the A10G’s. It features 18,176 shading units, 568 texture mapping units, and 192 raster operation pipelines. The A10G has 9,216 shading units, 288 TMUs, and 96 ROPs, roughly half the L40’s resources across the board. The ray tracing and tensor core counts follow the same pattern: the L40 has 142 RT cores and 568 tensor cores, while the A10G has 72 RT cores and 288 tensor cores. These specifications translate directly into the FP32 and FP16 performance figures: the L40 delivers 90.52 TFLOPS in both precision modes, while the A10G delivers 31.52 TFLOPS. The L40’s 1:1 FP16 to FP32 ratio is notable, indicating it does not rely on reduced precision for its tensor operations.

Memory is another major differentiator. The L40 comes with 48 GB of GDDR6 memory on a 384-bit bus, providing 864.0 GB/s of bandwidth. The A10G has 24 GB of GDDR6 on the same 384-bit bus width, but its memory bandwidth is only 600.2 GB/s. The L40’s memory clock runs at 2250 MHz (18 Gbps effective), while the A10G’s runs at 1563 MHz (12.5 Gbps effective). This 44% bandwidth advantage for the L40 is critical for memory-bound workloads like large model inference or high-resolution texture streaming.

The clock speeds tell an interesting story. The A10G actually has higher base and boost clocks, 1320 MHz and 1710 MHz respectively, compared to the L40’s 735 MHz base and 2490 MHz boost. The L40’s lower base clock but much higher boost clock indicates a wider dynamic range, allowing it to ramp up aggressively under load while idling more efficiently. This is a hallmark of the Ada Lovelace architecture’s improved power management.

Head-to-Head Benchmarks

The Geekbench OpenCL test provides the most dramatic separation between the two cards. The L40 scores 330,926, which is more than double the A10G’s 158,063. This 109.4% delta is the largest in the entire benchmark suite and reflects the L40’s overwhelming advantage in raw compute throughput. To contextualize this, the L40’s average benchmark score of 284,111 puts it 13.1% ahead of the NVIDIA L20 and just 1.1% behind the NVIDIA RTX 6000 Ada Generation. The A10G’s average score of 151,963, by contrast, places it 5.4% behind the AMD Radeon Pro W6800X and 9.3% ahead of the AMD Instinct MI100.

In the Geekbench Vulkan test, the L40 maintains its lead but with a narrower margin. The L40 scores 237,295 against the A10G’s 145,863, a 62.7% delta. This test exercises graphics and compute APIs more heavily, and while the L40 still wins decisively, the smaller gap suggests that the A10G’s Ampere architecture handles Vulkan workloads relatively better than OpenCL. Still, the L40’s advantage is substantial enough that no interpretation of the data can make the A10G competitive in this metric.

The L40’s nearest rivals in the overall benchmark standings further illustrate its positioning. It sits 3.9% behind the NVIDIA L40S and 10.7% behind the AMD Instinct MI300X, indicating that it is at the upper edge of its performance tier but not the absolute fastest server GPU available. The A10G, meanwhile, is 1.1% ahead of the Tesla V100 PCIe 32 GB and 6.5% behind the A100 PCIe 40 GB, showing that it competes with the previous generation’s high-end parts rather than the current one.

FAQ

Q: Which GPU has higher raw compute performance?

A: The NVIDIA L40 delivers 90.52 TFLOPS in both FP32 and FP16 modes, while the NVIDIA A10G delivers 31.52 TFLOPS in both modes. The L40’s compute capacity is roughly three times higher.

Q: How do their memory configurations differ?

A: The L40 has 48 GB of GDDR6 memory with 864.0 GB/s bandwidth on a 384-bit bus. The A10G has 24 GB of GDDR6 with 600.2 GB/s bandwidth on a 384-bit bus. The L40 also runs its memory at 2250 MHz (18 Gbps effective) versus the A10G’s 1563 MHz (12.5 Gbps effective).

Q: What is the performance difference in OpenCL benchmarks?

A: The L40 scores 330,926 in Geekbench OpenCL, which is 109.4% higher than the A10G’s 158,063. This is the largest performance gap between the two cards across all tested benchmarks.

Q: Are both cards similar in physical size and power requirements?

A: Both cards are 267 mm in length and approximately 111-112 mm in height. However, the L40 is dual-slot with a 300 W TDP and requires a 700 W power supply, while the A10G is single-slot with a 150 W TDP and requires a 450 W power supply.

Q: How do these cards compare to their closest rivals?

A: The L40’s average score of 284,111 places it 13.1% ahead of the NVIDIA L20 and 1.1% behind the RTX 6000 Ada Generation. The A10G’s average score of 151,963 places it 1.1% ahead of the Tesla V100 PCIe 32 GB and 6.5% behind the A100 PCIe 40 GB.

Q: Which card has more shading units and tensor cores?

A: The L40 has 18,176 shading units, 568 TMUs, 192 ROPs, 142 RT cores, and 568 tensor cores. The A10G has 9,216 shading units, 288 TMUs, 96 ROPs, 72 RT cores, and 288 tensor cores.

Specification Differences

| Specification | NVIDIA L40 | NVIDIA A10G |

|---|---|---|

| Architecture | Ada Lovelace | Ampere |

| Process Node | 5 nm (TSMC) | 8 nm (Samsung) |

| Transistors | 76,300 million | 28,300 million |

| Die Size | 609 mm² | 628 mm² |

| Transistor Density | 125.3M / mm² | 45.1M / mm² |

| Base Clock | 735 MHz | 1320 MHz |

| Boost Clock | 2490 MHz | 1710 MHz |

| Memory Clock | 2250 MHz (18 Gbps effective) | 1563 MHz (12.5 Gbps effective) |

| Memory Size | 48 GB GDDR6 | 24 GB GDDR6 |

| Memory Bandwidth | 864.0 GB/s | 600.2 GB/s |

| Shading Units | 18,176 | 9,216 |

| TMUs | 568 | 288 |

| ROPs | 192 | 96 |

| RT Cores | 142 | 72 |

| Tensor Cores | 568 | 288 |

| Pixel Rate | 478.1 GPixel/s | 164.2 GPixel/s |

| Texture Rate | 1,414.3 GTexel/s | 492.5 GTexel/s |

| FP32 / FP16 | 90.52 TFLOPS (1:1) | 31.52 TFLOPS (1:1) |

| TDP | 300 W | 150 W |

| Slot Width | Dual-slot | Single-slot |

| Power Connector | 1x 16-pin | 8-pin EPS |

| Suggested PSU | 700 W | 450 W |

| Display Outputs | 4x DisplayPort 1.4a | No outputs |

| Release Date | 2022-10-12 | 2021-04-11 |

The Verdict

The data makes the decision straightforward for most use cases: the NVIDIA L40 is the superior performer in every benchmark category, with an average score 87% higher than the A10G. Its 48 GB memory capacity, 90.52 TFLOPS compute throughput, and 99th percentile ranking among all GPUs make it the appropriate choice for compute-heavy server workloads where maximum performance is the priority. The L40’s nearest rival is the RTX 6000 Ada Generation, which edges it out by just 1.1%, and it sits 13.1% ahead of the L20, confirming it is near the top of its performance tier.

The A10G, however, should not be dismissed. Its 150 W TDP and single-slot design make it a more efficient option for dense server configurations where power budgets and physical space are limited. Its performance is competitive with the Tesla V100 PCIe 32 GB and within 6.5% of the A100 PCIe 40 GB, meaning it can handle many production workloads from the previous generation without issue. The A10G’s 97th percentile ranking is respectable, and its lower power requirements could make it the more practical choice in multi-GPU deployments where the L40’s 300 W TDP would multiply across many cards.

The verdict: choose the L40 if your workloads are compute-bound and you need maximum throughput, memory capacity, or ray tracing performance. Choose the A10G if you prioritize power efficiency, density, or if your workloads are already satisfied by the A10G’s 31.52 TFLOPS and 24 GB memory. The L40 wins on raw performance, but the A10G wins on efficiency per watt. The benchmark data cannot tell you which trade-off is right for your specific deployment, but it clearly shows that the L40 is the more powerful accelerator in absolute terms.

DETAILED SPECIFICATIONS

SPECIFICATION
A10G
L40
Core Specs
Shading Units
9,216
18,176 +97.2%
Shaders
9,216
18,176 +97.2%
TMUs
288
568 +97.2%
ROPs
96
192 +100.0%
SM Count
72
142 +97.2%
Clocks
Base Clock
1320 MHz
735 MHz
Boost Clock
1710 MHz
2490 MHz
Memory Clock
1563 MHz 12.5 Gbps effective
2250 MHz 18 Gbps effective
Memory
Memory Size
24 GB
48 GB
VRAM (MB)
24,576
49,152 +100.0%
Memory Type
GDDR6
GDDR6
Memory Bus
384 bit
384 bit
Bandwidth
600.2 GB/s
864.0 GB/s
Cache
L1 Cache
128 KB (per SM)
128 KB (per SM)
L2 Cache
6 MB
96 MB
Performance
Pixel Rate
164.2 GPixel/s
478.1 GPixel/s
Texture Rate
492.5 GTexel/s
1,414.3 GTexel/s
FP32 (TFLOPS)
31.52 TFLOPS
90.52 TFLOPS
FP64 (TFLOPS)
985.0 GFLOPS (1:32)
1,414.3 GFLOPS (1:64)
FP16 (TFLOPS)
31.52 TFLOPS (1:1)
90.52 TFLOPS (1:1)
AI/RT
RT Cores
72
142 +97.2%
Tensor Cores
288
568 +97.2%
Power
TDP
150 W
300 W
TDP (W)
150
300 +100.0%
Suggested PSU
450 W
700 W
Power Connectors
8-pin EPS
1x 16-pin
Architecture
Architecture
Ampere
Ada Lovelace
GPU Name
GA102
AD102
Generation
Server Ampere (Axx)
Server Ada (Lxx)
Process Size
8 nm
5 nm
Transistors
28,300 million
76,300 million
Die Size
628 mm²
609 mm²
Foundry
Samsung
TSMC
Density
45.1M / mm²
125.3M / mm²
API Support
DirectX
12 Ultimate (12_2)
12 Ultimate (12_2)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
3.0
3.0
CUDA
8.6
8.9
Shader Model
6.8
6.8
Physical
Slot Width
Single-slot
Dual-slot
Length
267 mm 10.5 inches
267 mm 10.5 inches
Height
112 mm 4.4 inches
111 mm 4.4 inches
Outputs
No outputs
4x DisplayPort 1.4a
Bus Interface
PCIe 4.0 x16
PCIe 4.0 x16
Other
Production
End-of-life
End-of-life
Predecessor
Tesla Turing
Server Ampere
Successor
Server Ada
Server Hopper
View A10G Details View L40 Details