NVIDIA A10M vs NVIDIA L40S Comparison

NVIDIA
GEFORCE

NVIDIA A10M

CORE STATE GA102
VRAM 20 GB
CLOCK SPEED 1635 MHz
TDP 150 W
BUS WIDTH 320 bit
ARCHITECTURE Ampere
nm
PROCESS 8 nm
LAUNCH DATE
VS
NVIDIA
GEFORCE

L40S

CORE STATE AD102
VRAM 48 GB
CLOCK SPEED 2520 MHz
TDP 300 W
BUS WIDTH 384 bit
ARCHITECTURE Ada Lovelace
nm
PROCESS 5 nm
LAUNCH DATE 2022

PERFORMANCE BENCHMARKS

geekbench_opencl
135,230
330,727
geekbench_vulkan
N/A
260,799

Analysis: NVIDIA A10M vs NVIDIA L40S

Head-to-Head Benchmarks

The recorded data contains exactly one shared benchmark between the NVIDIA L40S and the NVIDIA A10M: Geekbench OpenCL. The L40S scores 330,727, while the A10M scores 135,230. That is a delta of 144.6% in favor of the L40S, meaning the L40S more than doubles the A10M’s OpenCL throughput. This is not a marginal gap; it is a categorical difference in compute capacity.

Context from the database’s nearest-rival tables sharpens the picture. The L40S averages 295,763 across its two recorded benchmarks (Geekbench OpenCL and Geekbench Vulkan), and it sits in the 99th percentile of all GPUs. Its closest rivals include the NVIDIA H200 NVL at 334,891 (11.7% ahead) and the AMD Instinct MI300X at 317,994 (7% ahead). Meanwhile, the A10M averages 135,230, placing it in the 96th percentile, with nearest rivals all within 0.9%: the NVIDIA RTX 4000 Ada Generation at 135,218, the AMD Radeon PRO W6800 at 135,396, the AMD Radeon Pro W6800X Duo at 135,774, and the AMD Radeon PRO V620 at 136,472. So the A10M is effectively tied with a cluster of midrange workstation cards, while the L40S is competing in a tier that includes flagship accelerators.

The wins tally is lopsided: the L40S records 1 win, the A10M records 0 wins. There is only one head-to-head test in the database, but the magnitude of that win is decisive. The L40S’s Vulkan score of 260,799 also exceeds the A10M’s OpenCL score by 92.9%, even though Vulkan is not a shared test. That internal comparison reinforces the notion that the L40S is operating in a completely different performance class.

What does the 144.6% delta imply in practical terms? For OpenCL workloads, the L40S delivers roughly 2.45 times the raw score of the A10M. Given that both cards share the same PCIe 4.0 x16 interface, the difference is not bus-bound; it reflects the underlying compute resources. The data suggests that any workload that scales with shading units, tensor cores, or memory bandwidth will see a dramatic improvement on the L40S.

FAQ

Q: Which GPU wins the only shared benchmark in the database?

A: The NVIDIA L40S wins Geekbench OpenCL with a score of 330,727 versus the NVIDIA A10M’s 135,230, a 144.6% advantage.

Q: How does the L40S compare to its nearest rivals?

A: The L40S is 3% ahead of the NVIDIA RTX 6000 Ada Generation (287,237), 4.1% ahead of the NVIDIA L40 (284,111), 7% behind the AMD Instinct MI300X (317,994), and 11.7% behind the NVIDIA H200 NVL (334,891). Its average benchmark score is 295,763.

Q: How does the A10M compare to its nearest rivals?

A: The A10M is effectively tied with its nearest rivals: 0% delta versus the NVIDIA RTX 4000 Ada Generation (135,218), -0.1% versus the AMD Radeon PRO W6800 (135,396), -0.4% versus the AMD Radeon Pro W6800X Duo (135,774), and -0.9% versus the AMD Radeon PRO V620 (136,472). Its average score is 135,230.

Q: What percentile do these GPUs occupy in the database?

A: The L40S is in the 99th percentile of all GPUs, while the A10M is in the 96th percentile. Despite the A10M’s high percentile, the L40S’s average score is more than double.

Q: Does the A10M have any benchmark where it beats the L40S?

A: No. The database records 1 win for the L40S and 0 wins for the A10M.

Q: What is the memory difference between the two?

A: The L40S has 48 GB of GDDR6 on a 384-bit bus with 864.0 GB/s bandwidth and 18 Gbps effective memory speed. The A10M has 20 GB of GDDR6 on a 320-bit bus with 500.2 GB/s bandwidth and 12.5 Gbps effective speed.

Architecture Differences

The L40S is built on NVIDIA’s Ada Lovelace architecture, using the AD102 chip fabricated by TSMC on a 5 nm process. The A10M uses the Ampere architecture with the GA102 chip, fabricated by Samsung on an 8 nm process. This node difference is significant: the L40S packs 76,300 million transistors into a 609 mm² die, yielding a transistor density of 125.3 million per mm². The A10M contains 28,300 million transistors on a 628 mm² die, for a density of 45.1 million per mm². The L40S achieves nearly three times the transistor density on a slightly smaller die.

The compute cores tell the same story. The L40S has 18,176 shading units, 568 TMUs, 192 ROPs, 142 RT cores, and 568 tensor cores. The A10M has 7,168 shading units, 224 TMUs, 80 ROPs, 56 RT cores, and 224 tensor cores. In every category, the L40S has between 2.5 and 2.6 times the resources of the A10M. That ratio aligns closely with the 144.6% OpenCL delta, suggesting the benchmark difference is driven by raw execution width.

Both GPUs support DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4, so API compatibility is identical. The L40S offers display outputs (1x HDMI 2.1 and 3x DisplayPort 1.4a), while the A10M has no display outputs at all. The L40S is a dual-slot card with a 16-pin power connector and a 300 W TDP, while the A10M is a single-slot card with an 8-pin EPS connector and a 150 W TDP. The L40S suggests a 700 W power supply, the A10M suggests 450 W.

The generations differ as well. The L40S belongs to the Server Ada generation (Lxx), succeeding Server Ampere and preceding Server Hopper. The A10M belongs to the Server Ampere generation (Axx), succeeding Tesla Turing and preceding Server Ada. This places the two cards on opposite sides of a generational boundary, which explains the architectural leap.

Specification Differences

The table below isolates only the fields where the two GPUs differ, based on the recorded data:

| Field | NVIDIA L40S | NVIDIA A10M |

|---|---|---|

| Architecture | Ada Lovelace | Ampere |

| Chip | AD102 | GA102 |

| Generation | Server Ada (Lxx) | Server Ampere (Axx) |

| Process node | 5 nm | 8 nm |

| Foundry | TSMC | Samsung |

| Transistors | 76,300 million | 28,300 million |

| Die size | 609 mm² | 628 mm² |

| Transistor density | 125.3M / mm² | 45.1M / mm² |

| Base clock | 1110 MHz | 975 MHz |

| Boost clock | 2520 MHz | 1635 MHz |

| Memory clock | 2250 MHz, 18 Gbps effective | 1563 MHz, 12.5 Gbps effective |

| Memory size | 48 GB | 20 GB |

| Memory bus width | 384 bit | 320 bit |

| Memory bandwidth | 864.0 GB/s | 500.2 GB/s |

| Shading units | 18,176 | 7,168 |

| TMUs | 568 | 224 |

| ROPs | 192 | 80 |

| RT cores | 142 | 56 |

| Tensor cores | 568 | 224 |

| Pixel rate | 483.8 GPixel/s | 130.8 GPixel/s |

| Texture rate | 1,431.4 GTexel/s | 366.2 GTexel/s |

| FP32 | 91.61 TFLOPS | 23.44 TFLOPS |

| FP16 | 91.61 TFLOPS (1:1) | 23.44 TFLOPS (1:1) |

| TDP | 300 W | 150 W |

| Slot width | Dual-slot | Single-slot |

| Power connectors | 1x 16-pin | 8-pin EPS |

| Suggested PSU | 700 W | 450 W |

| Display outputs | 1x HDMI 2.1, 3x DisplayPort 1.4a | No outputs |

| Release date | 2022-10-12 | (not recorded) |

| Predecessor | Server Ampere | Tesla Turing |

| Successor | Server Hopper | Server Ada |

| Geekbench OpenCL | 330,727 | 135,230 |

| Geekbench Vulkan | 260,799 | (not recorded) |

| Average benchmark score | 295,763 | 135,230 |

| Percentile vs all GPUs | 99 | 96 |

Notable differences beyond raw counts: the L40S has a much higher boost clock (2,520 MHz versus 1,635 MHz), which compounds its core-count advantage. The FP32 and FP16 rates are identical within each card (1:1 ratio), but the L40S delivers 91.61 TFLOPS in both, versus 23.44 TFLOPS for the A10M. The L40S also has a higher pixel rate (483.8 GPixel/s versus 130.8 GPixel/s) and texture rate (1,431.4 GTexel/s versus 366.2 GTexel/s). The A10M is physically thinner (single-slot versus dual-slot) and draws half the power, but it lacks any display connectivity.

The Verdict

The data points to a clear conclusion: the L40S is the superior compute card by a wide margin. Its average benchmark score of 295,763 is 118.7% higher than the A10M’s 135,230. In the only shared test, the L40S leads by 144.6%. The L40S also sits in the 99th percentile versus the A10M’s 96th, and its nearest rivals include the H200 NVL and Instinct MI300X, whereas the A10M’s nearest rivals are midrange workstation cards.

Who should pick the L40S? Anyone running OpenCL-heavy workloads, high-resolution rendering, or large tensor operations. The 48 GB memory pool and 864 GB/s bandwidth provide ample headroom for large datasets, and the 300 W TDP is modest relative to the performance. The dual-slot form factor and 16-pin connector are standard for server deployments.

Who should pick the A10M? The single-slot design and 150 W TDP make it attractive for dense multi-GPU configurations where power and space are constrained. Its 96th percentile ranking means it is still a strong performer in absolute terms, and it is effectively tied with the RTX 4000 Ada Generation and the Radeon PRO W6800. For workloads that fit within 20 GB of memory and do not require the L40S’s extreme throughput, the A10M remains a viable option.

But the data does not support any scenario where the A10M outperforms the L40S. The wins column is 1 to 0. The L40S wins in every recorded metric, and its architectural advantages (5 nm node, more cores, faster memory) are decisive. The only reasons to choose the A10M are physical constraints, not performance.

Where Each One Wins

NVIDIA L40S wins in:

  • All compute-heavy tasks: OpenCL and Vulkan benchmarks show the L40S far ahead, with a 144.6% OpenCL advantage and a Vulkan score (260,799) that alone exceeds the A10M’s total average.
  • Memory-intensive workloads: 48 GB versus 20 GB, 864 GB/s versus 500.2 GB/s, and a 384-bit bus versus 320-bit.
  • High-precision floating point: FP32 and FP16 both at 91.61 TFLOPS versus 23.44 TFLOPS.
  • Ray tracing and tensor operations: 142 RT cores versus 56, and 568 tensor cores versus 224.
  • Display-connected tasks: the L40S has HDMI 2.1 and DisplayPort 1.4a outputs, while the A10M has none.
  • Larger datasets that require more than 20 GB of VRAM.

NVIDIA A10M wins in:

  • Power efficiency per watt: at 150 W versus 300 W, the A10M delivers 135,230 points at half the power draw. The L40S delivers roughly 2.45 times the score at twice the power, so the efficiency ratio is similar, but the A10M is the better choice for power-capped environments.
  • Physical density: single-slot versus dual-slot, allowing more cards per chassis.
  • Power connector compatibility: 8-pin EPS versus 16-pin, which may match existing infrastructure better.
  • Lower PSU requirement: 450 W versus 700 W.
  • Any workload that is memory-light and compute-light but benefits from a 96th-percentile GPU without the L40S’s cost in space and power.

The database does not show a single benchmark where the A10M wins outright, but its form factor and power profile carve out a niche. For a system builder prioritizing card count per server or staying within a 450 W power budget, the A10M is the logical pick. For anyone prioritizing raw compute, the L40S is the only choice supported by the measurements.

DETAILED SPECIFICATIONS

SPECIFICATION
A10M
L40S
Core Specs
Shading Units
7,168
18,176 +153.6%
Shaders
7,168
18,176 +153.6%
TMUs
224
568 +153.6%
ROPs
80
192 +140.0%
SM Count
56
142 +153.6%
Clocks
Base Clock
975 MHz
1110 MHz
Boost Clock
1635 MHz
2520 MHz
Memory Clock
1563 MHz 12.5 Gbps effective
2250 MHz 18 Gbps effective
Memory
Memory Size
20 GB
48 GB
VRAM (MB)
20,480
49,152 +140.0%
Memory Type
GDDR6
GDDR6
Memory Bus
320 bit
384 bit
Bandwidth
500.2 GB/s
864.0 GB/s
Cache
L1 Cache
128 KB (per SM)
128 KB (per SM)
L2 Cache
6 MB
48 MB
Performance
Pixel Rate
130.8 GPixel/s
483.8 GPixel/s
Texture Rate
366.2 GTexel/s
1,431.4 GTexel/s
FP32 (TFLOPS)
23.44 TFLOPS
91.61 TFLOPS
FP64 (TFLOPS)
732.5 GFLOPS (1:32)
1,431.4 GFLOPS (1:64)
FP16 (TFLOPS)
23.44 TFLOPS (1:1)
91.61 TFLOPS (1:1)
AI/RT
RT Cores
56
142 +153.6%
Tensor Cores
224
568 +153.6%
Power
TDP
150 W
300 W
TDP (W)
150
300 +100.0%
Suggested PSU
450 W
700 W
Power Connectors
8-pin EPS
1x 16-pin
Architecture
Architecture
Ampere
Ada Lovelace
GPU Name
GA102
AD102
Generation
Server Ampere (Axx)
Server Ada (Lxx)
Process Size
8 nm
5 nm
Transistors
28,300 million
76,300 million
Die Size
628 mm²
609 mm²
Foundry
Samsung
TSMC
Density
45.1M / mm²
125.3M / mm²
API Support
DirectX
12 Ultimate (12_2)
12 Ultimate (12_2)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
3.0
3.0
CUDA
8.6
8.9
Shader Model
6.8
6.8
Physical
Slot Width
Single-slot
Dual-slot
Length
267 mm 10.5 inches
267 mm 10.5 inches
Height
112 mm 4.4 inches
111 mm 4.4 inches
Outputs
No outputs
1x HDMI 2.13x DisplayPort 1.4a
Bus Interface
PCIe 4.0 x16
PCIe 4.0 x16
Other
Production
End-of-life
End-of-life
Predecessor
Tesla Turing
Server Ampere
Successor
Server Ada
Server Hopper
View A10M Details View L40S Details