NVIDIA A10M vs NVIDIA L40S Comparison
NVIDIA A10M
L40S
PERFORMANCE BENCHMARKS
Analysis: NVIDIA A10M vs NVIDIA L40S
Head-to-Head Benchmarks
The recorded data contains exactly one shared benchmark between the NVIDIA L40S and the NVIDIA A10M: Geekbench OpenCL. The L40S scores 330,727, while the A10M scores 135,230. That is a delta of 144.6% in favor of the L40S, meaning the L40S more than doubles the A10M’s OpenCL throughput. This is not a marginal gap; it is a categorical difference in compute capacity.
Context from the database’s nearest-rival tables sharpens the picture. The L40S averages 295,763 across its two recorded benchmarks (Geekbench OpenCL and Geekbench Vulkan), and it sits in the 99th percentile of all GPUs. Its closest rivals include the NVIDIA H200 NVL at 334,891 (11.7% ahead) and the AMD Instinct MI300X at 317,994 (7% ahead). Meanwhile, the A10M averages 135,230, placing it in the 96th percentile, with nearest rivals all within 0.9%: the NVIDIA RTX 4000 Ada Generation at 135,218, the AMD Radeon PRO W6800 at 135,396, the AMD Radeon Pro W6800X Duo at 135,774, and the AMD Radeon PRO V620 at 136,472. So the A10M is effectively tied with a cluster of midrange workstation cards, while the L40S is competing in a tier that includes flagship accelerators.
The wins tally is lopsided: the L40S records 1 win, the A10M records 0 wins. There is only one head-to-head test in the database, but the magnitude of that win is decisive. The L40S’s Vulkan score of 260,799 also exceeds the A10M’s OpenCL score by 92.9%, even though Vulkan is not a shared test. That internal comparison reinforces the notion that the L40S is operating in a completely different performance class.
What does the 144.6% delta imply in practical terms? For OpenCL workloads, the L40S delivers roughly 2.45 times the raw score of the A10M. Given that both cards share the same PCIe 4.0 x16 interface, the difference is not bus-bound; it reflects the underlying compute resources. The data suggests that any workload that scales with shading units, tensor cores, or memory bandwidth will see a dramatic improvement on the L40S.
FAQ
Q: Which GPU wins the only shared benchmark in the database?
A: The NVIDIA L40S wins Geekbench OpenCL with a score of 330,727 versus the NVIDIA A10M’s 135,230, a 144.6% advantage.
Q: How does the L40S compare to its nearest rivals?
A: The L40S is 3% ahead of the NVIDIA RTX 6000 Ada Generation (287,237), 4.1% ahead of the NVIDIA L40 (284,111), 7% behind the AMD Instinct MI300X (317,994), and 11.7% behind the NVIDIA H200 NVL (334,891). Its average benchmark score is 295,763.
Q: How does the A10M compare to its nearest rivals?
A: The A10M is effectively tied with its nearest rivals: 0% delta versus the NVIDIA RTX 4000 Ada Generation (135,218), -0.1% versus the AMD Radeon PRO W6800 (135,396), -0.4% versus the AMD Radeon Pro W6800X Duo (135,774), and -0.9% versus the AMD Radeon PRO V620 (136,472). Its average score is 135,230.
Q: What percentile do these GPUs occupy in the database?
A: The L40S is in the 99th percentile of all GPUs, while the A10M is in the 96th percentile. Despite the A10M’s high percentile, the L40S’s average score is more than double.
Q: Does the A10M have any benchmark where it beats the L40S?
A: No. The database records 1 win for the L40S and 0 wins for the A10M.
Q: What is the memory difference between the two?
A: The L40S has 48 GB of GDDR6 on a 384-bit bus with 864.0 GB/s bandwidth and 18 Gbps effective memory speed. The A10M has 20 GB of GDDR6 on a 320-bit bus with 500.2 GB/s bandwidth and 12.5 Gbps effective speed.
Architecture Differences
The L40S is built on NVIDIA’s Ada Lovelace architecture, using the AD102 chip fabricated by TSMC on a 5 nm process. The A10M uses the Ampere architecture with the GA102 chip, fabricated by Samsung on an 8 nm process. This node difference is significant: the L40S packs 76,300 million transistors into a 609 mm² die, yielding a transistor density of 125.3 million per mm². The A10M contains 28,300 million transistors on a 628 mm² die, for a density of 45.1 million per mm². The L40S achieves nearly three times the transistor density on a slightly smaller die.
The compute cores tell the same story. The L40S has 18,176 shading units, 568 TMUs, 192 ROPs, 142 RT cores, and 568 tensor cores. The A10M has 7,168 shading units, 224 TMUs, 80 ROPs, 56 RT cores, and 224 tensor cores. In every category, the L40S has between 2.5 and 2.6 times the resources of the A10M. That ratio aligns closely with the 144.6% OpenCL delta, suggesting the benchmark difference is driven by raw execution width.
Both GPUs support DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4, so API compatibility is identical. The L40S offers display outputs (1x HDMI 2.1 and 3x DisplayPort 1.4a), while the A10M has no display outputs at all. The L40S is a dual-slot card with a 16-pin power connector and a 300 W TDP, while the A10M is a single-slot card with an 8-pin EPS connector and a 150 W TDP. The L40S suggests a 700 W power supply, the A10M suggests 450 W.
The generations differ as well. The L40S belongs to the Server Ada generation (Lxx), succeeding Server Ampere and preceding Server Hopper. The A10M belongs to the Server Ampere generation (Axx), succeeding Tesla Turing and preceding Server Ada. This places the two cards on opposite sides of a generational boundary, which explains the architectural leap.
Specification Differences
The table below isolates only the fields where the two GPUs differ, based on the recorded data:
| Field | NVIDIA L40S | NVIDIA A10M |
|---|---|---|
| Architecture | Ada Lovelace | Ampere |
| Chip | AD102 | GA102 |
| Generation | Server Ada (Lxx) | Server Ampere (Axx) |
| Process node | 5 nm | 8 nm |
| Foundry | TSMC | Samsung |
| Transistors | 76,300 million | 28,300 million |
| Die size | 609 mm² | 628 mm² |
| Transistor density | 125.3M / mm² | 45.1M / mm² |
| Base clock | 1110 MHz | 975 MHz |
| Boost clock | 2520 MHz | 1635 MHz |
| Memory clock | 2250 MHz, 18 Gbps effective | 1563 MHz, 12.5 Gbps effective |
| Memory size | 48 GB | 20 GB |
| Memory bus width | 384 bit | 320 bit |
| Memory bandwidth | 864.0 GB/s | 500.2 GB/s |
| Shading units | 18,176 | 7,168 |
| TMUs | 568 | 224 |
| ROPs | 192 | 80 |
| RT cores | 142 | 56 |
| Tensor cores | 568 | 224 |
| Pixel rate | 483.8 GPixel/s | 130.8 GPixel/s |
| Texture rate | 1,431.4 GTexel/s | 366.2 GTexel/s |
| FP32 | 91.61 TFLOPS | 23.44 TFLOPS |
| FP16 | 91.61 TFLOPS (1:1) | 23.44 TFLOPS (1:1) |
| TDP | 300 W | 150 W |
| Slot width | Dual-slot | Single-slot |
| Power connectors | 1x 16-pin | 8-pin EPS |
| Suggested PSU | 700 W | 450 W |
| Display outputs | 1x HDMI 2.1, 3x DisplayPort 1.4a | No outputs |
| Release date | 2022-10-12 | (not recorded) |
| Predecessor | Server Ampere | Tesla Turing |
| Successor | Server Hopper | Server Ada |
| Geekbench OpenCL | 330,727 | 135,230 |
| Geekbench Vulkan | 260,799 | (not recorded) |
| Average benchmark score | 295,763 | 135,230 |
| Percentile vs all GPUs | 99 | 96 |
Notable differences beyond raw counts: the L40S has a much higher boost clock (2,520 MHz versus 1,635 MHz), which compounds its core-count advantage. The FP32 and FP16 rates are identical within each card (1:1 ratio), but the L40S delivers 91.61 TFLOPS in both, versus 23.44 TFLOPS for the A10M. The L40S also has a higher pixel rate (483.8 GPixel/s versus 130.8 GPixel/s) and texture rate (1,431.4 GTexel/s versus 366.2 GTexel/s). The A10M is physically thinner (single-slot versus dual-slot) and draws half the power, but it lacks any display connectivity.
The Verdict
The data points to a clear conclusion: the L40S is the superior compute card by a wide margin. Its average benchmark score of 295,763 is 118.7% higher than the A10M’s 135,230. In the only shared test, the L40S leads by 144.6%. The L40S also sits in the 99th percentile versus the A10M’s 96th, and its nearest rivals include the H200 NVL and Instinct MI300X, whereas the A10M’s nearest rivals are midrange workstation cards.
Who should pick the L40S? Anyone running OpenCL-heavy workloads, high-resolution rendering, or large tensor operations. The 48 GB memory pool and 864 GB/s bandwidth provide ample headroom for large datasets, and the 300 W TDP is modest relative to the performance. The dual-slot form factor and 16-pin connector are standard for server deployments.
Who should pick the A10M? The single-slot design and 150 W TDP make it attractive for dense multi-GPU configurations where power and space are constrained. Its 96th percentile ranking means it is still a strong performer in absolute terms, and it is effectively tied with the RTX 4000 Ada Generation and the Radeon PRO W6800. For workloads that fit within 20 GB of memory and do not require the L40S’s extreme throughput, the A10M remains a viable option.
But the data does not support any scenario where the A10M outperforms the L40S. The wins column is 1 to 0. The L40S wins in every recorded metric, and its architectural advantages (5 nm node, more cores, faster memory) are decisive. The only reasons to choose the A10M are physical constraints, not performance.
Where Each One Wins
NVIDIA L40S wins in:
- All compute-heavy tasks: OpenCL and Vulkan benchmarks show the L40S far ahead, with a 144.6% OpenCL advantage and a Vulkan score (260,799) that alone exceeds the A10M’s total average.
- Memory-intensive workloads: 48 GB versus 20 GB, 864 GB/s versus 500.2 GB/s, and a 384-bit bus versus 320-bit.
- High-precision floating point: FP32 and FP16 both at 91.61 TFLOPS versus 23.44 TFLOPS.
- Ray tracing and tensor operations: 142 RT cores versus 56, and 568 tensor cores versus 224.
- Display-connected tasks: the L40S has HDMI 2.1 and DisplayPort 1.4a outputs, while the A10M has none.
- Larger datasets that require more than 20 GB of VRAM.
NVIDIA A10M wins in:
- Power efficiency per watt: at 150 W versus 300 W, the A10M delivers 135,230 points at half the power draw. The L40S delivers roughly 2.45 times the score at twice the power, so the efficiency ratio is similar, but the A10M is the better choice for power-capped environments.
- Physical density: single-slot versus dual-slot, allowing more cards per chassis.
- Power connector compatibility: 8-pin EPS versus 16-pin, which may match existing infrastructure better.
- Lower PSU requirement: 450 W versus 700 W.
- Any workload that is memory-light and compute-light but benefits from a 96th-percentile GPU without the L40S’s cost in space and power.
The database does not show a single benchmark where the A10M wins outright, but its form factor and power profile carve out a niche. For a system builder prioritizing card count per server or staying within a 450 W power budget, the A10M is the logical pick. For anyone prioritizing raw compute, the L40S is the only choice supported by the measurements.