GPU Comparison
NVIDIA A100 PCIe 40 GB
L20
PERFORMANCE BENCHMARKS
Analysis: NVIDIA A100 PCIe 40 GB vs NVIDIA L20
The NVIDIA L20 and NVIDIA A100 PCIe 40 GB are both dual-slot server accelerators built for datacenter workloads, but they target very different points in the performance and feature spectrum. The L20 is an active Ada Lovelace-generation part, while the A100 is an end-of-life Ampere design. Benchmark data shows the L20 holds a commanding lead in both recorded tests, but the A100’s strengths lie in memory bandwidth and FP16 throughput, which the available benchmark scores do not directly measure.
Where Each One Wins
The L20 wins decisively in the two benchmark tests recorded. In Geekbench OpenCL, the L20 scores 274,276 against the A100’s 178,627, a 53.5% advantage. In Geekbench Vulkan, the L20 scores 228,018 against 146,380, a 55.8% lead. These are not narrow margins; they represent a generational leap in raw compute throughput. The L20’s average benchmark score of 251,147 places it in the 99th percentile of all GPUs, while the A100’s 162,504 average is in the 97th percentile. The L20 also holds a 2–0 win count in the head-to-head comparison.
The A100’s wins are not in the benchmark suite but in its specification sheet. It offers 1.56 TB/s of memory bandwidth, versus the L20’s 864.0 GB/s. That is a 44.6% higher bandwidth figure for the A100. The A100 also delivers 77.97 TFLOPS of FP16 performance with a 4:1 ratio, compared to the L20’s 59.35 TFLOPS with a 1:1 ratio. For workloads that rely on memory-bound operations or heavily favor FP16 tensor math, the A100’s specifications suggest it retains an edge, despite losing every recorded benchmark.
The Verdict
For general compute and graphics workloads that show up in Geekbench, the NVIDIA L20 is the clear pick. It beats the A100 by over 50% in both OpenCL and Vulkan, and its 99th percentile standing among all GPUs is two points higher than the A100’s. The L20 also has a higher boost clock (2520 MHz vs 1410 MHz), more shading units (11776 vs 6912), and double the memory capacity (48 GB vs 40 GB). If your work involves rendering, simulation, or any task that scales with raw shader throughput, the data points squarely at the L20.
The A100 is the choice only if your workload is specifically optimized for HBM2e memory bandwidth or FP16 tensor operations. Its 1.56 TB/s bandwidth is unmatched by the L20, and its 77.97 TFLOPS FP16 figure exceeds the L20’s 59.35 TFLOPS. The A100 also has more tensor cores (432 vs 368) and more ROPs (160 vs 128). For inference or training tasks that are memory-bound, the A100’s architecture may still be relevant, but its end-of-life status and lower benchmark scores make it a hard sell for new deployments.
Head-to-Head Benchmarks
The Geekbench OpenCL result is the largest single gap. The L20’s score of 274,276 is 95,649 points higher than the A100’s 178,627. That 53.5% delta is substantial and reflects the L20’s newer architecture and higher clock speeds. For context, the L20’s nearest rival below it, the NVIDIA PG506-232, scores 225,124, which is 11.6% lower. The A100’s closest rival, the AMD Radeon Pro W6800X, scores 160,671, just 1.1% lower than the A100. This suggests the A100 is tightly grouped with its immediate peers, while the L20 is in a different performance tier.
The Vulkan result tells a similar story. The L20 scores 228,018, which is 81,638 points above the A100’s 146,380, a 55.8% delta. The L20’s Vulkan score is closer to its OpenCL score, while the A100’s Vulkan score is significantly lower than its OpenCL score, indicating the A100 may not be as well optimized for Vulkan workloads. The L20’s nearest rival above it, the NVIDIA L40, scores 284,111, which is 11.6% higher, while the NVIDIA RTX 6000 Ada Generation scores 287,237, 12.6% higher. The A100’s nearest rival above it, the NVIDIA RTX 4500 Ada Generation, scores 166,094, just 2.2% higher, again showing the A100’s close competition at its performance level.
FAQ
Q: Which GPU has the higher average benchmark score?
A: The NVIDIA L20 has an average benchmark score of 251,147, which is 54.5% higher than the A100’s 162,504. The L20 sits in the 99th percentile of all GPUs, while the A100 is in the 97th.
Q: Does the A100 have any performance advantage at all?
A: Yes, in specifications. The A100 offers 1.56 TB/s memory bandwidth versus the L20’s 864.0 GB/s, and 77.97 TFLOPS FP16 performance versus the L20’s 59.35 TFLOPS. These figures are not reflected in the Geekbench scores.
Q: Which GPU has more memory?
A: The L20 has 48 GB of GDDR6 memory, while the A100 has 40 GB of HBM2e. The L20’s memory is on a 384-bit bus, while the A100 uses a 5120-bit bus.
Q: What are the clock speed differences?
A: The L20 has a base clock of 1440 MHz and a boost clock of 2520 MHz. The A100 has a base clock of 765 MHz and a boost clock of 1410 MHz. The L20’s boost clock is 78.7% higher.
Q: Which GPU has more shading units?
A: The L20 has 11,776 shading units, 368 TMUs, and 128 ROPs. The A100 has 6,912 shading units, 432 TMUs, and 160 ROPs. The L20 has 70.4% more shading units, but the A100 has 17.4% more TMUs and 25% more ROPs.
Q: Is the A100 still in production?
A: No. The A100 PCIe 40 GB is marked as end-of-life, while the L20 is active. The A100 was released in June 2020, while the L20 came in November 2023.
Architecture Differences
The L20 is built on the Ada Lovelace architecture using the AD102 chip, fabricated on a 5 nm process at TSMC. It packs 76,300 million transistors on a 609 mm² die, yielding a transistor density of 125.3 million per mm². The A100 uses the Ampere architecture with the GA100 chip, built on a 7 nm process, also at TSMC. It contains 54,200 million transistors on a larger 826 mm² die, giving a lower density of 65.6 million per mm². The L20’s newer node allows for more transistors in a smaller area, which explains its higher clock speeds and shader counts.
The L20 includes 92 RT cores and 368 tensor cores. The A100 has 432 tensor cores but no RT cores listed. The L20 supports DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4, while the A100 has no API support listed. The L20 also has four DisplayPort 1.4a outputs, whereas the A100 has no display outputs at all. The L20’s FP32 throughput is 59.35 TFLOPS, and its FP16 throughput is the same at 59.35 TFLOPS with a 1:1 ratio. The A100’s FP32 is 19.49 TFLOPS, but its FP16 jumps to 77.97 TFLOPS with a 4:1 ratio. The L20’s pixel rate is 322.6 GPixel/s and texture rate is 927.4 GTexel/s, compared to the A100’s 225.6 GPixel/s and 609.1 GTexel/s.
Specification Differences
The most obvious difference is memory. The L20 has 48 GB of GDDR6 on a 384-bit bus, while the A100 has 40 GB of HBM2e on a 5120-bit bus. The A100’s bandwidth is 1.56 TB/s, far exceeding the L20’s 864.0 GB/s. Clock speeds differ greatly: the L20 runs at 1440 MHz base and 2520 MHz boost, while the A100 runs at 765 MHz base and 1410 MHz boost. The L20’s memory clock is 2250 MHz (18 Gbps effective), while the A100’s is 1215 MHz (2.4 Gbps effective). The L20 has 11,776 shading units, 368 TMUs, and 128 ROPs, versus the A100’s 6,912 shading units, 432 TMUs, and 160 ROPs.
Power consumption is similar: the L20 has a TDP of 275 W, and the A100 is 250 W. Both are dual-slot cards with the same dimensions (267 mm length, 111 mm height). The L20 uses a 1x 16-pin power connector, while the A100 uses an 8-pin EPS connector. Both have a suggested PSU of 600 W and use a PCIe 4.0 x16 interface. The L20 has four DisplayPort 1.4a outputs; the A100 has none. The L20 is active and was released in November 2023, while the A100 is end-of-life and was released in June 2020. The L20’s predecessor is Server Ampere, and its successor is Server Hopper. The A100’s predecessor is Tesla Turing, and its successor is Server Ada.