NVIDIA L20 vs NVIDIA Quadro GP100 Comparison
NVIDIA L20
Quadro GP100
PERFORMANCE BENCHMARKS
Analysis: NVIDIA L20 vs NVIDIA Quadro GP100
Head-to-Head Benchmarks
The benchmark data records a single head-to-head comparison between the NVIDIA L20 and the NVIDIA Quadro GP100, and the result is decisive. In the Geekbench OpenCL test, the L20 scores 274,276 points, while the Quadro GP100 scores 87,445 points. That is a delta of 213.7% in favor of the L20, meaning the L20 delivers more than three times the raw compute performance of the older Quadro card in this workload. There is no benchmark in the database where the Quadro GP100 comes out ahead; the L20 wins the only recorded head-to-head test, giving it a clean 1-0 record.
To put the L20's OpenCL score into context, the database shows it sits in the 99th percentile among all GPUs. Its average benchmark score across all recorded tests is 251,147. The nearest rivals around that average include the NVIDIA L40 at 284,111 (which is 11.6% faster), the NVIDIA RTX 6000 Ada Generation at 287,237 (12.6% faster), the NVIDIA PG506-232 at 225,124 (where the L20 is 11.6% ahead), and the AMD Radeon PRO W7900D at 219,827 (where the L20 is 14.2% ahead). So while the L20 is not the absolute top of the server GPU stack, it clearly outperforms a large portion of the field, including the Quadro GP100 by a wide margin.
The Quadro GP100, by contrast, scores 87,445 in OpenCL, which places it in the 93rd percentile of all GPUs. Its nearest rivals in the database are the AMD Radeon PRO W7600 at 87,108 (only 0.4% slower), the NVIDIA CMP 40HX at 85,637 (2.1% slower), the NVIDIA RTX A4500 Mobile at 91,134 (4% faster), and the NVIDIA RTX A4500 at 91,671 (4.6% faster). The Quadro GP100 is therefore in a completely different performance tier. The 213.7% delta between the L20 and the Quadro GP100 dwarfs the differences between either card and its nearest rivals; it is a generational leap, not a minor spec bump.
The Geekbench Vulkan result reinforces the L20's position. The L20 scores 228,018 in Vulkan, which is lower than its OpenCL score but still a very strong result. The Quadro GP100 has no recorded Vulkan benchmark in the database, so no direct comparison is possible there. However, the OpenCL result alone is sufficient to establish the L20 as the clear winner in raw compute.
Where Each One Wins
The L20 wins every category that has recorded data. In OpenCL, the L20's 274,276 is more than triple the Quadro GP100's 87,445. In Vulkan, the L20's 228,018 stands alone, as the Quadro GP100 has no Vulkan score. There is no workload in the database where the Quadro GP100 beats the L20.
For compute-heavy tasks such as rendering, simulation, or machine learning inference, the L20's massive FP32 throughput of 59.35 TFLOPS versus the Quadro GP100's 10.34 TFLOPS means the L20 is the obvious choice. The L20 also offers 48 GB of GDDR6 memory with 864.0 GB/s of bandwidth, compared to the Quadro GP100's 16 GB of HBM2 with 732.2 GB/s. While the Quadro GP100's HBM2 memory has a wider 4096-bit bus, the L20's higher clocked GDDR6 delivers more raw bandwidth. For memory capacity, the L20 offers triple the VRAM, which matters for large datasets and models that do not fit in 16 GB.
The only area where the Quadro GP100 shows a comparative advantage is in its FP16 performance ratio. The Quadro GP100 achieves 20.69 TFLOPS FP16 with a 2:1 ratio relative to FP32, whereas the L20 achieves 59.35 TFLOPS FP16 with a 1:1 ratio. In absolute terms, the L20 still delivers nearly three times the FP16 throughput, but the Quadro GP100's 2:1 ratio means it does not lose half its rate when switching to FP16, which is a minor efficiency point. Still, the L20's absolute numbers are so far ahead that this does not change the verdict.
Architecture Differences
The two cards are built on fundamentally different architectures from different eras. The L20 uses the AD102 chip based on Ada Lovelace architecture, fabricated on a 5 nm process at TSMC. The Quadro GP100 uses the GP100 chip based on Pascal architecture, fabricated on a 16 nm process, also at TSMC. The process node difference is stark: 5 nm versus 16 nm, which explains much of the performance and efficiency gap.
The transistor counts reflect this generational divide. The L20 packs 76,300 million transistors on a 609 mm² die, giving a transistor density of 125.3 million transistors per square millimeter. The Quadro GP100 has 15,300 million transistors on a 610 mm² die, yielding a density of 25.1 million per square millimeter. The die sizes are nearly identical (609 mm² vs 610 mm²), but the L20 crams five times more transistors into the same area thanks to the advanced process node.
The L20 features hardware that the Quadro GP100 simply does not have. The L20 includes 92 ray tracing cores and 368 tensor cores, while the Quadro GP100 has none of either. This means the L20 can accelerate ray-traced workloads and tensor-based operations (such as deep learning training and inference) in hardware, whereas the Quadro GP100 must rely on traditional shader cores for those tasks. The L20 also supports DirectX 12 Ultimate (12_2), while the Quadro GP100 only reaches DirectX 12 (12_1). Vulkan support is also newer on the L20 (version 1.4 versus 1.3 on the Quadro GP100). OpenGL support is identical at version 4.6.
The memory technology differs completely. The L20 uses GDDR6 with a 384-bit bus, while the Quadro GP100 uses HBM2 with a 4096-bit bus. HBM2's wide bus was an advanced feature in 2016, but the L20's GDDR6 runs at a much higher effective clock speed (18 Gbps effective versus 1430 Mbps effective), which is why the L20 achieves higher total bandwidth despite the narrower bus.
Specification Differences
The two cards differ across nearly every specification field. The L20 has 11,776 shading units, 368 texture mapping units, and 128 ROPs. The Quadro GP100 has 3,584 shading units, 224 TMUs, and 96 ROPs. The L20's pixel rate is 322.6 GPixel/s versus 138.5 GPixel/s for the Quadro GP100. The texture rate is 927.4 GTexel/s versus 323.2 GTexel/s.
Clock speeds also diverge significantly. The L20 has a base clock of 1440 MHz and a boost clock of 2520 MHz. The Quadro GP100 has a base clock of 1304 MHz and a boost clock of 1443 MHz. The L20's boost clock is nearly 75% higher.
Memory configuration: the L20 has 48 GB of GDDR6 on a 384-bit bus with 864.0 GB/s bandwidth. The Quadro GP100 has 16 GB of HBM2 on a 4096-bit bus with 732.2 GB/s bandwidth.
Power and connectivity: the L20 has a TDP of 275 W with a single 16-pin power connector and a suggested PSU of 600 W. The Quadro GP100 has a TDP of 235 W with a single 8-pin connector and a suggested PSU of 550 W. The L20 uses PCIe 4.0 x16, while the Quadro GP100 uses PCIe 3.0 x16. The L20 has 4x DisplayPort 1.4a outputs; the Quadro GP100 has 1x DVI plus 4x DisplayPort 1.4a outputs.
The production status differs as well: the L20 is listed as Active, while the Quadro GP100 is End-of-life. The release dates are far apart: the L20 was released in November 2023, while the Quadro GP100 was released in September 2016. The L20's predecessor is Server Ampere and its successor is Server Hopper. The Quadro GP100's predecessor is Quadro Maxwell and its successor is Quadro Volta.
FAQ
Q: Which card is faster in OpenCL?
A: The NVIDIA L20 scores 274,276 in Geekbench OpenCL, which is 213.7% higher than the Quadro GP100's 87,445.
Q: Does the Quadro GP100 win any benchmark?
A: No. The database records one head-to-head benchmark (Geekbench OpenCL), and the L20 wins it. The Quadro GP100 has no Vulkan score recorded.
Q: How much memory does each card have?
A: The L20 has 48 GB of GDDR6. The Quadro GP100 has 16 GB of HBM2.
Q: Does the Quadro GP100 support ray tracing or tensor cores?
A: No. The Quadro GP100 has no ray tracing cores and no tensor cores. The L20 has 92 ray tracing cores and 368 tensor cores.
Q: What is the transistor count difference?
A: The L20 has 76,300 million transistors, while the Quadro GP100 has 15,300 million transistors. Both are built by TSMC, but the L20 uses a 5 nm process versus 16 nm for the Quadro GP100.
Q: Which card has a higher boost clock?
A: The L20 boosts to 2520 MHz, while the Quadro GP100 boosts to 1443 MHz.
Q: What is the FP32 performance of each?
A: The L20 delivers 59.35 TFLOPS FP32. The Quadro GP100 delivers 10.34 TFLOPS FP32.
The Verdict
The data is unambiguous. The NVIDIA L20 is the superior card in every measured dimension. It is 213.7% faster in the only shared benchmark, it has triple the memory capacity, it has dramatically higher compute throughput in both FP32 and FP16, it adds ray tracing and tensor core hardware, it uses a modern 5 nm process, and it supports newer API versions. The L20 sits in the 99th percentile of all GPUs, while the Quadro GP100 sits in the 93rd percentile, which sounds close until you look at the actual scores: 251,147 average versus 87,445 average.
The Quadro GP100 is an end-of-life product from 2016. It has no ray tracing, no tensor cores, and a 16 nm process that limits its efficiency. Its only comparative advantages are a wider memory bus (4096-bit versus 384-bit) and a lower TDP (235 W versus 275 W). Neither of those translates into a performance win.
For any workload that stresses compute, memory capacity, or modern API features, the L20 is the clear choice. The Quadro GP100 might still be adequate for legacy applications that do not need much VRAM or throughput, but the database shows that even then, it is only 0.4% faster than an AMD Radeon PRO W7600, which is a much more modern competitor. The L20 is not the absolute fastest GPU in the database (the RTX 6000 Ada Generation is 12.6% faster), but it is far ahead of the Quadro GP100 and represents a sensible modern workhorse. Users who need maximum performance should look at the L40 or RTX 6000 Ada Generation; users who need a capable server GPU with strong compute and 48 GB of memory should choose the L20. The Quadro GP100 should only be considered if legacy software requires it, and even then, its end-of-life status is a risk.