NVIDIA L40 vs NVIDIA Quadro GP100 Comparison
NVIDIA L40
Quadro GP100
PERFORMANCE BENCHMARKS
Analysis: NVIDIA L40 vs NVIDIA Quadro GP100
The NVIDIA L40 dominates the NVIDIA Quadro GP100 in the recorded data, and the gap is not close. In the Geekbench OpenCL test the two cards share, the L40 posts a score of 330926 against the GP100's 87445, a delta of 278.4 percent in the L40's favor. Both cards are end-of-life professional GPUs from NVIDIA, but they represent different eras of the company's workstation and server roadmap, and the database shows exactly how far that gap stretches in measurable compute performance.
Where Each One Wins
The L40 wins the only recorded head-to-head benchmark, and it wins decisively. That single data point, the Geekbench OpenCL result, carries the entire comparison: 330926 for the L40 versus 87445 for the GP100. With one win to zero in the recorded head-to-head, there is no benchmark category in the database where the Quadro GP100 comes out ahead.
The GP100's strengths are visible in its specifications rather than its scores. It runs at a lower TDP of 235 W compared with the L40's 300 W, needs only a single 8-pin power connector rather than the L40's 16-pin connector, and carries a lower suggested PSU rating of 550 W versus 700 W. Its HBM2 memory on a 4096-bit bus delivers 732.2 GB/s of bandwidth, which is remarkably close to the L40's 864.0 GB/s despite the L40 using far more modern GDDR6. For deployments where power delivery infrastructure is limited, the GP100's requirements are objectively lighter.
The L40, by contrast, wins everywhere performance matters. Its FP32 throughput of 90.52 TFLOPS is nearly nine times the GP100's 10.34 TFLOPS. Its FP16 throughput of 90.52 TFLOPS runs at a 1:1 ratio with FP32, while the GP100's FP16 at 20.69 TFLOPS uses a 2:1 ratio. The L40 also brings hardware the GP100 simply lacks: 142 RT cores and 568 tensor cores, compared with none of either on the Pascal-based card. Any workload involving ray tracing or tensor operations is uncontested.
Memory capacity is another clear L40 advantage: 48 GB versus 16 GB. For datasets, models, or scenes that exceed 16 GB, the GP100 cannot participate at all, regardless of its bandwidth credentials.
FAQ
Q: How much faster is the L40 than the Quadro GP100 in benchmarks?
A: In the Geekbench OpenCL test recorded in the database, the L40 scores 330926 against the GP100's 87445, a 278.4 percent advantage for the L40.
Q: Which card has more memory?
A: The L40, with 48 GB of GDDR6 on a 384-bit bus delivering 864.0 GB/s. The GP100 has 16 GB of HBM2 on a 4096-bit bus delivering 732.2 GB/s, so the L40 leads in both capacity and bandwidth.
Q: Do both cards support ray tracing and AI tensor operations?
A: Only the L40. It has 142 RT cores and 568 tensor cores. The GP100 has neither, as its Pascal architecture predates these dedicated units.
Q: Which card uses less power?
A: The GP100, with a 235 W TDP, a single 8-pin connector, and a suggested 550 W PSU. The L40 runs at 300 W, uses a 16-pin connector, and suggests a 700 W PSU.
Q: How do the two compare against other GPUs in the database?
A: The L40 sits in the 99th percentile against all GPUs, while the GP100 sits in the 93rd percentile. The L40's average score of 284111 places it within about one percent of the RTX 6000 Ada Generation, while the GP100's average of 87445 lands it essentially level with the AMD Radeon PRO W7600.
Q: Are both cards still in production?
A: No. Both are listed as end-of-life. The GP100 was released in 2016 and succeeded by the Quadro Volta generation, while the L40 was released in 2022 and succeeded by the Server Hopper generation.
Head-to-Head Benchmarks
The database contains one direct head-to-head result, and it is lopsided. In geekbench_opencl, the L40 scores 330926 and the GP100 scores 87445, making the L40 the winner by 278.4 percent. To put that in perspective, the GP100's score is not merely lower, it is closer to zero than to the L40's result by a wide margin.
The L40's separate Vulkan result of 237295 provides additional context. Even if the GP100 had matched the L40's Vulkan score, it would still trail the L40's OpenCL figure substantially, and its actual recorded OpenCL result of 87445 sits far below both. The GP100 has no Vulkan result recorded in the database, which itself reflects its older API support tier: Vulkan 1.3 versus the L40's Vulkan 1.4, and DirectX 12 feature level 12_1 versus the L40's 12 Ultimate (12_2).
The percentile data reinforces the scale of the gap. The L40's average benchmark score of 284111 puts it in the 99th percentile of all GPUs in the database, and its nearest rivals are all modern professional cards: the RTX 6000 Ada Generation at 287237 (a delta of -1.1 percent, meaning the L40 is within about one percent of it), the L40S at 295763 (-3.9 percent), the L20 at 251147 (13.1 percent in the L40's favor), and the AMD Instinct MI300X at 317994 (-10.7 percent). The L40 is a top-tier performer by any measure in the data.
The GP100's average score of 87445 places it in the 93rd percentile, which is respectable across the full database but reflects the breadth of entries more than competitive positioning at the top. Its rivals are telling: the AMD Radeon PRO W7600 at 87108 (0.4 percent), the NVIDIA CMP 40HX at 85637 (2.1 percent), the RTX A4500 Mobile at 91134 (-4 percent), and the RTX A4500 at 91671 (-4.6 percent). The GP100 trades blows with mid-tier professional and mining cards of more recent vintage, not with the L40's competitive set. The two cards occupy entirely different tiers of the performance stack.
Specification Differences
The specification deltas run in one direction. Shading units: 18176 on the L40 versus 3584 on the GP100. TMUs: 568 versus 224. ROPs: 192 versus 96. Pixel rate: 478.1 GPixel/s versus 138.5 GPixel/s. Texture rate: 1414.3 GTexel/s versus 323.2 GTexel/s. Every rendering-related metric favors the L40 by a factor of roughly three to four.
Memory differs in philosophy as much as in magnitude. The L40 uses 48 GB of GDDR6 across a 384-bit bus at 18 Gbps effective, producing 864.0 GB/s. The GP100 uses 16 GB of HBM2 across a 4096-bit bus at 1430 Mbps effective, producing 732.2 GB/s. The older card's exotic memory technology keeps it within reach on bandwidth, but capacity is tripled on the L40.
Clocks favor the L40 on boost: 2490 MHz versus 1443 MHz, though the GP100's base clock of 1304 MHz is higher than the L40's 735 MHz. FP32 throughput is 90.52 TFLOPS for the L40 versus 10.34 TFLOPS for the GP100; FP16 is 90.52 TFLOPS at a 1:1 ratio versus 20.69 TFLOPS at 2:1.
Power and platform differ. The L40 draws 300 W through a 16-pin connector over PCIe 4.0 x16 with a 700 W suggested PSU; the GP100 draws 235 W through an 8-pin connector over PCIe 3.0 x16 with a 550 W suggested PSU. Display outputs differ slightly: 4x DisplayPort 1.4a on the L40 versus 1x DVI plus 4x DisplayPort 1.4a on the GP100. Both are dual-slot cards of identical footprint, 267 mm long and 111 mm tall, so physical clearance requirements are the same.
Architecture Differences
The architectural distance between these cards is generational in every sense. The GP100 is built on TSMC's 16 nm process with 15,300 million transistors on a 610 mm² die, yielding a density of 25.1M per mm². The L40 uses TSMC's 5 nm node with 76,300 million transistors on a 609 mm² die, a density of 125.3M per mm². Nearly identical die sizes, but roughly five times the transistor count and density, which explains the performance chasm.
The L40 is an Ada Lovelace design (AD102) from the Server Ada generation, following Server Ampere and preceding Server Hopper. The GP100 is a Pascal design from the Quadro Pascal generation, following Quadro Maxwell and preceding Quadro Volta. Architecturally, the L40 adds dedicated RT cores (142) and tensor cores (568), neither of which exists on the GP100. API support reflects the era gap: DirectX 12 Ultimate (12_2), Vulkan 1.4, and OpenGL 4.6 on the L40, versus DirectX 12 (12_1), Vulkan 1.3, and OpenGL 4.6 on the GP100.
The Verdict
For any compute workload measured in the database, the L40 is the only rational choice: a 278.4 percent OpenCL advantage, roughly nine times the FP32 throughput, triple the memory capacity, and 99th-percentile positioning against the GP100's 93rd. Workloads needing tensor cores or RT cores exclude the GP100 entirely.
The GP100's case rests on lighter power requirements, a 235 W TDP with a single 8-pin connector and a 550 W PSU suggestion, and HBM2 bandwidth that remains competitive per the recorded 732.2 GB/s figure. Both cards are end-of-life, so neither represents a current purchase path. But in the data, the L40 wins every performance dimension recorded, and the GP100 wins none. The verdict is unanimous.