NVIDIA L40S vs NVIDIA Quadro GP100 Comparison

NVIDIA
GEFORCE

NVIDIA L40S

CORE STATE AD102
VRAM 48 GB
CLOCK SPEED 2520 MHz
TDP 300 W
BUS WIDTH 384 bit
ARCHITECTURE Ada Lovelace
nm
PROCESS 5 nm
LAUNCH DATE 2022
VS
NVIDIA
GEFORCE

Quadro GP100

CORE STATE GP100
VRAM 16 GB
CLOCK SPEED 1443 MHz
TDP 235 W
BUS WIDTH 4096 bit
ARCHITECTURE Pascal
nm
PROCESS 16 nm
LAUNCH DATE 2016

PERFORMANCE BENCHMARKS

geekbench_opencl
330,727
87,445
geekbench_vulkan
260,799
N/A

Analysis: NVIDIA L40S vs NVIDIA Quadro GP100

Head-to-Head Benchmarks

The recorded database contains a single common benchmark for these two accelerators: Geekbench OpenCL. In that test, the NVIDIA L40S scores 330,727, while the NVIDIA Quadro GP100 scores 87,445. This yields a delta of 278.2% in favor of the L40S. That is not a marginal lead; it is a decisive, multi-generational gap. The L40S delivers roughly 3.78 times the raw OpenCL compute throughput of the GP100. For context, the L40S sits at the 99th percentile among all GPUs in the database, while the GP100 sits at the 93rd percentile. Even though both are high-percentile parts, the absolute distance between them is enormous.

The L40S's nearest rivals in the database include the NVIDIA RTX 6000 Ada Generation (average score 287,237, delta 3% behind the L40S), the NVIDIA L40 (284,111, 4.1% behind), the AMD Instinct MI300X (317,994, 7% ahead of the L40S), and the NVIDIA H200 NVL (334,891, 11.7% ahead of the L40S). This shows the L40S is competitive with the top tier of current server accelerators, trading blows with AMD's MI300X and coming within roughly 12% of the H200 NVL. In contrast, the GP100's nearest rivals are far less demanding: the AMD Radeon PRO W7600 (87,108, 0.4% behind), the NVIDIA CMP 40HX (85,637, 2.1% behind), the NVIDIA RTX A4500 Mobile (91,134, 4% ahead), and the NVIDIA RTX A4500 (91,671, 4.6% ahead). The GP100 is essentially level with those mid-range workstation parts, which underscores how far behind the L40S it is in this specific workload.

There are no other overlapping benchmark entries in the database, so the head-to-head comparison rests entirely on this one OpenCL result. The win count is 1 for the L40S and 0 for the GP100. The data does not support any other quantitative comparison. The magnitude of the delta, however, is sufficient to characterize the performance relationship between the two cards: the L40S is in a different league for general-purpose compute as measured by OpenCL.

Architecture Differences

The L40S is built on the Ada Lovelace architecture, using the AD102 chip, fabricated on a 5 nm process at TSMC. It contains 76,300 million transistors on a 609 mm² die, yielding a transistor density of 125.3 million transistors per square millimeter. The GP100 uses the older Pascal architecture, with the GP100 chip, fabricated on a 16 nm process, also at TSMC. It contains 15,300 million transistors on a 610 mm² die, for a density of 25.1 million transistors per square millimeter. The die sizes are nearly identical (609 mm² vs. 610 mm²), but the L40S packs roughly five times as many transistors into the same physical area. That is the clearest architectural distinction: process node and transistor density.

The L40S features 18,176 shading units, 568 texture mapping units, 192 ROPs, 142 ray tracing cores, and 568 tensor cores. The GP100 has 3,584 shading units, 224 TMUs, 96 ROPs, and no ray tracing cores or tensor cores at all. The GP100 predates both RT and tensor core hardware, so it lacks those dedicated units entirely. This is a fundamental capability difference: the L40S can accelerate ray tracing and tensor workloads in hardware, while the GP100 cannot.

The FP32 and FP16 throughput numbers reflect the same gap. The L40S delivers 91.61 TFLOPS FP32 and 91.61 TFLOPS FP16 (at a 1:1 ratio). The GP100 delivers 10.34 TFLOPS FP32 and 20.69 TFLOPS FP16 (at a 2:1 ratio). So the L40S is about 8.9 times faster in FP32 and about 4.4 times faster in FP16, even accounting for the GP100's doubled FP16 rate. The L40S's FP16 performance matches its FP32 performance, while the GP100 trades off FP32 throughput to achieve its FP16 number.

Memory architecture also differs sharply. The L40S uses 48 GB of GDDR6 on a 384-bit bus, with 864.0 GB/s of bandwidth. The GP100 uses 16 GB of HBM2 on a 4096-bit bus, with 732.2 GB/s of bandwidth. The GP100 has a wider bus (4096-bit vs. 384-bit), but the L40S still achieves higher bandwidth thanks to faster memory clocks (2250 MHz with 18 Gbps effective vs. 715 MHz with 1430 Mbps effective). The L40S also has three times the memory capacity, which matters for large models and datasets.

Pixel and texture rates follow the same trend. The L40S reaches 483.8 GPixel/s and 1,431.4 GTexel/s, while the GP100 reaches 138.5 GPixel/s and 323.2 GTexel/s. The L40S is roughly 3.5 times faster in pixel fill and 4.4 times faster in texture fill.

The GP100 does have one advantage: its base clock is higher at 1304 MHz versus 1110 MHz for the L40S, and its boost clock is lower at 1443 MHz versus 2520 MHz. The L40S's boost clock is substantially higher, which is part of why its throughput numbers are so far ahead.

The Verdict

The data points to a clear conclusion: the NVIDIA L40S is the superior compute accelerator in every measurable aspect within the database. The OpenCL benchmark shows a 278.2% lead, and the architectural specifications confirm that lead across shading units, tensor cores, ray tracing cores, memory capacity, bandwidth, and process technology. The L40S is a modern server part aimed at high-end compute, while the GP100 is an older workstation part from the Pascal generation.

For anyone choosing between these two strictly on the recorded data, the L40S is the only rational pick for compute-intensive workloads. The GP100's only potential appeal is its lower power draw (235 W vs. 300 W) and its older HBM2 memory, which might be relevant in legacy systems. But the performance gap is so large that the L40S's higher power consumption is trivial in comparison. The GP100 also uses a simple 8-pin power connector and a 550 W suggested PSU, while the L40S uses a 16-pin connector and a 700 W suggested PSU. That is a real system requirement difference, but it does not offset a 278.2% benchmark deficit.

The GP100 could be considered only if the workload is limited to FP16 with a 2:1 ratio, but even then the L40S's FP16 output (91.61 TFLOPS) is 4.4 times higher than the GP100's (20.69 TFLOPS). There is no benchmark scenario in the database where the GP100 wins. The verdict is unambiguous.

Specification Differences

The following fields differ between the two cards, based on the recorded data:

  • Chip: AD102 (L40S) vs. GP100 (Quadro GP100)
  • Architecture: Ada Lovelace (L40S) vs. Pascal (Quadro GP100)
  • Generation: Server Ada (Lxx) for the L40S vs. Quadro Pascal (Px000) for the GP100
  • Process node: 5 nm (L40S) vs. 16 nm (GP100)
  • Transistors: 76,300 million (L40S) vs. 15,300 million (GP100)
  • Die size: 609 mm² (L40S) vs. 610 mm² (GP100)
  • Transistor density: 125.3M / mm² (L40S) vs. 25.1M / mm² (GP100)
  • Base clock: 1110 MHz (L40S) vs. 1304 MHz (GP100)
  • Boost clock: 2520 MHz (L40S) vs. 1443 MHz (GP100)
  • Memory clock: 2250 MHz, 18 Gbps effective (L40S) vs. 715 MHz, 1430 Mbps effective (GP100)
  • Memory size: 48 GB (L40S) vs. 16 GB (GP100)
  • Memory type: GDDR6 (L40S) vs. HBM2 (GP100)
  • Bus width: 384 bit (L40S) vs. 4096 bit (GP100)
  • Bandwidth: 864.0 GB/s (L40S) vs. 732.2 GB/s (GP100)
  • Shading units: 18,176 (L40S) vs. 3,584 (GP100)
  • TMUs: 568 (L40S) vs. 224 (GP100)
  • ROPs: 192 (L40S) vs. 96 (GP100)
  • RT cores: 142 (L40S) vs. none (GP100)
  • Tensor cores: 568 (L40S) vs. none (GP100)
  • Pixel rate: 483.8 GPixel/s (L40S) vs. 138.5 GPixel/s (GP100)
  • Texture rate: 1,431.4 GTexel/s (L40S) vs. 323.2 GTexel/s (GP100)
  • FP32: 91.61 TFLOPS (L40S) vs. 10.34 TFLOPS (GP100)
  • FP16: 91.61 TFLOPS, 1:1 (L40S) vs. 20.69 TFLOPS, 2:1 (GP100)
  • TDP: 300 W (L40S) vs. 235 W (GP100)
  • Power connectors: 1x 16-pin (L40S) vs. 1x 8-pin (GP100)
  • Suggested PSU: 700 W (L40S) vs. 550 W (GP100)
  • Bus interface: PCIe 4.0 x16 (L40S) vs. PCIe 3.0 x16 (GP100)
  • Display outputs: 1x HDMI 2.1, 3x DisplayPort 1.4a (L40S) vs. 1x DVI, 4x DisplayPort 1.4a (GP100)
  • DirectX support: 12 Ultimate (12_2) (L40S) vs. 12 (12_1) (GP100)
  • Vulkan support: 1.4 (L40S) vs. 1.3 (GP100)
  • Release date: 2022-10-12 (L40S) vs. 2016-09-30 (GP100)
  • Predecessor: Server Ampere (L40S) vs. Quadro Maxwell (GP100)
  • Successor: Server Hopper (L40S) vs. Quadro Volta (GP100)

Fields that are identical include manufacturer (NVIDIA), foundry (TSMC), slot width (Dual-slot), dimensions (267 mm length, 111 mm height), OpenGL version (4.6), and production status (End-of-life).

FAQ

Q: Which GPU has the higher OpenCL benchmark score?

A: The NVIDIA L40S scores 330,727 in Geekbench OpenCL, while the NVIDIA Quadro GP100 scores 87,445. The L40S leads by 278.2%.

Q: Does the Quadro GP100 have ray tracing cores?

A: No. The GP100 has no ray tracing cores and no tensor cores. The L40S has 142 ray tracing cores and 568 tensor cores.

Q: What is the memory capacity difference?

A: The L40S has 48 GB of GDDR6 memory, while the GP100 has 16 GB of HBM2 memory. The L40S also has higher bandwidth at 864.0 GB/s versus 732.2 GB/s.

Q: How do their FP32 compute performances compare?

A: The L40S delivers 91.61 TFLOPS FP32, while the GP100 delivers 10.34 TFLOPS FP32. The L40S is roughly 8.9 times faster in FP32.

Q: Are both cards still in production?

A: No. Both the L40S and the GP100 are marked as End-of-life in the database.

Q: Which card has the newer architecture?

A: The L40S uses the Ada Lovelace architecture on a 5 nm process, released on 2022-10-12. The GP100 uses the Pascal architecture on a 16 nm process, released on 2016-09-30.

Where Each One Wins

The L40S wins in every recorded benchmark and specification comparison. There is exactly one head-to-head test in the database, Geekbench OpenCL, and the L40S wins it decisively. The L40S also wins on compute throughput (FP32 and FP16), memory capacity, memory bandwidth, shading units, TMUs, ROPs, pixel rate, texture rate, and the presence of dedicated RT and tensor cores. It supports PCIe 4.0, DirectX 12 Ultimate, and Vulkan 1.4, all of which are newer than the GP100's PCIe 3.0, DirectX 12 (12_1), and Vulkan 1.3.

The GP100 does not win any benchmark in the database. It has a lower TDP (235 W versus 300 W), a lower suggested PSU (550 W versus 700 W), and a simpler 8-pin power connector. It also has a higher base clock (1304 MHz versus 1110 MHz), though its boost clock is far lower (1443 MHz versus 2520 MHz). The GP100's 4096-bit memory bus is wider, but it does not translate into higher bandwidth. For any use case that favors lower power draw or legacy PCIe 3.0 compatibility, the GP100 is technically distinct, but the data shows no workload where it outperforms the L40S.

In short, the L40S is the choice for modern compute, AI inference, ray tracing, and large-memory workloads. The GP100 is only relevant in legacy systems where its lower power draw or older interface is a requirement, and even then, the performance penalty is extreme. The database records one winner, and it is the L40S across the board.

DETAILED SPECIFICATIONS

SPECIFICATION
L40S
Quadro GP100
Core Specs
Shading Units
18,176
3,584 -80.3%
Shaders
18,176
3,584 -80.3%
TMUs
568
224 -60.6%
ROPs
192
96 -50.0%
SM Count
142
56 -60.6%
Clocks
Base Clock
1110 MHz
1304 MHz
Boost Clock
2520 MHz
1443 MHz
Memory Clock
2250 MHz 18 Gbps effective
715 MHz 1430 Mbps effective
Memory
Memory Size
48 GB
16 GB
VRAM (MB)
49,152
16,384 -66.7%
Memory Type
GDDR6
HBM2
Memory Bus
384 bit
4096 bit
Bandwidth
864.0 GB/s
732.2 GB/s
Cache
L1 Cache
128 KB (per SM)
24 KB (per SM)
L2 Cache
48 MB
4 MB
Performance
Pixel Rate
483.8 GPixel/s
138.5 GPixel/s
Texture Rate
1,431.4 GTexel/s
323.2 GTexel/s
FP32 (TFLOPS)
91.61 TFLOPS
10.34 TFLOPS
FP64 (TFLOPS)
1,431.4 GFLOPS (1:64)
5.172 TFLOPS (1:2)
FP16 (TFLOPS)
91.61 TFLOPS (1:1)
20.69 TFLOPS (2:1)
AI/RT
RT Cores
142
Tensor Cores
568
Power
TDP
300 W
235 W
TDP (W)
300
235 -21.7%
Suggested PSU
700 W
550 W
Power Connectors
1x 16-pin
1x 8-pin
Architecture
Architecture
Ada Lovelace
Pascal
GPU Name
AD102
GP100
Generation
Server Ada (Lxx)
Quadro Pascal (Px000)
Process Size
5 nm
16 nm
Transistors
76,300 million
15,300 million
Die Size
609 mm²
610 mm²
Foundry
TSMC
TSMC
Density
125.3M / mm²
25.1M / mm²
API Support
DirectX
12 Ultimate (12_2)
12 (12_1)
OpenGL
4.6
4.6
Vulkan
1.4
1.3
OpenCL
3.0
3.0
CUDA
8.9
6.0
Shader Model
6.8
6.0
Physical
Slot Width
Dual-slot
Dual-slot
Length
267 mm 10.5 inches
267 mm 10.5 inches
Height
111 mm 4.4 inches
111 mm 4.4 inches
Outputs
1x HDMI 2.13x DisplayPort 1.4a
1x DVI4x DisplayPort 1.4a
Bus Interface
PCIe 4.0 x16
PCIe 3.0 x16
Other
Production
End-of-life
End-of-life
Predecessor
Server Ampere
Quadro Maxwell
Successor
Server Hopper
Quadro Volta
View L40S Details View Quadro GP100 Details