NVIDIA L40S vs NVIDIA RTX 6000 Ada Generation Comparison

NVIDIA
GEFORCE

NVIDIA L40S

CORE STATE AD102
VRAM 48 GB
CLOCK SPEED 2520 MHz
TDP 300 W
BUS WIDTH 384 bit
ARCHITECTURE Ada Lovelace
nm
PROCESS 5 nm
LAUNCH DATE 2022
VS
NVIDIA
GEFORCE

RTX 6000 Ada Generation

CORE STATE AD102
VRAM 48 GB
CLOCK SPEED 2505 MHz
TDP 300 W
BUS WIDTH 384 bit
ARCHITECTURE Ada Lovelace
nm
PROCESS 5 nm
LAUNCH DATE 2022

PERFORMANCE BENCHMARKS

geekbench_opencl
330,727
311,629
geekbench_vulkan
260,799
262,845

Analysis: NVIDIA L40S vs NVIDIA RTX 6000 Ada Generation

The NVIDIA L40S and NVIDIA RTX 6000 Ada Generation are two remarkably similar GPUs, both built on the AD102 chip and Ada Lovelace architecture. The benchmark data shows a near-tie in average performance, but the specific test results reveal distinct strengths that could influence a purchasing decision. The L40S edges ahead in one major compute workload, while the RTX 6000 Ada counters with a narrow win in another, making the choice dependent on the exact application.

Head-to-Head Benchmarks

The most significant divergence appears in the Geekbench OpenCL test, where the L40S delivers a score of 334,437 compared to the RTX 6000 Ada's 311,629. This translates to a 7.3% advantage for the L40S, a substantial margin that suggests its higher boost clock and memory speed translate into measurable gains in OpenCL compute tasks. The L40S's average benchmark score of 292,603 further reinforces this, sitting 3.8% above the RTX 6000 Ada's average of 281,932. In the broader context of the L40S's nearest rivals, this OpenCL result is a key differentiator, as the L40S only trails the NVIDIA H200 NVL (average score 305,608) by 4.3% while leading the L40 by 3.9% and the L20 by 9.8%.

However, the story reverses in the Geekbench Vulkan test. Here, the RTX 6000 Ada Generation scores 252,235, narrowly beating the L40S's 250,769 by a mere 0.6%. While this delta is small, it is a consistent win for the workstation card. This suggests that in Vulkan-based workloads, the RTX 6000 Ada's slightly different clock behavior and memory configuration might offer a marginal performance edge. The RTX 6000 Ada's average score of 281,932 places it just 0.1% behind the L40 and 5.8% ahead of the L20, but the head-to-head data shows it is the L40S that poses the biggest challenge, with the RTX 6000 Ada trailing it by 3.6% in average score.

The wins are split evenly at one apiece, which is a telling statistic. It indicates that neither card is categorically faster; instead, the performance hierarchy shifts depending on the API or workload type. The L40S's larger win in OpenCL (7.3%) outweighs the RTX 6000 Ada's slim victory in Vulkan (0.6%) in terms of raw margin, but the Vulkan result shows the RTX 6000 Ada is not a pushover. For users focused on OpenCL compute, the L40S is clearly the stronger option, but for Vulkan rendering or compute, the RTX 6000 Ada holds its ground.

Architecture Differences

At their core, these two GPUs are near-identical twins. Both use the AD102 chip fabricated on a 5 nm process at TSMC, with exactly 76,300 million transistors and a die size of 609 mm². This results in a transistor density of 125.3M per mm² for both. The shading unit count is also identical at 18,176, as are the texture mapping units (568), raster output pipelines (192), ray tracing cores (142), and tensor cores (568). The fundamental compute resources are the same, meaning the architectural potential is theoretically equal.

The differences emerge in the clock speeds and memory configuration. The L40S has a base clock of 1110 MHz, which is notably higher than the RTX 6000 Ada's 915 MHz. However, the boost clocks are much closer, with the L40S at 2520 MHz and the RTX 6000 Ada at 2505 MHz. This higher base clock on the L40S likely contributes to its OpenCL advantage, as it maintains a higher frequency under sustained load. The memory clocks also differ: the L40S runs at 2250 MHz (18 Gbps effective), while the RTX 6000 Ada runs at 2500 MHz (20 Gbps effective). This gives the RTX 6000 Ada a memory bandwidth advantage of 960.0 GB/s compared to the L40S's 864.0 GB/s, a 96 GB/s difference that could explain its Vulkan win.

Both cards feature 48 GB of GDDR6 memory on a 384-bit bus, so capacity is not a differentiator. The pixel rate is nearly identical (483.8 GPixel/s for L40S vs 481.0 GPixel/s for RTX 6000 Ada), as is the texture rate (1,431.4 GTexel/s vs 1,422.8 GTexel/s). The FP32 and FP16 compute figures are also close, with the L40S at 91.61 TFLOPS and the RTX 6000 Ada at 91.06 TFLOPS. The L40S's slight edge here is again attributable to its higher boost clock. The production status is identical (End-of-life), and both cards are dual-slot with a 300 W TDP and a single 16-pin power connector, requiring a 700 W power supply.

Where Each One Wins

The L40S is the clear winner for OpenCL-based compute workloads. Its 7.3% lead in Geekbench OpenCL is the most significant performance gap between the two cards, and its higher base clock suggests it can sustain higher throughput in long-running compute tasks. The data indicates that for applications relying on OpenCL—such as certain scientific simulations, machine learning inference, or general-purpose GPU computing—the L40S is the more capable choice. Its average score of 292,603 also positions it as the stronger overall performer in the immediate comparison, as it leads the RTX 6000 Ada by 3.8% and the L40 by 3.9%. This makes it a better pick for users who need maximum raw compute throughput in OpenCL-accelerated environments.

The RTX 6000 Ada Generation wins in Vulkan, albeit by a narrow 0.6% margin. This could be significant for graphics rendering, game development, or real-time visualization workloads that leverage the Vulkan API. The card's higher memory bandwidth (960.0 GB/s) may be a factor here, as Vulkan tasks often involve heavy data movement. While the average score is lower than the L40S, the RTX 6000 Ada's performance in this specific test shows it is not outclassed. For workstation users who prioritize Vulkan-based applications, the RTX 6000 Ada offers a slight edge that could matter in professional environments where every frame or computation counts.

The Verdict

The data presents a nuanced picture. The NVIDIA L40S is the better choice for users whose primary workloads are OpenCL-based, as it offers a substantial 7.3% performance advantage in that test. Its higher base clock and slightly higher FP32 compute (91.61 vs 91.06 TFLOPS) make it the more attractive option for compute-heavy tasks. The L40S also holds a 3.8% lead in average benchmark score over the RTX 6000 Ada, reinforcing its position as the stronger overall performer in this head-to-head.

The NVIDIA RTX 6000 Ada Generation is the better pick for Vulkan-centric applications, where its 0.6% win and superior memory bandwidth (960.0 GB/s) could provide tangible benefits. It is also the more balanced card, with a narrower average score gap (3.6% behind the L40S) compared to the L40S's lead over other rivals. The RTX 6000 Ada's launch MSRP is 6,799 USD, which is a factor to consider, but the benchmark data shows it is a capable competitor.

Ultimately, the choice hinges on the software stack. If OpenCL performance is critical, the L40S is the data-backed winner. If Vulkan performance or memory bandwidth is more important, the RTX 6000 Ada is the safer bet. Both cards are end-of-life, but they remain powerful options in their respective niches.

FAQ

Q: Which GPU has a higher average benchmark score?

A: The NVIDIA L40S has an average benchmark score of 292,603, which is 3.8% higher than the NVIDIA RTX 6000 Ada Generation's average score of 281,932.

Q: How do the two cards compare in the Geekbench OpenCL test?

A: The L40S scores 334,437 in Geekbench OpenCL, while the RTX 6000 Ada scores 311,629, giving the L40S a 7.3% win.

Q: What is the difference in memory bandwidth between the two GPUs?

A: The RTX 6000 Ada Generation has a memory bandwidth of 960.0 GB/s, while the L40S has 864.0 GB/s, a difference of 96 GB/s in favor of the RTX 6000 Ada.

Q: Are the core counts identical between the L40S and RTX 6000 Ada?

A: Yes, both GPUs have 18,176 shading units, 568 TMUs, 192 ROPs, 142 ray tracing cores, and 568 tensor cores.

Q: Which GPU has a higher boost clock?

A: The L40S has a boost clock of 2520 MHz, which is slightly higher than the RTX 6000 Ada's boost clock of 2505 MHz.

Q: What is the production status of both cards?

A: Both the NVIDIA L40S and the NVIDIA RTX 6000 Ada Generation are listed as End-of-life.

Specification Differences

| Specification | NVIDIA L40S | NVIDIA RTX 6000 Ada Generation |

|---|---|---|

| Generation | Server Ada (Lxx) | Workstation Ada (x000A) |

| Base Clock | 1110 MHz | 915 MHz |

| Boost Clock | 2520 MHz | 2505 MHz |

| Memory Clock | 2250 MHz (18 Gbps effective) | 2500 MHz (20 Gbps effective) |

| Memory Bandwidth | 864.0 GB/s | 960.0 GB/s |

| FP32 Compute | 91.61 TFLOPS | 91.06 TFLOPS |

| FP16 Compute | 91.61 TFLOPS (1:1) | 91.06 TFLOPS (1:1) |

| Pixel Rate | 483.8 GPixel/s | 481.0 GPixel/s |

| Texture Rate | 1,431.4 GTexel/s | 1,422.8 GTexel/s |

| Display Outputs | 1x HDMI 2.1, 3x DisplayPort 1.4a | 4x DisplayPort 1.4a |

| Height | 111 mm | 112 mm |

| Release Date | 2022-10-12 | 2022-12-02 |

| Predecessor | Server Ampere | Workstation Ampere |

| Successor | Server Hopper | Blackwell PRO W |

DETAILED SPECIFICATIONS

SPECIFICATION
L40S
RTX 6000 Ada Generation
Core Specs
Shading Units
18,176
18,176 0.0%
Shaders
18,176
18,176 0.0%
TMUs
568
568 0.0%
ROPs
192
192 0.0%
SM Count
142
142 0.0%
Clocks
Base Clock
1110 MHz
915 MHz
Boost Clock
2520 MHz
2505 MHz
Memory Clock
2250 MHz 18 Gbps effective
2500 MHz 20 Gbps effective
Memory
Memory Size
48 GB
48 GB
VRAM (MB)
49,152
49,152 0.0%
Memory Type
GDDR6
GDDR6
Memory Bus
384 bit
384 bit
Bandwidth
864.0 GB/s
960.0 GB/s
Cache
L1 Cache
128 KB (per SM)
128 KB (per SM)
L2 Cache
48 MB
96 MB
Performance
Pixel Rate
483.8 GPixel/s
481.0 GPixel/s
Texture Rate
1,431.4 GTexel/s
1,422.8 GTexel/s
FP32 (TFLOPS)
91.61 TFLOPS
91.06 TFLOPS
FP64 (TFLOPS)
1,431.4 GFLOPS (1:64)
1,422.8 GFLOPS (1:64)
FP16 (TFLOPS)
91.61 TFLOPS (1:1)
91.06 TFLOPS (1:1)
AI/RT
RT Cores
142
142 0.0%
Tensor Cores
568
568 0.0%
Power
TDP
300 W
300 W
TDP (W)
300
300 0.0%
Suggested PSU
700 W
700 W
Power Connectors
1x 16-pin
1x 16-pin
Architecture
Architecture
Ada Lovelace
Ada Lovelace
GPU Name
AD102
AD102
Generation
Server Ada (Lxx)
Workstation Ada (x000A)
Process Size
5 nm
5 nm
Transistors
76,300 million
76,300 million
Die Size
609 mm²
609 mm²
Foundry
TSMC
TSMC
Density
125.3M / mm²
125.3M / mm²
API Support
DirectX
12 Ultimate (12_2)
12 Ultimate (12_2)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
3.0
3.0
CUDA
8.9
8.9
Shader Model
6.8
6.8
Physical
Slot Width
Dual-slot
Dual-slot
Length
267 mm 10.5 inches
267 mm 10.5 inches
Height
111 mm 4.4 inches
112 mm 4.4 inches
Outputs
1x HDMI 2.13x DisplayPort 1.4a
4x DisplayPort 1.4a
Bus Interface
PCIe 4.0 x16
PCIe 4.0 x16
Other
Launch Price
6,799 USD
Production
End-of-life
End-of-life
Predecessor
Server Ampere
Workstation Ampere
Successor
Server Hopper
Blackwell PRO W
View L40S Details View RTX 6000 Ada Generation Details