NVIDIA L40S vs NVIDIA RTX 4000 Ada Generation Comparison

NVIDIA
GEFORCE

NVIDIA L40S

CORE STATE AD102
VRAM 48 GB
CLOCK SPEED 2520 MHz
TDP 300 W
BUS WIDTH 384 bit
ARCHITECTURE Ada Lovelace
nm
PROCESS 5 nm
LAUNCH DATE 2022
VS
NVIDIA
GEFORCE

RTX 4000 Ada Generation

CORE STATE AD104
VRAM 20 GB
CLOCK SPEED 2175 MHz
TDP 130 W
BUS WIDTH 160 bit
ARCHITECTURE Ada Lovelace
nm
PROCESS 5 nm
LAUNCH DATE 2023

PERFORMANCE BENCHMARKS

geekbench_opencl
330,727
146,593
geekbench_vulkan
260,799
123,842

Analysis: NVIDIA L40S vs NVIDIA RTX 4000 Ada Generation

Where Each One Wins

The recorded benchmark data splits cleanly along compute-density lines. The NVIDIA L40S wins both recorded head-to-head tests, and by a very wide margin. In Geekbench OpenCL, the L40S scores 330727 against the RTX 4000 Ada Generation's 146593, a 125.6% advantage. In Geekbench Vulkan, the L40S posts 260799 versus 123842, a 110.6% lead. There is no recorded test where the RTX 4000 Ada Generation comes out ahead.

That said, the two cards are not aimed at the same workload profile, and the win tally does not tell the whole story. The L40S sits at the 99th percentile among all GPUs in the database, with an average benchmark score of 295763. Its nearest rivals are the NVIDIA H200 NVL (334891, 11.7% higher), the AMD Instinct MI300X (317994, 7% higher), the NVIDIA RTX 6000 Ada Generation (287237, 3% lower), and the NVIDIA L40 (284111, 4.1% lower). This places the L40S in the upper tier of server-class accelerators, trading blows with some of the largest memory-pooled compute parts available.

The RTX 4000 Ada Generation, by contrast, sits at the 95th percentile with an average score of 135218. Its nearest rivals are the NVIDIA A10M (135230, essentially tied at 0% delta), the AMD Radeon PRO W6800 (135396, 0.1% higher), the AMD Radeon Pro W6800X Duo (135774, 0.4% higher), and the AMD Radeon PRO V620 (136472, 0.9% higher). This is a tightly clustered midrange workstation segment where the RTX 4000 Ada Generation is effectively at parity with its immediate competitors. The data suggests this card wins on efficiency and form factor rather than raw throughput.

The practical use-case split is therefore: the L40S wins wherever raw compute density, memory capacity, and render throughput dominate, while the RTX 4000 Ada Generation wins in environments where the physical footprint, power envelope, and thermal load are the binding constraints. The benchmark results show that the L40S is more than twice as fast in both recorded APIs, but the RTX 4000 Ada Generation still occupies a distinct niche as a single-slot, low-power workstation card that can fit into dense chassis configurations.

Architecture Differences

Both GPUs are built on the Ada Lovelace architecture and both are fabricated by TSMC on a 5 nm process node. The similarities end at the chip level. The L40S uses the AD102 die, which contains 76,300 million transistors on a 609 mm² die, yielding a transistor density of 125.3M per mm². The RTX 4000 Ada Generation uses the AD104 die, with 35,800 million transistors on a 294 mm² die, yielding a density of 121.8M per mm². The L40S therefore packs more than double the transistor count and more than double the die area.

The compute resources differ by a factor of roughly three. The L40S has 18,176 shading units, 568 texture mapping units, 192 ROPs, 142 ray tracing cores, and 568 tensor cores. The RTX 4000 Ada Generation has 6,144 shading units, 192 TMUs, 64 ROPs, 48 RT cores, and 192 tensor cores. In every compute category, the L40S has approximately three times the hardware resources. This directly explains the benchmark deltas.

Clock behavior differs in an interesting way. The RTX 4000 Ada Generation has a higher base clock at 1500 MHz versus 1110 MHz on the L40S. However, the L40S has a higher boost clock at 2520 MHz versus 2175 MHz. The lower base clock on the L40S is consistent with a larger die that needs to manage thermal density, while the higher boost clock indicates that under load with sufficient cooling, the L40S can sustain a higher peak frequency. The memory clock is identical at 2250 MHz with 18 Gbps effective on both cards.

The memory subsystem diverges sharply. The L40S ships with 48 GB of GDDR6 on a 384 bit bus, delivering 864.0 GB/s of bandwidth. The RTX 4000 Ada Generation ships with 20 GB of GDDR6 on a 160 bit bus, delivering 360.0 GB/s. The L40S has 2.4 times the memory capacity and 2.4 times the bandwidth. Both use GDDR6 rather than GDDR6X or HBM, which keeps the memory architecture simpler and more cost predictable, but the bus width difference is the dominant factor in the bandwidth gap.

The production status and release cadence also differ. The L40S is marked as end-of-life, having been released on 2022-10-12, with a predecessor of Server Ampere and a successor of Server Hopper. The RTX 4000 Ada Generation is active, released on 2023-08-08, with a predecessor of Workstation Ampere and a successor of Blackwell PRO W. The L40S belongs to the Server Ada generation (Lxx branding), while the RTX 4000 Ada Generation belongs to the Workstation Ada generation (x000A branding). The series field for the RTX 4000 Ada Generation is listed as GeForce 40-series, which is notable given its workstation positioning.

FAQ

Q: Which GPU has the higher average benchmark score?

A: The NVIDIA L40S has an average benchmark score of 295763, while the NVIDIA RTX 4000 Ada Generation has an average score of 135218. The L40S is more than double the RTX 4000 Ada Generation's average.

Q: How much faster is the L40S in OpenCL compute?

A: The L40S scores 330727 in Geekbench OpenCL, versus 146593 for the RTX 4000 Ada Generation. This is a 125.6% advantage for the L40S.

Q: What is the memory capacity difference?

A: The L40S has 48 GB of GDDR6 memory on a 384 bit bus, while the RTX 4000 Ada Generation has 20 GB of GDDR6 on a 160 bit bus. The L40S also has 864.0 GB/s of bandwidth versus 360.0 GB/s.

Q: Are these GPUs on the same architecture?

A: Yes, both are built on the Ada Lovelace architecture and fabricated by TSMC on a 5 nm process. However, they use different dies: the L40S uses AD102, while the RTX 4000 Ada Generation uses AD104.

Q: What is the production status of each card?

A: The L40S is end-of-life, released on 2022-10-12. The RTX 4000 Ada Generation is active, released on 2023-08-08.

Q: How does the RTX 4000 Ada Generation compare to its nearest rivals?

A: The RTX 4000 Ada Generation is effectively at parity with the NVIDIA A10M (0% delta), and within 0.1% to 0.9% of the AMD Radeon PRO W6800, the AMD Radeon Pro W6800X Duo, and the AMD Radeon PRO V620. Its average score of 135218 sits in a very tight competitive cluster.

Specification Differences

The two cards differ across nearly every major specification category. The process node is the same (5 nm, TSMC), and both support PCIe 4.0 x16, use a 1x 16-pin power connector, and support DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4. The display outputs differ: the L40S has 1x HDMI 2.1 and 3x DisplayPort 1.4a, while the RTX 4000 Ada Generation has 4x DisplayPort 1.4a.

The die-level differences are substantial. The L40S has a 76,300 million transistor count on a 609 mm² die, while the RTX 4000 Ada Generation has 35,800 million transistors on a 294 mm² die. The transistor density is slightly higher on the L40S at 125.3M per mm² versus 121.8M per mm².

The compute configuration scales by roughly three times across the board. The L40S has 18,176 shading units, 568 TMUs, 192 ROPs, 142 RT cores, and 568 tensor cores. The RTX 4000 Ada Generation has 6,144 shading units, 192 TMUs, 64 ROPs, 48 RT cores, and 192 tensor cores.

Clock speeds show a mixed picture. The L40S has a base clock of 1110 MHz and a boost clock of 2520 MHz. The RTX 4000 Ada Generation has a base clock of 1500 MHz and a boost clock of 2175 MHz. The memory clock is identical at 2250 MHz with 18 Gbps effective on both.

The memory subsystem is a major differentiator. The L40S has 48 GB of GDDR6 on a 384 bit bus with 864.0 GB/s bandwidth. The RTX 4000 Ada Generation has 20 GB of GDDR6 on a 160 bit bus with 360.0 GB/s bandwidth.

The power and physical specifications diverge strongly. The L40S has a TDP of 300 W, requires a 700 W suggested PSU, is dual-slot, and measures 267 mm in length and 111 mm in height. The RTX 4000 Ada Generation has a TDP of 130 W, requires a 300 W suggested PSU, is single-slot, and measures 245 mm in length and 112 mm in height. The RTX 4000 Ada Generation is shorter, same height, thinner, and draws less than half the power.

The production lifecycle also differs. The L40S is end-of-life with a release date of 2022-10-12, a predecessor of Server Ampere, and a successor of Server Hopper. The RTX 4000 Ada Generation is active with a release date of 2023-08-08, a predecessor of Workstation Ampere, and a successor of Blackwell PRO W.

Head-to-Head Benchmarks

The recorded head-to-head data is limited to two Geekbench tests, and the L40S wins both decisively. The first test, Geekbench OpenCL, shows the L40S at 330727 and the RTX 4000 Ada Generation at 146593. The delta is 125.6%, meaning the L40S is more than double the RTX 4000 Ada Generation's score. This is the larger of the two gaps, and it reflects the raw compute throughput advantage of the L40S's 18,176 shading units versus 6,144, as well as its 864.0 GB/s memory bandwidth versus 360.0 GB/s. OpenCL workloads often scale with both compute unit count and memory bandwidth, so the L40S's advantages in both areas compound.

The second test, Geekbench Vulkan, shows the L40S at 260799 and the RTX 4000 Ada Generation at 123842. The delta is 110.6%, still a dominant win but slightly narrower than the OpenCL gap. Vulkan workloads can be more sensitive to driver scheduling and draw-call overhead, which may explain why the percentage advantage is slightly smaller. Even so, the L40S maintains a lead of more than 2.1 times in absolute score.

Contextualizing these scores against the nearest rivals helps clarify where each card lands. The L40S's average score of 295763 is 3% below the NVIDIA RTX 6000 Ada Generation (287237, note the delta is listed as 3% in favor of the L40S, so the L40S is actually 3% higher), 4.1% above the NVIDIA L40 (284111), 7% below the AMD Instinct MI300X (317994), and 11.7% below the NVIDIA H200 NVL (334891). This places the L40S in a tight grouping with the RTX 6000 Ada Generation and the L40, while trailing the larger memory-pooled accelerators from AMD and NVIDIA's H series.

The RTX 4000 Ada Generation's average score of 135218 is essentially tied with the NVIDIA A10M (135230, 0% delta), 0.1% below the AMD Radeon PRO W6800 (135396), 0.4% below the AMD Radeon Pro W6800X Duo (135774), and 0.9% below the AMD Radeon PRO V620 (136472). The competitive cluster around the RTX 4000 Ada Generation is extremely tight, with all four nearest rivals within 1% of its score. This suggests that in the midrange workstation segment, the RTX 4000 Ada Generation does not have a clear performance edge over its peers; it competes on other attributes such as power draw, slot width, and ecosystem fit.

The head-to-head deltas between the L40S and the RTX 4000 Ada Generation are far larger than the deltas between either card and its nearest rivals. The 125.6% OpenCL gap and 110.6% Vulkan gap dwarf the single-digit percentage differences seen in the rival comparisons. This is consistent with the two cards being aimed at different tiers of the market. The L40S is a server-class accelerator in the top percentile of all GPUs, while the RTX 4000 Ada Generation is a workstation card in the 95th percentile, competing against a tight pack of similarly performing alternatives.

In summary, the data shows a clean hierarchy. The L40S wins both recorded benchmarks by a margin of over 100%, and it sits in a performance tier with other high-end server accelerators. The RTX 4000 Ada Generation wins no benchmarks in this comparison, but it holds its own within its own competitive set, where all rivals are within 1% of its average score. The choice between the two is therefore not a question of which is faster, the L40S is unambiguously faster in every recorded metric, but rather which fits the deployment context. The RTX 4000 Ada Generation's single-slot design, 130 W TDP, and 300 W suggested PSU make it suitable for space- and power-constrained environments, while the L40S's 48 GB memory, 864.0 GB/s bandwidth, and 300 W TDP make it the appropriate choice for compute-heavy server workloads where raw throughput is the priority.

DETAILED SPECIFICATIONS

SPECIFICATION
L40S
RTX 4000 Ada Generation
Core Specs
Shading Units
18,176
6,144 -66.2%
Shaders
18,176
6,144 -66.2%
TMUs
568
192 -66.2%
ROPs
192
64 -66.7%
SM Count
142
48 -66.2%
Clocks
Base Clock
1110 MHz
1500 MHz
Boost Clock
2520 MHz
2175 MHz
Memory Clock
2250 MHz 18 Gbps effective
2250 MHz 18 Gbps effective
Memory
Memory Size
48 GB
20 GB
VRAM (MB)
49,152
20,480 -58.3%
Memory Type
GDDR6
GDDR6
Memory Bus
384 bit
160 bit
Bandwidth
864.0 GB/s
360.0 GB/s
Cache
L1 Cache
128 KB (per SM)
128 KB (per SM)
L2 Cache
48 MB
48 MB
Performance
Pixel Rate
483.8 GPixel/s
139.2 GPixel/s
Texture Rate
1,431.4 GTexel/s
417.6 GTexel/s
FP32 (TFLOPS)
91.61 TFLOPS
26.73 TFLOPS
FP64 (TFLOPS)
1,431.4 GFLOPS (1:64)
417.6 GFLOPS (1:64)
FP16 (TFLOPS)
91.61 TFLOPS (1:1)
26.73 TFLOPS (1:1)
AI/RT
RT Cores
142
48 -66.2%
Tensor Cores
568
192 -66.2%
Power
TDP
300 W
130 W
TDP (W)
300
130 -56.7%
Suggested PSU
700 W
300 W
Power Connectors
1x 16-pin
1x 16-pin
Architecture
Architecture
Ada Lovelace
Ada Lovelace
GPU Name
AD102
AD104
Generation
Server Ada (Lxx)
Workstation Ada (x000A)
Process Size
5 nm
5 nm
Transistors
76,300 million
35,800 million
Die Size
609 mm²
294 mm²
Foundry
TSMC
TSMC
Density
125.3M / mm²
121.8M / mm²
API Support
DirectX
12 Ultimate (12_2)
12 Ultimate (12_2)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
3.0
3.0
CUDA
8.9
8.9
Shader Model
6.8
6.8
Physical
Slot Width
Dual-slot
Single-slot
Length
267 mm 10.5 inches
245 mm 9.6 inches
Height
111 mm 4.4 inches
112 mm 4.4 inches
Outputs
1x HDMI 2.13x DisplayPort 1.4a
4x DisplayPort 1.4a
Bus Interface
PCIe 4.0 x16
PCIe 4.0 x16
Other
Production
End-of-life
Active
Predecessor
Server Ampere
Workstation Ampere
Successor
Server Hopper
Blackwell PRO W
View L40S Details View RTX 4000 Ada Generation Details